//02 Agents and orchestration · Series: The Token Economy
Multi-agent workflows in production: who is actually running thousands of agents (and who is pretending)
49 independent publications in three days say everyone wants multi-agent. The real question: when does it work, and how much does the supervisor cost that nobody measures.
TL;DR. Most teams claiming “multi-agent in production” run 3 to 6 specialized agents with a coordinator. Nobody on the largest industry forum could show thousands of agents working. The real problem is not scale: it is the supervisor LLM that re-sends all accumulated context between each step, spends 69% of tokens on routing, and exposes the system to 14 attack vectors that a single-agent setup does not have. I measured the amplification on a synthetic 4-step pipeline: 3.2x more tokens than the equivalent with a state machine. This week’s decision is whether your coordinator actually needs to be a model.
Context
In three days of September 2026, the lab’s trend engine logged 49 independent publications about multi-agent, from 30 distinct authors, across 4 sources. Velocity was 4.35x the average. The question that started it all appeared on Hacker News: “Multi-agent workflows in production; Where people using 1000s of agents?”
The lab has published on this before. In May, “Building multi-agent systems that don’t collapse” described the base patterns: contract per agent, lightweight coordinator, token budget. This article does not repeat that work. It starts from the question the previous one did not answer: in real production, how many agents do people actually have, how much does the part nobody measures cost, and what new risks appear when you go from one agent to several.
The short answer: nobody showed thousands
The Hacker News thread accumulated dozens of replies. None contained a concrete number of agents in production above roughly a dozen. A founder at a YC F26 startup dedicated to agent swarms asked the question and got back variations of “I’d love examples of it actually working but right now all it’s seemed to be is hype” and “very little organic takeup.”
What the responses did have was something else. One engineer described the real pain point: “most of the pain at scale isn’t the agents themselves, it’s observability.” Another mentioned “OpenTelemetry + Prometheus” and “a massive K8s cluster spinning up pods per N agents.” Nobody complained about having too few agents. Everyone complained about not seeing what the agents do.
The absence of an answer is the answer. Real deployments use 3 to 6 specialized agents with a coordinator, and the hard work is making that legible, not multiplying by a thousand.
The supervisor is the tax nobody measures
The most common architecture in a multi-agent system is the coordinator with a supervisor LLM: a central model receives each specialist’s output, decides who goes next, and passes context along. Intuitive. Also expensive in a way nobody sees on the invoice.
Two studies from September 2026 measure this cost from different angles.
Token waste. An article published on dev.to describes a team that replaced the supervisor LLM with a typed state machine (XState). Reported numbers: median latency from 44.8 seconds to 16.2 seconds, infinite loops from 8.2% to 0%, completion rate from 82% to 94%, and a 71.4% reduction in total token consumption. The reason: every call to the supervisor included all accumulated context from previous steps. With the state machine, transitions are deterministic and consume zero tokens.
These numbers are reported by the author, not independently verified. But the mechanism is reproducible.
Shared memory cost. Singh, Priyam and Bhowmick published the Total Cost of Agency (TCA) concept on arXiv, decomposing multi-agent workflow cost into five components. The key finding: memory-injected tokens represent 13.6% of controllable variable costs. At workflow depth 1, close to zero. At depth 6, 27.6%. Growth is linear (R-squared 0.9974), meaning each additional level adds the same slice. And the good news: reducing retrieval capacity from 32 to 2 entries lowered injected tokens by 28.7% with no measurable accuracy loss.
What I measured
I built two equivalent pipelines for processing a synthetic invoice in four steps: classify, extract line items, validate VAT, format for ERP. Each specialist has a system prompt and returns JSON. The only difference is who decides the next step.
Pipeline A (supervisor LLM). After each specialist, the supervisor receives its own system prompt, the full invoice, and all accumulated outputs up to that point. It decides which agent runs next. Four routing calls plus four work calls.
Pipeline B (state machine). The sequence classify-extract-validate-format is fixed in code. Each specialist receives exactly what it needs. Four work calls, zero routing calls.
Token count (estimated via word*1.3 because tiktoken was not installed in the environment; absolute numbers are approximate, proportions hold):
| Supervisor | State machine | |
|---|---|---|
| Routing tokens | 708 (69.1%) | 0 |
| Work tokens | 317 (30.9%) | 317 (100%) |
| Total | 1,025 | 317 |
| Amplification | 3.23x | 1x |
The supervisor consumes 3.23 times more tokens than the state machine for the same work. And routing cost grows with depth: at each step, the supervisor re-sends the full invoice plus all previous outputs. At depth 6, each routing call carries about 628 tokens of accumulated context alone.
This does not contradict the dev.to article (71.4% reduction). My pipeline has 4 steps; theirs had more steps and larger outputs, which amplifies the effect. The point is the same: most of the cost of a supervisor LLM is not the decision, it is the repetition of context.
The attack surface multi-agent opens
The second cost does not appear on the invoice because it is a risk, not a charge.
Paul and Nandy published in September 2026 (arXiv 2609.22949) the first threat model dedicated to prompt injection in multi-agent systems. They tested 14 attack vectors in 4 groups against a 6-agent system:
- Direct injection via user input (3 vectors)
- Indirect injection via tool outputs (4 vectors)
- Inter-agent injection via message passing (4 vectors)
- Cascading injection via orchestrator manipulation (3 vectors)
Groups 3 and 4 do not exist in a single-agent system. They are a direct consequence of the multi-agent architecture.
Results: 67% of agents vulnerable to at least one scope violation even with system-prompt-level guardrails; baseline injection rate of 31.2%; indirect injection via tool output succeeding in 43% of attempts.
The four defenses they tested reduced the rate from 31.2% to 4.2%:
- Message signing with provenance tracking (91% reduction in inter-agent injection)
- Input/output sanitization at agent boundaries (78% in indirect injection)
- Tool access scoped by role (eliminated privilege escalation)
- Anomaly detection on inter-agent communication (caught 84% of cascading attempts)
None of these defenses is free. Each adds latency, complexity, and tokens to the system. The question is not “should I protect?”; it is “what did the multi-agent architecture give me that justifies this cost?”
The decision: when multi-agent is justified
Three conditions, all required:
1. The specialists need incompatible contexts. If the classifier needs a system prompt that conflicts with the validator’s, they are different agents. If they fit in the same prompt without confusion, they are steps of the same agent.
2. Fan-out is bounded and contracts are typed. Each agent has an input schema and an output schema, and the transition is deterministic or governed by explicit rules. If the only reason for having a supervisor is “the LLM decides,” you are paying 3x for tokens and opening 14 attack vectors for flexibility you probably do not need.
3. Observability is in place before you scale. The consensus in the Hacker News thread was not about how many agents to have. It was about seeing what they do. Without per-agent tracing with a correlation_id, you are debugging blind, and more agents make you more blind, not less.
If all three conditions are not met, a single agent with tools is almost always cheaper, faster, and more secure.
Limits of this analysis
The amplification measurement uses approximate word counting (no tiktoken), on a 4-step pipeline with simulated outputs. In a real pipeline, specialist outputs are larger, which amplifies the effect (in favor of the thesis), but the supervisor could be more selective about the context it re-sends (against the thesis). The dev.to numbers and the TCA paper are from sources with declared interests. The Paul and Nandy paper is a preprint. The 49-publication count from the trend engine reflects conversation volume, not adoption. Nobody in this article, including me, measured a system with thousands of agents in production because, as far as we could determine, nobody has one.
This week’s decision
- Open your multi-agent coordinator and count how many calls are routing (the supervisor deciding the next step) versus work (the specialists producing output). If routing is above 40%, test a typed state machine on the same pipeline.
- For each agent you add, write the contract (input and output schema) and the test that verifies the output meets it. Without a written contract, it is not an agent; it is an expensive while loop.
- If your system has more than 3 agents, implement message signing and sanitization at boundaries. The cost is marginal compared to the risk of a cascading injection that a single-agent system does not have.
Questions this article answers
4- How do you cut 70% of wasted tokens in a multi-agent workflow?
- Replace the supervisor LLM with a typed state machine. The supervisor repeats accumulated context on every routing call. A state machine transitions without calling the model.
- Is anyone actually running thousands of agents in production?
- No concrete answer appeared in the Hacker News thread. Real examples use 3 to 6 specialized agents, not thousands.
- Are multi-agent systems more vulnerable to prompt injection?
- Yes. A September 2026 paper tested 14 vectors across 4 groups and found 67% of agents vulnerable to at least one scope violation, with a 31.2% baseline injection rate.
- What is the hidden cost of shared memory between agents?
- At depth 1, close to zero. At depth 6, memory-injected tokens represent 27.6% of costs. Growth is linear, not exponential.