Running two AI coding agents as a coordinated team sounds like a productivity unlock. The reality involves hard trade-offs in token costs, debugging transparency, and how much control a developer retains over their own codebase.
A project called OpenRig recently made this concrete. It's an open-source harness that lets developers run Claude Code and Codex together as one system, managed through a single YAML config and a lead agent that coordinates work across specialists. The pitch: stop juggling terminal sessions and start treating AI agents like a team. The catch, which the project's documentation only partly addresses, is that every layer of orchestration adds cost, opacity, and risk that compounds in ways solo-agent setups don't.
This isn't a review of OpenRig, but an examination of what happens when the multi-agent pattern hits real codebases, where the gap between demo and production is wider than most developers expect.
How Multi-Agent Harnesses Actually Work
OpenRig's architecture is straightforward in concept. You define a team of agents in YAML, specifying which models handle which roles. A lead agent receives your high-level intent, delegates tasks to specialists, and surfaces results and decisions that need human attention. As described on the project's GitHub repository, the system requires Node.js 22 or 24 and tmux on macOS or Linux, and it writes provider hooks and workspace trust settings during setup.
The interesting design choice is how it handles permissions. On first launch, the system asks whether agents should be allowed to run commands without repeated permission prompts. The recommended answer is yes. This makes sense for workflow speed, but it also means the developer is granting blanket execution authority to multiple agents simultaneously, each with its own context window, its own interpretation of the task, and its own capacity for error.
This is the core tension. A single agent running Claude Code or Codex alone already requires trust. A multi-agent rig multiplies the trust surface without proportionally increasing visibility into what each agent is doing at any given moment.
The Token Overhead Tax on AI Coding Agent Orchestration
Running multiple agents isn't just conceptually more complex. It's measurably more expensive in ways that aren't obvious from the outside.
Research published by Systima measured what happens at the API boundary when Claude Code processes even simple requests. Their analysis found that Claude Code uses roughly 33,000 tokens of system prompt, tool schemas, and scaffolding before a user's actual prompt even arrives. By comparison, OpenCode, a leaner alternative harness, used about 7,000 tokens for the same baseline.
That overhead gets worse with real-world configurations. A production repository's instruction file (like CLAUDE.md) adds an average of 20,000 tokens per request, Systima found. Five modest MCP servers add another 5,000 to 7,000. Before a developer types a single word, the system is 75,000 to 85,000 tokens deep.
Now multiply that across agents. Systima found that a small task costing 121,000 tokens when handled directly ballooned to 513,000 tokens when fanned out to two subagents, because each subagent re-reads its own system prompt and tools on every turn. In a multi-agent harness running both Claude Code and Codex, you're paying multiple full baselines simultaneously, each with its own cache behavior and billing profile.
Systima's measurements also showed Claude Code re-writing tens of thousands of prompt-cache tokens mid-session, up to 54 times more cache tokens than OpenCode on the same task. Cache writes are billed at a premium. In a multi-agent setup, this behavior compounds across every active agent.
As we explored in our coverage of Claude Opus 5's effort settings, Anthropic has been working to make long-running agent sessions more efficient and sustainable. Opus 5's tunable effort setting and lower per-task costs help. But efficiency gains at the model level can be eaten alive by orchestration overhead at the harness level, especially when multiple agents are each maintaining their own context windows.
When Agents Go Wrong, Complexity Makes It Worse
The stakes of reduced oversight aren't theoretical. TechRadar reported on a developer who lost 48,000 files in 103 seconds when Claude Code, tasked with rebuilding a project mirror, decided the best approach was to create its own cleanup script. The script protected junction-level folders but left nested directories exposed, and the agent removed 55,550 files before posting a message that read, "Craig — stop and read this. I broke something."
The damage extended to .git internals, making recovery impossible. This happened with a single agent operating under a single developer's supervision.
In a multi-agent system, the failure surface expands in two directions. First, more agents means more independent decision-making happening in parallel, each capable of destructive actions. Second, debugging becomes harder because the causal chain crosses agent boundaries. If Agent A modifies a file that Agent B then uses as input for a destructive operation, tracing the failure requires reconstructing the interaction between two separate context windows, each with its own reasoning history.
OpenRig's lead-agent architecture partially addresses this by routing decisions through a coordinator. But the coordinator is itself an AI agent, subject to the same context limitations and reasoning errors as the agents it manages. The developer's role shifts from direct supervisor to reviewer of a coordinator's summaries, adding another layer of abstraction between human judgment and file-system changes.
The Autonomy Gradient
The competitive landscape suggests this pattern will only accelerate. As WIRED reported in its profile of OpenAI's efforts to catch up in coding agents, Claude Code accounts for nearly a fifth of Anthropic's business, representing more than $2.5 billion in annualized revenue. Codex trailed at just over $1 billion in annualized revenue by the end of January, according to a person with direct knowledge cited by WIRED. Both companies are investing heavily in making their agents more capable and more autonomous.
More capable agents are genuinely useful. The question is what happens to developer autonomy as the stack grows. With a single agent, a developer reads the agent's output, reviews diffs, and decides what to commit. With a multi-agent rig, the developer reviews a lead agent's summary of what multiple specialists did, often across multiple files and multiple steps. The information asymmetry between what happened and what the developer sees grows with every additional agent.
This isn't a binary choice between full control and no control. It's a gradient. At one end, you have a developer using an AI autocomplete in their editor, reviewing every suggestion inline. At the other end, you have a multi-agent system running autonomously for hours, surfacing only final results. Most real usage falls somewhere in between, but multi-agent harnesses push developers further toward the autonomous end of the spectrum whether they intend it or not.
What Gets Lost
Several specific things degrade as agent count increases. Diff review quality drops because the volume of changes exceeds what a developer can meaningfully inspect. Context awareness declines because no single agent, including the coordinator, holds the full picture of the codebase. Error attribution becomes ambiguous when multiple agents touch the same files. And rollback becomes harder when changes from different agents are interleaved in the commit history.
Where This Leaves Developers
Multi-agent coding systems solve a real problem. Complex projects have multiple workstreams, and coordinating AI agents across them is a natural extension of how development teams already operate. But the analogy to human teams breaks down in important ways. Human team members share institutional knowledge, communicate in natural language with full context, and can exercise judgment about when to stop. AI agents share none of these properties reliably.
The practical path forward likely involves better tooling for the gaps that multi-agent setups expose:
- Granular per-agent audit logs
- Automatic diff segmentation by agent
- Cost projections before fan-out
- Hard limits on destructive operations regardless of permission settings
OpenRig and projects like it are useful precisely because they make the multi-agent pattern concrete enough to study. The pattern's costs are real, and developers adopting it should understand exactly what they're trading away.