At Straight Up AI we’ve built an internal control plane for orchestrating coding agents. How much latitude agents are given is driven by three factors:
-
Blast radius of a mistake. Mature production systems often have offline consequences whereas 0-1 MVPs do not.
-
Project context. Legacy projects have less embedded knowledge that coding agents make probabilistic judgements on.
-
Project maturity. Greenfield projects are more likely to follow established design patterns that coding agents can follow more easily.
This list is missing one factor; the cost of building. We’re a small consultancy. When not using client allocations we run the control plane through a Claude Max account. As this is priced at £200 monthly, the marginal cost of bad decision making is time.
That is a fine place to prove a control plane works. Developing it further without reviewing its sustainability is dangerous. Nothing tells us whether we have built something that we could genuinely afford to run. So I went and analysed seven weeks of the control plane in action.
Our Utilisation
Between 15 July and 4 September we analysed 44 development cycles across our portfolio.
To calculate the API equivalent bill I converted every token to input-token equivalents with cached reads at 0.1x, cached writes at 1.25x, and output tokens at 5x.
At API rates we would be paying 22 times more for our utilisation. As our consultancy grows this quickly becomes untenable, and is almost the cost of a mid engineer’s salary.
The second thing bothering me was time rather than money. Several parts of a run felt unnecessarily slow. My finger pointed at our usage of adversarial review, which spins up a second coding agent to attack every commit/increment. That intuition turned out to be roughly right, though not for the reason I assumed.
|
Adversarial review |
Implementation |
Ratio |
|
|
Agents dispatched |
242 |
204 |
1.19x |
|
Cost units |
234.8M |
339.6M |
69% |
|
Agent-hours |
28.0 |
39.2 |
71% |
|
Median per agent |
825k units, 5.6 min |
1.12M units, 7.8 min |
0.74x, 0.72x |
The adversarial review overhead is coming from the volume of reviewers dispatched. It is not one singular, expensive review agent. We submit more reviewers than implementers with 26% of them demanding changes. When changes are requested a remediation agent is dispatched to implement the fixes, which spins up another adversarial review cycle. Similar to regular code review, if a solution is not found by the second or third round more code isn’t the solution, and instead more structural changes are required. Adversarial review in this context risks rabbit-holing rather than taking a step back and considering what changes may be required that aren’t necessarily in the commit scope.

The Control Plane

A run has one controller and many workers. The controller is the session you talk to. It’s job is to:
-
Read the build plan and increments
-
Dispatch one worker per increment
-
Validate the increment state (completed, reviewed, failed)
-
Dispatch the next task
It is the only agent that lives throughout the entire job. This architecture addresses a limitation with plan and build setups. When a worker both narrates progress and advances a plan, incorrect progress statements cascade to incorrect increments. The controller architecture stops this by acting as an independent arbiter of the truthfulness of the implementer. Our execute-plan contract states the rule once and every variant inherits it:
The controller is also required to stay awake for the duration, which matters a great deal to the cost figures further down:
Once an increment is implemented and verified, it goes to adversarial review.
A fresh agent receives the repository root, the diff scope (head vs base commit) and is told to explicitly find failure scenarios. To ensure a fair test the reviewer cannot edit code; if it thinks there is a logic or type error it must generate a test to prove it. The end is an adversarial report which is then fed into a remediator agent which implements the fixes.
Adversarial review is a blunt instrument by design. The decision to hand it to an expensive coding agent (Fable 5 high), is also intentional. It acts as a gate keeper for all code that could be progressed. A cheaper review with lower recall will cascade failures throughout the system. And in scoping the review to only the associated diffs we limit the review to only the code that has changed. This ensures the review agent does not explore transient issues or known/accepted project risks.
Where The Money Goes
The most expensive component in the system is the one that writes no code. The controller costs more than implementation, planning and review combined.

This is where I expected the investigation to end quickly. We suspected context over long sessions would be a problem and designed the system to include per-commit compaction. In reality however this was not happening.
|
Compaction events across all 44 controllers |
34 |
|
Controller API calls |
25,878 |
|
Calls per compaction |
~760 |
|
Controllers that compacted at all |
12 of 44 |
|
Largest context observed |
996,659 tokens |
Across 75% of runs the controllers were never compacted. And when we analysed the context windows of the controllers we found that context increased (more or less) linearly with session duration. This points to our intuition that compaction would be needed to maintain context, but it was not being implemented reliably.

|
API calls |
Share of calls |
Cache-read tokens |
Share |
Median context per call |
|
|
Controllers |
25,878 |
30.7% |
9.31bn |
58.7% |
312k |
|
Workers |
58,319 |
69.3% |
6.54bn |
41.3% |
101k |
The cost of our system was not derived from implementation or review. It was the re-reading and maintenance of the session state which increased on every turn.
In reviewing the system we found aggressive compaction, being implemented in the wrong place. The individual workers were compacting as they opened on average at 38,000 tokens and closed at 103,000 tokens. The controller was not, and ended up carrying the state of every worker in memory. This was a problem not just from a cost perspective. It also posed a problem for the reliability of the system as the degradation of performance as context windows saturate is a known limitation of LLMs.
An Honest Look At The Controller
From both a cost and reliability perspective, we needed to know what was going on. We started off with the planning system, which was the most obvious target.
We plan in layers. An architecture specification becomes a detailed technical specification, which becomes an implementation scope and a build plan of numbered tasks. The controller reads all of it. Although this chain provides a granular execution and evaluation path, we thought maintaining this ledger across the entire run conflicted with compaction.
On measuring the documents against an average controller increment of 312,000 tokens we found specification stacks to be 2% of the overhead. This was the wrong path. It is also the single most reliable technique for keeping LLMs on task and inventing solutions and implementation paths. A lot of time is spent reviewing these specifications prior to dispatch, and having these scopes make it clear what we are checking the agents against.
|
Artifact |
Median length |
Approximate tokens |
|
Architecture specification |
3,304 words |
4,600 |
|
Implementation scope |
1,322 words |
1,850 |
|
Build plan |
1,041 words |
1,460 |
What actually fills the controller is its own output.

The largest single component of a controller’s context is its tool arguments. And a quarter of those were the dispatch briefs used for sending work to agents. Across the corpus there are 698 of them, with a median length of about 1,390 tokens and a longest of over 6,000.
Scoped planning is paid for here once, and is marginal. The brief derived from the plan is written for every increment. The length is necessitated as it contains the task overview, acceptance criteria, base commit and the constraints.
The trouble is that these were persisted. On every subsequent increment these briefs are re-read until the run ends. This is not needed as by design, the briefs are scoped only to the increments they are being measured against. Therefore our system was ingesting a slowly incrementing stack of briefs rather than treating the briefs as ephemeral and marking completion against the build plans.
In reviewing our internal contract this was made worse in two places. These are deliberate design decisions that are fine in isolation. The first is the controller must validate every worker result itself. This means that all state must be evaluated by the controller, necessitating the evidence being persisted to the controller context.
The second is task tracking. We request the controller to re-render the task list on every state agent for easy human review. The task list therefore also needs to be maintained. This, alongside task metadata, naturally grows across longer sessions.
Post the full list at run start… whenever any task changes state (batched: one refreshed list per chunk of work, not per tool call), and in the final message of every turn that leaves work outstanding.
Combined, these decisions result in a coding agent that accumulates metadata that is not shed, that increases on every increment completed by one of the workers.
Having Our Cake and Eating It?

The controller is full of things it’s already finished with. Three changes would ensure they are not re-read on every increment.
Pass briefs by reference, not by value.
If you’ve done some C++/low level language programming you will be familiar with this concept. Objects that are expensive to create are not passed from one function to another. Instead, we pass pointers or references to them so that they can be reutilized downstream.
The run record exists in a persistent state that is incremented on every turn pass. The controller does not need to know about the entire record. It just needs to know the most recent state. Applied recursively this guarantees that all previous states are in the appropriate state before a decision is made. This ensures the controller only needs to hold a small amount of metadata and a pointer to the run state.
Compact at increment boundaries.
Once an increment reaches a terminal state, nothing in its implementation transcript changes the next decision.
The controller can rehydrate from the build plan and run record, which are the source of truth. This reduces context overhead from 65,000 to 487,000 tokens into a sawtooth that resets every increment.
Across the corpus’s roughly sixteen dispatches per session, it would put a typical controller turn near 100,000 tokens rather than 360,000. I want to be plain that this is a projection from what I measured, not a result I have run.
Scope the reviewer by attempt, not by increment.
The first review of an increment should see the whole change. The second, after remediation, should see the remediation diff and the findings still open against it. Like an effective human code reviewer, this ensures that an increasingly narrow set of in scope items go through the review funnel. In our system it will reduce the ratio between reviewers and executors.
Conclusion
We’ve developed a system that has become critical to both client delivery and internal work. The system however is not affordable, and leans on the discrepancy between Claude Max and API pricing, which could change quickly.
So far we’ve been focused on whether the system delivers. This is the first time we’ve asked whether we can sustain it. If we address the context bloat we think we can. The architecture is generally sound. The problems are derived from holding increment state in the global run context. Fixing this reduces both the cost and the propensity for incorrect path following.
Our hypothesis on the bloat was also completely wrong, twice. We believed we were compacting aggressively at every increment. We then assumed the planning chain was the weight, since the controller reads an architecture spec, a technical spec, an implementation scope and a build plan on every run. That entire stack is under 8,000 tokens, roughly 2% of a median controller turn.
What actually filled the controller was the material it generated itself. Tool-call arguments are 37.7% of everything it accumulates, and a quarter of those are the 698 dispatch briefs it wrote and never put down. The exercise has demonstrated the importance of completing analysis without priors, and being honest about your work.

