Introduction
Part 1 of this series, “Why a green test suite can mean nothing”, started from the golden rule I worked to on General Motors’ Super Cruise and Ultra Cruise programs: the person who builds the system must never be the person who verifies it. It then described a new workflow that enforces that rule between two AI agents, by splitting one specification in two and never showing the coding agent the acceptance criteria as part of my new open source Python project qikly, that I developed for spec‑driven test automation at: https://github.com/gal-a/qikly
This part follows a run. One real task, one real failure, the repair that followed, and the mistake I made writing the spec the first time. I come back to how we applied the rule at GM near the end, once the machinery is on the table.
One run, followed to the end
Here is an extract from one real run rather than a description of one. The task from part 1 asks for code that checks following distance: read logged radar samples, reject the invalid ones, compute the time headway to the car ahead, and warn when it drops under two seconds.
The coding agent, working only from the requirements, wrote a range check that kept the upper limit and let a gap of zero through:
One row in the fixture CSV, S005, is a radar sample with a gap of exactly 0 m. The test suite, written from criteria the coding agent never saw, required that row to be rejected. It was not:
That line was everything the coding agent got back: which test failed and what it asserted, with no hint about the rule behind it. The model concluded from it that a gap of exactly zero is not a measurement.
It then produced a PATCH, the diff stage described in Part 1, changing one character:

The whole run, from writing the tests to a passing suite at every stage, took under a minute and fewer than ten model calls on a small model, for a fraction of a cent. What really matters is not the speed. It is that the agent closed the bottom of the range only because a test it had never seen put a zero-metre gap in the rejected pile, not because a criterion told it zero was invalid. Crucially, the coding agent tightened the lower bound only because an unseen test built by an independent test agent forced it to, and not because any written criterion told it directly that zero was invalid.
If it had been given that criterion directly, it would have written the correct check on the first attempt, the suite would have written that exact line on the first try, and that “green” result would have proven nothing, just like a student acing an exam they wrote themselves.
A common mistake: when a decision hides in the acceptance criteria
That first run as explained above actually comes from my second version of that task. The initial first version taught me something else, and it is worth noting because from the outside it might look like editing a spec until it passes.
In the first version, the spec’s requirement said only to warn under “the two-second rule”. Whether a headway of exactly 2.00 seconds raises a warning was settled only in the hidden acceptance criteria.
Now let’s ask the question from earlier again regarding that line:
Given only the requirement, could two competent developers disagree about exactly 2.00 seconds?
Easily. “… breaks the two-second rule” reads as strictly “below” just as naturally as “at or below”. That makes it a decision, and a decision hidden in the test criteria is exactly the mistake the split warns against and that we must avoid.
The first runs showed the cost. Across ten seeds, only three converged. In the other seven, the test‑writing agent resolved the ambiguity the wrong way and wrote a test that warns at exactly 2.00 seconds, contradicting the criterion it had been handed. That is worth pausing on: the test‑writing agent had both halves, the vague requirement and the sharp acceptance criterion, and still went the wrong way. A requirement loose enough to read either way does not just leave the coder guessing, it pulls the test writer off the criterion it was given. And no implementation can pass a test that contradicts its own criterion.
The legitimate fix was to move the decision, and only the decision. The criteria kept every consequence that follows from it, and the limits on speed and gap stayed exactly as they were, including that a gap of zero is rejected. Nothing was weakened and nothing was removed.
With that one sentence moved, on the first sweep the same ten seeds converged ten times out of ten, and the zero‑gap repair above still happened in seven of them, caught by a test the coding agent never saw.
This exemplifies the fine line between fixing a specification and lowering a bar for the acceptance criteria. Moving a decision the coder needs is a correction the spec always needed, and it is justified by the two competent developers question on its own terms, without looking at any code. Deleting the criterion, loosening it, or copying its boundary values into the requirements would also have made the runs pass, but it would have proven nothing.
When the code already exists
Everything above assumes the implementation is written during the run. That is the case where the separation can be enforced rather than promised: one process ran both agents, and the acceptance criteria were stripped from one side’s prompt in code. Most code is not like that. It already exists, written by a person or by an assistant, months ago, and what a team wants is that code checked.
My open-source project allows the same workflow to point at code you already have, and the test suite is still written from the criteria by an agent that never reads the code implementation. What changes is not the suite. It is what each result is worth.
|
Code written during the run |
Code that already existed |
|
|---|---|---|
|
A test fails |
a real finding |
still a real finding |
|
A test passes |
the code satisfied a bar it never saw |
the code agrees with a bar its author may have seen |
A failure is sound either way, because the test was written from the criteria without reading the implementation, so a disagreement is a real disagreement between code and specification. Only the passes weaken, and they weaken for a reason that has nothing to do with the tool: whoever wrote that code may have had the acceptance criteria open in the next window.
That is a claim about a timeline, so it is worth drawing one. The workflow diagram above shows who sees what. This one shows when, and what has to be true before anything else means anything:
For code that already exists, nobody can enforce the verification after the fact. What can be done is to ask version control, because it has been keeping notes:
A run against existing code now asks exactly that, and prints the answer with its bounds attached, because the bounds are most of the truth. A commit date is not a writing date. Code can sit uncommitted in a working tree for weeks, so this order narrows the risk that the bar was reverse engineered from the code without closing it. And criteria committed first say nothing at all about whether the developer was given access to read them before implementing.
So the evidence is worth most in the opposite situation, when it comes out badly. Criteria revised after the code landed is a question somebody should answer. That asymmetry is the same one as the table above, and I would rather ship a check that explicitly says so than one that prints a green tick and lets a reader assume more than it said.
What version control does know, and states plainly, is who committed each side. Two different people is the separation of duties the whole method is borrowed from.
What I’m measuring next
Here is where the evidence stands today, and what comes next.
Across three sweeps totalling 967 runs on the example tasks that ship with the tool, on a small, inexpensive model (gemini-3.5-flash-lite) chosen so the sweeps could be repeated affordably, roughly eight runs in ten produce code that passes every integration and system test, and roughly six in ten pass everything, including the unit stage. A fourth sweep of 390 runs, taken after I hardened the process and in two places made the tasks strictly harder, reproduced the same two figures. I report it beside the others rather than pooled with them: one number describing two different systems is how a result gets withdrawn later. These runs used a cheap model, so read the figures as a performance floor.
When a run does not converge, it exits non-zero, names the tests that blocked it, and ships nothing. The gap between those two figures is mostly the unit stage, which runs last and is the strictest. Two questions cover most of the rest. Near-identical FIX and PATCH every iteration means the agent is guessing a decision nobody told it. And a rule that no row in your sample data can trigger gives you a test that passes whatever the code does, so the fix is in the data and not the code. The repository’s troubleshooting page has the rest. The one thing not to do is loosen a criterion to get green.
The central research questions are whether withholding the acceptance criteria produces test suites that catch more real defects than a suite written with full sight of the code, and whether automatically refining the criteria sharpens the test suite. Both remain open questions for formal proof, but the working hypothesis is yes. Withholding the acceptance criteria from the coding agent is not a hypothesis: it is a property of the pipeline, enforced in code, and easily demonstrated with the tool.
The golden rule of V&V
At GM, I worked in both the Triage and the Functional Performance Analytics (FPA) teams, where verifying a component was kept structurally separate from the development teams who had built it. The engineer who wrote a piece of code was never the person who signed off that it met its requirements.
That separation of duties was the golden rule: the person who builds the system must never be the person who verifies it.
In the FPA team we were particularly concerned with edge cases. An example of a recurring problem class was a perception module that met its requirement exactly as the developer interpreted it, yet missed what we were measuring.
A detection‑accuracy requirement might state that the system had to hold within some bound. The implementing team would read that as accuracy across a logged drive, which seems like a natural interpretation when you are building the thing and watching it work, especially since their own tests agreed with them. Our evaluation was against GT (Ground Truth) for the same data, but applied according to the requirement rather than the developer’s interpretation, and that difference exposed failures their drive‑level accuracy never revealed.
For example, our FPA acceptance criteria asked whether it held on every frame, including the handful where a vehicle was cutting in at the edge of the sensor’s field of view, or some kind of occlusion occurred. Averaged over a drive, such a module passes comfortably. Frame by frame it does not, because thousands of easy frames drown the few dozen that determine whether the system is safe.
The requirement never said whether the bound was a mean or a worst case. The implementers picked the reading their code naturally produced, and they tried to honestly verify it themselves against that reading. In general, developers don’t have the time or resources to check all the edge cases themselves and so they typically pick a handful of important cases, verify the results and quickly move on to their next task.
That’s ODD. The Operational Design Domain.

Defining edge cases accurately and comprehensively is a lot of work, especially if you are trying to figure out all the potential real world failures in something as complex as ADAS (Advanced Driver Assistance Systems). This is referred to as the Operational Design Domain (ODD) and defines the specific operating conditions in ADAS. An independent team needs to do this work thoroughly, and arriving at the optimal test suite is usually an iterative process based on evolving test criteria based on past failures. Since there are so many parameter variations to test, the number of test scenarios increases exponentially and is not practical if we employ a brute force test coverage approach. Only an oracle can generate perfect test coverage across all possible scenarios. Therefore good “Design of Experiments” (DOE) principles need to be followed in order to optimally cover the ODD for the entire test space.
The same problem turns up in miniature in my own tool. When I planted faults to measure whether a sharper bar catches more, most of them were invisible to every test suite, because the acceptance criteria described inputs the fixture data never contained. A criterion nothing can exercise is not a bar. Choosing what the data must contain is the same DOE question, one scale down. If your test drives never include heavy rain, you learn nothing about rain, no matter how good your test team is.
Validation still needs a “human in the loop”
Validation is a different story. There is a limit to the automation process that needs to be stated plainly. Everything above has been about verification: given a written standard, did the code satisfy it, checked by someone who could not see the standard while building the thing it judges. None of it touches validation, which tackles the harder question of whether the standard itself describes what should actually happen.
A tool can enforce that a coding agent never sees the acceptance criteria. No tool can tell you the criteria were the right ones to write. For the time being, that still requires someone who understands the domain, the user, and what the system is actually for. It also requires writing a specific, checkable description of correct behaviour, including the edge cases a careless reading would miss. Of course AI can assist with all of this too, but AI still needs a human expert to oversee that project, approve its specification, and remain in the loop as the specification evolves.
The underlying question
Working on Super Cruise and Ultra Cruise taught me that Verification & Validation are not really two steps in a process. They are a professional mindset about who is allowed to answer which question, and why. I did not expect to find this mindset missing from a tool built by a completely different kind of engineer, one made of parameter weights rather than people, asked to write both an answer and the review that grades it.
The real issue is not simply about trust. It is about the underlying methodology and noticing that a check performed by the person who built the thing being checked is not really a check. That is true whether the coder is a person under deadline pressure or an LLM with a context window and a limited budget, and it will remain true for whoever writes the next generation of software where we should demand a higher level of quality assurance and auditing.
About the Author
Gal Arav is the author of Applied Statistics for Data Science (qikly.com) and maintains his new open‑source project for spec‑driven test automation at:
https://github.com/gal-a/qikly
Readers are encouraged to give it a spin, and any (human) feedback is much appreciated as the codebase continues to evolve.

