Just interested in the code for this project? Find it here.
I do not think it’s an overstatement to say that the worlds of software engineering, data science, and analysis are in the process of rapid and dramatic transformation due to AI tools. This transformation crept up on the community over the course of 2025 and has really exploded since the beginning of 2026 with the widespread adoption of coding harnesses like Claude Code, Codex, and Antigravity. These are agentic tools, making multiple Large Language Model (LLM) calls per request to reason through and solve complex problems. And they are amazing, if a little surreal, to watch as they go about their work. Similarly, the field of business analytics is seeing a shift towards AI analysts, which can explore complex datasets, write and issue SQL queries and generate polished reports complete with figures and recommendations. In summary, AI in the technology workplace is clearly transitioning from a productivity enhancing add-on to an indispensable tool.
This transformation is taking developers a step away from the inner workings of the software or queries that they write, ideally freeing up bandwidth for them to tackle higher level problems of design. For this to work well, we have to be able to trust that the AI is reliable — or at least have some way of measuring its reliability so that we can select an appropriate setting for the task at hand.
Why is reliability so important? AI agents are powered by LLMs, which are non-deterministic token generators. This non-determinism becomes unpredictable in systems with long reasoning chains, multiple LLM calls and/or usage of different models with settings that are not visible to the user. On the flip side, non-determinism is very useful — it’s what gives AI creativity, ability to reason through problems and adaptability. So it’s not necessarily bad, it’s just important to understand and measure its effect on the consistency of responses.
Consistency is defined as the ability of a system to generate reproducible outputs from identical inputs. Measuring consistency is especially critical for tasks where there is a correct answer and deviations from that could be misleading. Business analytics is a good example, but coding is too — there may be many valid ways of reaching a solution, but that solution should actually work as designed, or solve the problem that it was intended for. In real-world deployments, we rarely have ground-truth unit tests to verify an agent’s work, and different LLM-problem combinations show different consistency characteristics. How, then, do we evaluate reliability when we don’t know the answer ahead of time?
In this article we explore this problem by building a command line tool called Coding Agent Consistency (cca), which allows us to set up experiments where we call models multiple times and analyze the distribution of results. This is helpful because in the absence of ground truth, multi-sample consistency might be the only proxy for model reliability that we have.
Here we focus on well constrained coding problems and small models to keep cost manageable, but the concept is extensible to many other use cases. The cca package also enables us to explore different foundation models via API thanks to litellm, and can also interface with local models via ollama as well as coding agents via Omnigent. It therefore offers a tool to compare the consistency of multiple AI tools on a given problem or set of problems.
1. The problem of non-determinism
At their core, generative LLMs are next-token predictors. They generate output in a sequence of autoregressive steps, with each step being basically a classification problem where the objective is to choose the token with a high probability of being “correct”.
This means that LLMs cannot conceive of an entire sentence, code block or query in a single “thought”. At inference time their only objective is just to produce the next token: At a very high level, self attention is used to understand the contextual relationship between all the tokens in the prompt. This contextual representation is then passed through many feed-forward layers, where it is enriched with “learnings” from the billions of weights that were set during training. The final output is a set of scores (logits), one for each possible token. These are then converted into a probability distribution via softmax, and a sampling strategy is used to actually choose that token. The most straightforward strategy is simply to choose the token with the highest probability (greedy sampling), but in practice the token is selected at random from the distribution according to its probability. Techniques such as top-k or top-p sampling aim to truncate this distribution before it is sampled from, but the key concept is the same: We obtain a probability distribution and then sample from it to get our token. Then the procedure repeats to get the next token. It follows that a small change at the beginning of the sequence can have a big impact on what comes later!
The temperature parameter is critical here. Temperature is a scaling factor applied to the logits before the softmax turns them into probabilities, and it changes the shape of the distribution. As temperature approaches zero, the probability of the top-ranked token approaches 100%, which is why setting temperature = 0 is functionally the same as greedy sampling. As temperature increases beyond 1 the distribution becomes flatter, increasing the chance that the model will select a less expected token and therefore increasing its variability/creativity but decreasing its consistency.
It’s important to note that setting temperature to zero doesn’t guarantee absolute consistency in terms of the sequence generated. This is because of tiny rounding errors in the inference calculations as well as batching and parallelism effects (a lot more detail on this can be found in this excellent blog post). Also, and perhaps more importantly today as the industry shifts towards more complex models, neither reasoning models such as GPT5+ and Claude Opus nor coding harnesses like Claude Code, Codex or Antigravity even expose temperature or sampling controls to developers. This is because reasoning models need carefully selected (and non-zero) temperature and sampling settings to work well, and harnesses likely adjust these settings on the fly depending on the question and content being generated. All this is to say that as the models become more powerful, it is becoming even more difficult to understand or influence the characteristics of their non-determinism.
This behavior has benefits and drawbacks. It is what gives LLMs creativity and the ability to come up with new ideas. Human beings — whose activities have generated the training data for these systems — are also inconsistent, and LLM architecture is at its core inspired by biological brains that are the source of this inconsistency. But it does raise the specter of trustworthiness and reproducibility, especially when it comes to solving problems that have a clear “right answer”.
The field of business analytics presents many of these problems, and is one where AI is increasingly being used for writing the code and queries used to answer data questions. Imagine an executive at a retailer asking an AI assistant to provide a report about how customer LTV is distributed by age group. This is a somewhat vague question which could be interpreted in different ways, and requires the system to have access to business definitions so that the answer is presented in the way that the executive expects.
A trustworthy system would need to consistently write correct code, meaning that no matter how many times the executive asked their question (or semantically equivalent variation of it), the system would always produce exactly the same numbers. The exact words the agent uses to present the results are not a big issue here — those could vary without causing harm — but whatever code is written should produce exactly the same numerical results every time the question is asked. To do this with LLMs is challenging.
Of course, human analysts can be inconsistent and make mistakes, but the difference between human and AI assistant results is auditability. If an executive asks a business question to a human analyst, that analyst would presumably save their procedure and, after review, that would become the procedure of record for their team to answer that question. Thus even if the procedure produces the wrong answer, it is at least consistent and therefore not confusing. Conversely, asking an AI agent an analytics or coding question 10 times is more similar to asking 10 independent — or, depending on how conversation history and memory is handled — semi-independent analysts the question. If the question is complex or vague and temperature is non-zero (which we know is for reasoning models), we might expect to get different answers.
We should also expect there to be a relationship between consistency and accuracy. In the best case scenario, we might be able to use this to assess reliability in the case where we have no ground truth. For example, for a certain class of problem, let’s say we observe a model or coding harness writing different code each time we ask the question, but the code always generates the same result and that result is correct. Now suppose we have a new problem of the same class but without ground truth. If we observe the same pattern (i.e. running multiple times generates different solutions but they all evaluate to the same answer), then out confidence in the correctness of that answer would be higher. This article explores a framework for doing just this.
The measurement of LLM consistency and debate about about the pros and cons of non-determinism has a rich literature, with many academic studies examining the diversity of generates solutions under repeated trials and propsing metrics to quantify this. Some notable recent examples are here, here and here. I’m using similar ideas in this post, by using a simple combination of implementation and behavioral consistency as a two-dimensional way of looking at model behavior
2. Goal of this article
Here we are going to explore two dimensions of consistency measurement: Syntax consistency and output consistency. Syntax consistency focuses on the sequence of tokens that the model produces without considering what they do. In our business analytics case this would be the SQL queries or Python code generated, but in a chatbot use case it could also be the text.
A system with high syntax consistency would produce very similar sequences of tokens every time, as measured by some non-semantic metric. Output consistency focuses on what the generated result actually does. For example what numerical values the SQL queries generate, what the code generates when it executes or what the meaning or semantic classification of some text is. Intuitively, a problem-model combination with high output consistency and low syntax consistency might be a lot more reliable than one with low output consistency and high syntax consistency. Furthermore, if we have ground truth, we can test this intuition by looking at the relationship between the proportion of correct answers, syntax consistency and output consistency. To keep the scope manageable we’re just looking at Python problems inspired by the Mostly Basic Python Problems Dataset (Mbpp) here, but the framework should be easily extensible to other types of problems too.
The central idea is that if we run the same problem through an LLM-powered tool — be that a single model call or an agentic system such as a coding harness — multiple times, the results should fall into one of the following quadrants in the consistency space, shown in figure 1.
-
Quadrant 1: High syntax consistency, low output consistency: This is a problematic result because it suggests syntactically similar solutions whose variations are critical for output behavior. Results in this class would have similar implementations but materially different behavior, making them unreliable and possibly very sensitive to prompt changes.
-
Quadrant 2: High syntax consistency, high output consistency: The model generates solutions that are behaviorally and syntactically similar. In the best case scenario, this stability happens because the model is consistently following the optimal approach to arrive at the correct answer. In the worst case scenario the model is confidently wrong and/or is following a memorized approach learned durning its training.
-
Quadrant 3: Low syntax consistency, low output consistency: This is the class of problems about which the model is inconsistent: It tries various approaches and these approaches generate different results. By chance some proportion of these results may be correct, but the behavior is unreliable.
-
Quadrant 4: Low syntax consistency, high output consistency: The model generates syntactically variable solutions that are behaviorally consistent. This is evidence that the model is not using a memorized approach and has robust enough understanding of the problem that it can take multiple approaches and get the same solution.
Intuitively, for problems that have a single correct answer Quadrants 2 and 4 represent desirable behavior. My hypothesis going into this work was that Quadrant 4 is the most desirable outcome, since getting the same result from multiple different approaches should build confidence that the result is actually correct. Therefore, it may be possible to use this approach as a proxy for estimating reliability in the absence of ground truth.
Furthermore, it might be expected that varying the temperature setting would change the classification of a given problem. At low temperatures, we might expect to see solutions falling into Quadrants 1 and 2, while as temperature increases we should expect representation from Quadrants 3 or 4. For a given problem, it might even be possible to use this concept to estimate an “optimal” temperature setting by plotting its trajectory across the consistency space as temperature changes. This will be explored further in section 5.
To explore these ideas further, we need a simple harness that allows us to define the problems to be solved, the models to test and the analysis to generate. We need to be able to then run each problem multiple times through each model, measure consistency, save the results, conduct and present the analysis as an experiment. The harness should be model agnostic, allowing us to test both direct LLM calls and coding harnesses. The recent Omnigent project provides a way of easily switching between harnesses, so we incorporate that too.
3. How do we measure consistency?
In order to plot results in a 2D consistency space as shown in figure 1, we need a way of measuring syntax and output consistency in a way that creates a numerical score between zero and 1.
For our example Python problems, syntax consistency is probably the most challenging because we need to find a way of converting a group of solutions in code into an interpretable similar value. There are a number of ways of doing this, and for the coding problems here we choose a simple approach that involves canonicalization and Levenshtein distance. We also use clustering in an attempt to create groups of solutions, from which representative samples can be shown to the user so that they can understand the range of solutions generated.
We first need to make sure that the code is stripped from the model output, which may contain other tokens or text-based commentary. Next, we need to strip the code of non-algorithmic sources of variation such as docstrings, comments and whitespace. This is done by parsing the AST tree using the following code:
Optionally we might also want to give any local parameters and arguments generic names like arg_0, var_0 etc in order to prevent those from influencing the variation score (this is turned off by default in the package). Having stripped the non-algorithmic variation, we then proceed to compute the levenshtein (edit) distance between all pairs and compute syntax consistency as follows:
Agglomerative clustering with a sensible threshold (0.2 by default) can then be used to group the solutions into representative clusters, samples from which can be chosen to give a sense of the type of variability being seen.
For code, output consistency is more straightforward to calculate because we can execute the code and measure the distribution of results. For each problem, we define a set of test inputs. Then for each generated solution, we run it with those test inputs to generate a set of results known as the solution’s “fingerprint”. Output consistency then becomes a matter of understanding the spread of those fingerprints, which we can measure by clustering. Two solutions are placed in the same cluster if their fingerprints match across all test inputs, and at the end of this process we have one or more clusters. We then scale this cluster count to a consistency value 0 and 1 using normalized Shannon entropy, which is a standard measure of uncertainty. The essence of the algorithm looks like this:
This will return an output consistency of 1 if the results form a single cluster, and a value of zero if each generated solution generates a different fingerprint. It is very important to note that output consistency measure is only as good as the input test suite being used to generate it: For code, this metric only works for problems that have a comprehensive input suite that actually tests edge cases.
To see this in action, let’s consider a very simple Python problem, Mbpp #404.
What happens when we run it through a model 20 times with a temperature of 1? We’re using nemotron-30b from fireworks.ai here just as an example. To get a detailed log of the model’s progress, we can use the following try_runner.py script from the cca package:
Even with this really simple problem, we get 4 syntactically different clusters.
Technically this is a Q4 result, since it has (trivial) syntactical differences that create a syntax consistency score of 0.72 but they all evaluate to the same result, giving an output consistency of 1.
The choice of syntax and output consistency metrics is essential for the results of this analysis to be useful, and will differ by problem domain. For example, in a SQL-generating model, syntax consistency could still broadly follow the approach here, but output consistency might involve identifying the key metric(s) that the user wanted to see and evaluating whether their values were equivalent across the output tables. Equivalent results could then be clustered and normalized entropy used to convert those clusters into a score. Care would have to be taken to avoid penalizing the model for SQL queries that generate the same result in a different tabular format, or with different numbers of rows or columns. This, in my experience, is surprisingly tricky and perhaps deserves a blog post in its own right!
For a text-generating chatbot result, syntax consistency could be measured with a programmatic score like ROUGE or BLEU and output consistency would likely need a sentiment or topic classification approach, followed by clustering.
With these consistency concepts in hand, let’s now explore the cca package itself and how we might use it to start testing the quadrant hypothesis in figure 1.
4. Introducing the cca package
The repository associated with this article contains a basic Python toolkit for configuring and running consistency experiments. It was built with lots of assistance from Claude Code and Google Antigravity, and as such already has a nice README that you can point your coding assistant of choice at to understand it better. In this section we’ll look at its main functionality and how to use it. Note that LiteLLM standardizes the API schema across all underlying providers, simplifying multi-provider integration.
The main objective of the tool is to run consistency experiments. Experiments are defined as .yml files in the config folder, and let’s look at one as an example.
This describes an experiment where we compare the consistency of four OpenAI models (with temperature = 1) using 10 problems from the Mbpp dataset. The problems themselves are stored in here and more can be downloaded using the package. We’re running 20 samples per problem, saving the output to coding_agent_consistency/results/gpt_study and choosing a threshold of 0.8 when plotting the quadrants. This is an arbirary choice and just means that the dividing line between “low” and “high” consistency on both the output and syntax axes is set at that value. Adding also_median_split as True will further calculate the median scores across all runs and allow for use of those as the thresholds when plotting too.
To get started with running experiments, copy .env.example into a new .env file in the repo and paste your chosen model provider API keys in there. The package currently works with Gemini, OpenAI, Anthropic, Fireworks and Groq. To use it with coding harnesses you will also need to install Omnigent.
Our experiments can get expensive, so we should first estimate the token costs.
Next, we kick off the experiments with uv run cca run configs/gpt_series.yaml and a progress bar should track the runs.
To generate a consistency analysis we then run uv run cca analyze configs/gpt_series.yaml which will print a summary of the results and write the full table to results/gpt_study/results.csv.
Finally we run uv run cca plot configs/gpt_series.yaml to plot and save the results. Our GPT experiment produces the following plot

When I ran this for the first time, I was surprised that the pass rate wasn’t closer to 100% for all of these relatively simple Python problems with these powerful models. Clearly there is a large group that achieves 100% pass rate (spread over a range of syntax consistency, which will likely widen if we increase temperature), but by definition any solution suite whose output consistency is not 100% must have contained some wrong answers and there are many examples of this. There are also several problems with poor pass rates.
Some of the errors are explained by the model sometimes not writing the exact function name specified in the instructions. For example mbpp/126 asks for a function to find the sum of common divisors of two numbers and implies that function should be called “sum”. Understandably, the models don’t always comply, preferring objectively better names like “sum_common_divisors”. But this causes the output consistency checker to fail since it’s looking for a function with a particular name. We could of course try to make this more lenient or try to correct this behavior with prompting, but even this is interesting because it tells us about the reliability of the instruction following abilities of these models at a constant temperature setting.
Only gpt-4o-mini and gpt-5.4-nano have this function-name issue with mbpp/126, but the other two models still get the answer wrong in 4 of their 20 runs. gpt-5.6-luna makes the mistake of writing this in 20% of its runs.
Which is mathematically correct but fails because the inner sum calls Python’s built-in accumulator rather than the “sum” function recursively as is intended.
There is a surprising variety of the syntactical consistency of these solutions too, and we’ll return to some more examples in the next section. It’s fair to say that the quadrant plots such as figure 2 are a useful visualization of consistency but still hide important details that can only be found by carefully comparing the generated results.
Let’s leave analysis for now and continue with functionality. One natural question at this stage might be about how the models are prompted, because that will almost certainly affect the results. The only prompt provided here is a single user instruction defined in runners/llm.py .
This is intentional so as to not overcomplicate the experiments, but it would be very interesting to see how the consistency results vary if more specific system instructions like “you are a professional software engineer who consistently writes excellent Python code” were added.
It’s also easy to add new datasets with novel questions, both with and without ground truth. The datasets are just jsonl files with one problem per line in the following format.
Where harness_mode is either “function_call” or “stdio” (for use with Omnigent) and the reference contains the ground truth coded solution, if available. This is run on the values in the input bank to produce the ground truth output. Having ground truth is helpful for analysis purposes, but of course the main benefit of this consistentcy approach is that it can be applied in the absence of a reference. For those cases, we just leave reference as an empty string.
For an experience with more detailed printout of what’s going on for each problem, including cluster assignment from the consistency analysis, you can also use the script /scripts/try_runner.py or refer to the notebooks in coding_agent_consistency/notebooks for more functionality and examples.
What if we want to use this with a coding agent instead of a single model call? Omnigent is a powerful, open source tool which allows users to easily switch between coding harnesses. We can take advantage of this common interface for to extend cca to those too, allowing each “sample” to be a full run where the agent works in a temporary working directory and we read back any code it generated for analysis. As a caveat, this part has not been tested on questions beyond leetcode-style Python problems that only need a single function response. It would be possible to extend the consistency scoring to entire codebases, but this would be expensive to run and has not been implemented here.
To run an experiment with Omnigent calling Claude Code to solve a problem with ground truth, you can follow this example:
Unlabelled_hard_subset contains a single dynamic programming problem with input examples but no ground truth, which was invented by Claude Code. Let’s see if Claude Code can solve it consistently.
When running experiments with Omnigent it is possible to track their progress in the Omnigent app, which will show you the code being written and the model’s explanations. 20 attempts at this DP problem reveal that Claude is applying the same logic (which looks correct; Claude is a better coder than me and this problem is difficult, so I defer to Gemini for a logic check), but the syntax, documentation and organization of the code vary considerably, thus producing low syntax consistency.

Claude Code’s result falls squarely in the robust behavior quadrant, which is driven by the variety in its choice of syntax combined with an output consistency of 0.93. Looking at the output solutions themselves, we see that about 10% of the time the solutions don’t generate consistent results with all members of the input suite. As a single datapoint this is not very helpful, but if this was a very important problem we could run it across a suite of models, harness versions or reasoning effort levels. A picture of how these choices affect the consistency of solutions to this particular question would emerge.
5. Exploration of model-problem combinations
With the concept and package introduced, let’s proceed to look briefly at some interesting results. This will serve to give some indication of what the consistency quadrants method can and cannot tell us about model performance.
First, let’s try a small local model — mistral7b — on the mbpp problems with 20 samples per problem and a temperature of 1. You find the experiment configuration for that here. This might be expected to give quite a spread of results across the quadrants and provide some indication of how suitable these small models are for Python coding out of the box

Once again the quadrent threshold is set at 0.8, which admittedly is an arbitary choice and would need calibration if this method was to be run outside this toy problem setting.
From the plot, we can clearly see that there are some instruction following issues here that cause some solutions to have incorrect function names and therefore low pass rates because the output check just doesn’t run. But there is also some more interesting signal — a group of problems that have high (but not perfect) syntax consistency, perfect output consistency and 100% pass rate, and a group that has moderate syntax consistency, very low output consistency and poor pass rates. This is broadly supportive of the hypothesis that “good” answers lie in quadrants 2 and 4, but there is a lot of nuance here to continue exploring.

When we overlay a larger model’s results with these same problems (figure 5), we do also see a shift towards the upper right corresponding with an overall increase in correctness. It’s no surprise that larger models do better on these Python problems in terms of consistency and correctness, though it is important to note that since these problems are well known and freely available online, all models have likely seen these problems and their solutions multiple times during training. This is why I am still surprised to see instances of the models still getting the wrong answer.
What happens when we try with some genuinely novel problems that the models would not have seen? I asked Claude Code to generate five novel, leetcode style Python problems with inputs and solutions in the format of mbpp. These problems were then tested with nemotron-30b at a temperature of 1 and 20 samples per problem, with the results shown below.

As a final experiment, let’s see what happens when we run the same problem through the consistency framework at different temperatures. We’ll try this for mbpp/126 with nemotron-30b and 20 samples, and I have seen similar behavior with other problems too. None of the results are good, but temperature makes a big difference. As expected, syntax consistency decreases as temperature increases (and it does not start at 100% with temperature = 0!). Output consistency also initially decreases but there is a corresponding increase in pass rate, which might also be expected since we start from a pass rate of zero at temperature = 0, so more variability can only bring benefits here.
What’s interesting is that pass rate starts to decrease as temperature climbs above 0.5, and output consistency then jumps back up at high temperatures. This is an artifact of the fact that some solutions become invalid at temperatures of 1.5 and above and these fail with syntax errors, giving an artificially high output consistency. This actually points to a limitation in th output consistency metric which could arguably be fixed though proper accounting for unrunnable code, which the current framework doesn’t really do but could easily incorporate.
However, the general behavior points to the possibility that there may be an optimal temperature for any given model-problem combination, although exploring this further would start to stretch beyond my budget! With complex problems and large models or coding harnesses, this consistency evaluation work does become expensive so targeted evaluation is essential.


6. Final thoughts and learnings
Thank you for making it to the end! This article has been an interesting journey and the key take-away is that consistency is an important and seemingly unsolved problem in the world of generative AI. Ironically, the article itself would not have been possible without the powerful AI tools that allowed me to write the code in a reasonable time window. Clearly these tools are achieving widespread adoption and generating real productivity gains regardless of any consistency issues, and research continues at rapid pace to understand and improve this. A great example is the recent advent of “System One” models like Jev, which are still transformer-based models but generate output directly rather than token by token. These models don’t write code and are still not completely deterministic, but their consistency is considerably higher than regular LLMs at the classification tasks for which they were designed. As such “decision models” become more widespread inside the workings of coding harnesses, we could potentially use the quadrant method discussed here to assess the impact of these changes.
Although simple, the consistency quadrant method appears to give valuable insights into model behavior, especially for problems where solution variability is undesirable and needs to be quantified. I definitely foresee application of this in the world of text2SQL agents and as always would love to hear feedback from anyone willing to try out this approach in their work or refine it further.

