Agentic Engineering / September 2026
One Successful Agent Run Proves Almost Nothing
How I used parallel runs of the same engineering ticket to expose variance, compare agent behavior, and turn lucky successes into repeatable system behavior.

One of the easiest mistakes to make with an AI agent is to watch it succeed once and conclude that you have built something reliable.
I made the opposite assumption while building ARF.
If the pipeline solved a ticket once, I wanted to know why.
Then I wanted to know whether it could do it again.
And again.
Eventually I started running multiple ARF executions in parallel against the same engineering ticket and comparing them.
Same problem.
Same goal.
Different runs.
The differences were often more useful than the success itself.
The best run is not the result
Imagine several executions of the same ticket.
One reaches a clean reproduction, produces a solid fix, and survives independent verification.
Another eventually gets there but takes a long detour.
A third makes a plausible decision early and wastes most of its run pursuing it.
A fourth gets blocked by the environment.
It is tempting to look at the first run and say:
Great. Use that result.
But if I am developing the pipeline, that is not the interesting question.
The interesting question is:
Why was that run better?
Did the model reason better?
Did it already know how to use a tool?
Did the prompt give it a clearer goal?
Was the environment slightly different?
Did one expert have a skill that another path failed to use?
Did a deterministic helper prevent one run from wandering into ambiguity?
A good result may be evidence of a good system.
It may also be a lucky path through a probabilistic one.
My job was to tell the difference.
Let the pipeline compete against itself
I used Pipeline Fleet Architect, a separate engineering expert I worked with while developing ARF, to help with this.
I could have it launch multiple ARF runs against the same ticket, then analyze the resulting sessions and artifacts.
It had enough context to do more than compare final statuses.
It could inspect expert transcripts, virtual-filesystem deliverables, JSONL traces, pipeline state, Linear issues, repositories, and PRs.
So the comparison could ask:
Where did the runs first diverge?
What decision caused that divergence?
Was one path actually better, or did it only look better because it got lucky later?
What did the successful run know or do that the others did not?
Then came the important part:
Do not preserve the winning run. Preserve the reason it won.
Sometimes the answer is a skill
A useful example came from reproducing bugs affected by network conditions.
Suppose the hypothesis is that latency or unreliable connectivity triggers the failure.
The model now needs an environment where it can test that hypothesis.
One capable model may reason through the problem, install a proxy, configure it correctly, shape the network, and reproduce the behavior.
Excellent.
Another model may attack the same problem differently, consume a large number of tokens, and still fail to create the right conditions.
If I only compare outcomes, I might conclude that the first model is better.
Sometimes that is true.
But there is another possibility:
the system is asking the model to rediscover a procedure that should already be available.
Once we know how to create a controlled network-latency environment, I do not want every future agent reinventing that setup.
The reasoning problem remains:
Is network behavior a plausible mechanism worth testing?
The procedural problem becomes a skill:
Create the known test environment and apply the requested conditions.
Now the architecture has absorbed something that previously depended on one model being clever.
That is a repeatability improvement.
Sometimes the answer really is the model
Not every difference should be abstracted away.
Different model families have different strengths.
A model that is excellent at broad investigation may not be the one I want for implementation.
A model that writes convincing code may not be the one I want independently challenging that code.
One of the advantages of a specialist architecture is that the model choice can be local to the responsibility.
So if repeated runs show that one model consistently performs a particular role better, changing the model for that expert may be exactly the right answer.
The important part is that this becomes an architectural choice backed by observed behavior, not a global declaration that one model is "the best."
Tomorrow that choice may change.
The responsibility should survive it.
Sometimes the answer is the environment
Agent failures can look surprisingly intelligent.
An expert may produce several paragraphs of reasoning about why it cannot proceed.
The root cause may simply be that a tool is missing.
If parallel runs repeatedly hit the same environment limitation, I do not want to solve it with more persuasive prompting.
I want to fix the environment.
Install the tool.
Provide the dependency.
Make the capability part of the execution context.
Then rerun the experiment.
This is one of the broader lessons I took from building agentic systems:
A model can compensate for a weak environment, but that does not mean it should.
Every workaround you leave to model improvisation becomes another source of variance.
Sometimes the prompt is the problem
Prompts matter, but not in the magical way they are sometimes discussed.
A prompt is part of the system's specification.
If two reasonable interpretations of an expert's responsibility are possible, different runs may choose different ones.
That can be useful evidence.
Maybe the goal is too broad.
Maybe the stopping condition is unclear.
Maybe the expert does not know which evidence it must produce before handing work to the next stage.
Parallel runs expose those ambiguities quickly.
If one run interprets the task correctly and another takes an equally defensible but harmful path, I do not necessarily blame the second run.
I ask whether the contract was clear enough.
Then we tighten it and run again.
The experiment is the pipeline, not just the model
This is why I do not think of this kind of work as model evaluation alone.
I am evaluating a system.
That system includes models, prompts, skills, deterministic code, tools, environments, permissions, shared state, handoff contracts, and verification.
Changing any of them may change the result.
So when several runs behave differently, the useful question is not:
Which model hallucinated?
It is:
Which part of the system allowed these runs to diverge, and is that variance useful or harmful?
Some variance is exactly what I want.
Investigation requires judgment.
I do not want every expert mechanically following the same path through an unfamiliar problem.
But once a useful behavior becomes known and procedural, I want to move it out of chance.
That is the line I kept trying to find.
Pick the best run, then make it less special
There is a pattern I used repeatedly.
Run the same problem several times.
Compare the executions.
Pick the strongest result.
Then study it.
What did that run do differently?
If the advantage came from reusable knowledge, turn it into a skill.
If it came from a missing tool, improve the environment.
If it came from clearer reasoning in one specialist, consider a model change.
If the prompt created ambiguity, improve the contract.
If deterministic behavior was being improvised, move it into code.
Then run the ticket again.
And preferably another ticket after that.
The goal is not to force every run into an identical transcript.
The goal is to make the quality of the outcome less dependent on accidental choices.
Do not optimize for the best run. Reduce the variance between runs.
Beware of overfitting your pipeline
There is an obvious danger in repeatedly testing the same ticket.
You can make the pipeline extremely good at that ticket.
That proves very little.
So once a change seemed to improve the repeated case, I had to try it elsewhere.
Did the new skill help on another problem?
Did the tighter prompt accidentally constrain a legitimate investigation?
Did a model change improve one expert while making another class of work worse?
Did the environment improvement actually generalize?
This is very similar to ordinary software testing.
One case exposes a problem.
You fix it.
Then another case tells you whether you fixed the class of problem or merely the example.
Agentic systems need the same discipline.
A prompt can overfit.
A skill can encode a bad assumption.
A deterministic helper can make the wrong invariant permanent.
"More repeatable" is useful only if you are repeating the right behavior.
Independent verification changes the experiment
ARF also had a separate verifier using a different model family and a different goal from the agents producing the fix.
That matters when comparing runs.
A run that reaches a green test fastest is not automatically the best run.
Maybe the agent changed the test in a way that made it easier to pass.
Maybe the assertion never exercised the original failure.
Maybe the fix works against the synthetic reproduction but does not support the claim being made about the customer problem.
The verifier's job was to challenge those shortcuts.
That means a "winning" run had to do more than finish.
It had to leave evidence another specialist could attack.
This is one reason I care so much about architecture around the model.
If the same agent defines success, implements the change, writes the test, and judges the result, the system has very little resistance to self-confirmation.
Parallelism shows you variance.
Independent verification tells you whether the variance actually matters.
Observability makes this possible
None of this works well if the only artifact from an agent run is a final message.
To compare runs, I needed evidence.
What did the expert try?
What failed?
What tool did it use?
What did it produce?
Where did it hand off?
What state did the pipeline record?
What repository changes resulted?
ARF experts left structured artifacts and JSONL breadcrumbs, and the broader system kept enough state to reconstruct what happened.
Initially, that kind of observability is useful because debugging agents without history is painful.
Later, it becomes the basis for something more interesting:
experimentation.
You can compare runs because the runs leave something comparable behind.
And once you have that, the system can start teaching you where its own variability comes from.
Repeatability does not mean determinism
I do not expect an agentic engineering pipeline to execute identically every time.
That would defeat part of the reason to use models.
I want agents to interpret evidence, form hypotheses, adapt to unfamiliar code, and choose useful next actions.
Those are probabilistic tasks.
The architecture should not remove that.
It should constrain the places where creativity is not useful.
Known setup procedures.
State transitions.
Identity checks.
Evidence contracts.
Reproducible test environments.
Acceptance conditions.
Those are places where repeatability is more valuable than improvisation.
The art is deciding which side of that boundary a behavior belongs on.
A successful run is the beginning of the investigation
When one run succeeds and another fails, I do not see the failed run only as wasted compute.
The difference is information.
It tells me something about the system.
Sometimes the answer is a better model.
Sometimes it is a skill.
Sometimes an environment change.
Sometimes deterministic code.
Sometimes a clearer contract.
And sometimes the right answer is to leave the variability alone because the problem genuinely requires judgment.
That is why a single successful run proves almost nothing.
The work begins when you ask why it succeeded.
I wasn't trying to find the run that got lucky. I was trying to understand what made it better, then make that advantage part of the system.

