A personal notebook on technology, work, and curiosity.
← Back to the notebook

Agentic Engineering / August 2026

The Model Is Not the System

What building ARF taught me about turning capable coding models into a repeatable engineering system with specialist agents, deterministic skills, independent verification, environments, and versioned architecture.

ARF workflow showing specialist agents, independent verification, human review and supporting infrastructure.

AI can write code.

That turned out to be the easy part.

While building ARF — Autonomous Repro & Fix — at Augment Code, I spent an unreasonable amount of time feeding the same VS Code bug through an agentic engineering pipeline again and again.

The bug looked simple from the outside: some conversations appeared to disappear after the extension restarted.

The actual failure was buried deeper in the extension architecture. Part of the relevant state lived in the webview and Redux and had to be persisted through VS Code workspace storage. Under a failure condition, that write never happened. The running extension did not make the failure visible enough, so everything could look normal for a while. After reload, the newer conversations were simply not there because they had never been persisted.

A capable coding model could inspect that code and propose a fix.

My problem was different:

Could I build a system that would reach a defensible result repeatedly, without depending on one particular model having a particularly good day?

That question changed the way I thought about agentic software engineering.

Solving the ticket once was not the goal

ARF was not one coding agent with a very long prompt.

Its job was to connect the messy beginning of a software defect to something an engineer could actually review: intake, deduplication, issue tracking, diagnosis, reproduction, implementation, verification, automated review and CI feedback, human review, and ultimately a mergeable pull request.

There were human decision points in that flow. There were deterministic components. There were agents with different responsibilities. And there were places where the right answer was to stop rather than invent evidence that did not exist.

I am deliberately simplifying the internal architecture here. The important point is that the coding model was only one component of a larger engineering system.

The missing-conversations issue became one of the tickets I used to develop that system.

At first, I watched almost everything.

For the first five or six runs I followed the experts closely. I read what they were doing, watched where they hesitated, looked at which tools they chose, checked what they concluded and compared that with the actual engineering evidence.

Sometimes the model made a bad decision.

More interestingly, sometimes the model made a perfectly reasonable decision that exposed a bad decision in my architecture.

Maybe the goal was underspecified. Maybe a handoff had lost important context. Maybe I had given a model responsibility for something that should have been deterministic. Maybe the test technically passed without proving what I intended it to prove.

I would change the system and run the ticket again.

And again.

Eventually I stopped watching every move. I could let the pipeline work, come back later, inspect the artifacts and decide whether the result was trustworthy.

That was a much more meaningful milestone than watching an agent successfully generate a patch.

A model solving a ticket once is interesting. A system solving it predictably is engineering.

An agent is more than a model and a prompt

One of the easiest mistakes in agentic engineering is to think of an agent as a personality wrapped around an LLM.

“You are an expert software engineer. Investigate this bug and fix it.”

That can produce impressive results. It can also produce wildly different behavior between runs.

I found it more useful to think in terms of responsibilities.

A reproduction expert has a different goal from a fixing expert. A verifier has a different goal from both of them. An intake expert should not behave like a developer, and an agent diagnosing competing root-cause hypotheses should not silently turn its favorite hypothesis into a fact.

In practice, an expert becomes a package of things: a model, a goal, prompts and guidelines, skills, tools and MCP integrations, a session, an execution environment, permissions, access to shared artifacts and a defined output.

We built ARF on Augment's Cosmos platform, where those experts could operate in dedicated environments and collaborate across a larger workflow.

That architecture also meant I did not need to make one global decision such as “this system uses model X.”

Different responsibilities can use different models.

If one model family is better at a particular coding task and another is better at challenging the result, that choice can stay local to the expert. If the best model changes six months later, I want to replace a component, not redesign the whole system.

The responsibility should survive the model.

Sometimes the best AI component is a Python script

A bug-reproduction agent often has to do more than read source code.

Suppose the evidence suggests that a problem may depend on network latency. The agent now has to create the right conditions: install or use a proxy, configure the environment, introduce controlled network behavior and then reproduce the issue through that setup.

One model did this well. It reasoned through the problem, created the environment and got to useful evidence.

Another model facing the same kind of task chose a different path. It consumed tokens, did work that looked plausible, and still did not create the conditions correctly.

My conclusion was not that the second model was bad.

My conclusion was that I had asked both models to repeatedly rediscover something the system could already know.

Once I know how to create a controlled network-latency environment, I do not want an expensive reasoning model reinventing that procedure every time.

I want the model to decide:

Network behavior is a plausible hypothesis. I should test it.

But the known mechanics of setting up that test can become a reusable skill, backed by deterministic code.

That leads to one of the rules I kept coming back to:

Don't ask a model to reason about something you've already reduced to a procedure.

The inverse matters too.

It is possible to respond to model unpredictability by turning everything into workflow machinery. That is not automatically better. If the next useful action depends on evidence you could not know when designing the system, that is exactly the kind of judgment a model is good at.

The architectural work is deciding where that boundary belongs.

Models for judgment.

Code for invariants.

Skills for reusable capability.

The verifier should not want the fix to pass

There is another uncomfortable property of coding agents: they are very good at satisfying the thing you asked them to satisfy.

That is not always the same as solving the problem you meant.

Imagine an agent that fixes some code and another agent that writes or updates a test. Both are working toward the same visible outcome: green.

They do not intentionally cheat. But a capable model can find surprisingly creative ways to satisfy a measurable definition of success while missing the engineering intent behind it.

A test can pass because the important assertion never ran.

A regression test can be changed until it no longer represents the original failure.

A synthetic fault can prove that the software survives the injected fault without proving that the customer ever experienced that fault.

So I wanted verification to be genuinely independent.

In ARF, the verifier used a different model family from the agents doing the implementation work. It also had its own prompt, skills, constraints and objective.

Its job was not to help the fix succeed.

Its job was to challenge the claim that the fix was real.

For a typical verified result, I wanted something much stronger than “the test is green.” The regression test should pass with the fix present. The same test should fail against the unfixed baseline. And the wider configured regression checks should still succeed on the proposed revision.

That is a very different thing from asking a second model, “Does this PR look good to you?”

The separation of responsibilities matters because the builder should not be the only judge of what it built.

You have to test the pipeline, not just the code

Once the system had specialists, skills, environments and verification, development still looked surprisingly familiar.

Pick a real ticket.

Run it.

Inspect the result.

Find the weak point.

Change one thing.

Run it again.

Did the agent misunderstand the goal?

Did the environment differ from what I thought it was?

Did an expert choose a tool when a deterministic skill already existed?

Did the next expert lose some evidence in the handoff?

Did verification prove the original claim or merely produce a green signal?

Eventually the same ticket becomes boring.

Boring is good.

Then you take another ticket and discover which of your “general improvements” were actually assumptions you had accidentally optimized for in the first case.

So you repeat the process.

That is why I do not find the term prompt engineering sufficient for this kind of work. Prompts matter, but the prompt is only one variable in a system that also includes environments, tools, permissions, state, contracts, deterministic code, evidence and human decisions.

You are testing the architecture.

If a prompt can change production behavior, it belongs in version control

Once a system behaves differently because you changed three sentences in a prompt, those three sentences are part of the software.

The same applies to a skill, an expert definition, a helper script, a schema, an environment definition or a handoff contract.

I want those things in Git.

Not because Git somehow makes an AI system deterministic. It does not.

It gives me something just as important: history.

If Tuesday's version behaves worse than Monday's version, I want a diff.

What changed?

Can I reproduce the old behavior?

Can I roll back?

Can I branch an experiment instead of destroying a working configuration?

Can somebody else review the change?

Can I move the system to another environment without reconstructing its behavior from memory and UI settings?

This also matters for model independence.

Models will change. Platforms will change. The system should not exist only as a collection of conversations inside one product.

The durable parts should be explicit and portable: goals, prompts, skills, contracts, scripts, tests, environment definitions and the knowledge the experts rely on.

If those artifacts are versioned, changing the model underneath one expert becomes an engineering change rather than an identity crisis for the whole system.

Good autonomy means I can stop watching

There is a version of “autonomous agents” that sounds like this:

I started the agent and I have no idea what it is doing now.

That is not the kind of autonomy I want.

Early in ARF development I watched the runs because I needed to understand the failure modes. Later I could walk away because the system left enough evidence behind for me to reconstruct what had happened.

The experts were designed to leave structured breadcrumbs in JSONL artifacts: what they tried, what succeeded, what failed, and what they learned about the environment along the way.

Those breadcrumbs were useful for debugging, but they suggested something more interesting.

If the system encounters the same environment trap repeatedly, why should every future expert discover it from scratch?

That question eventually led into shared learning, short- and long-term memory, and a separate architecture for analyzing the pipeline itself.

That deserves its own article.

For this one, the important point is simpler:

Good autonomy is not the absence of visibility. It is being able to stop watching because the work remains inspectable.

The architecture has to outlive today's best model

Coding models will continue to improve.

That does not remove the architectural work. In some ways it makes that work more important.

When generating another hundred lines of code becomes cheap, the difficult questions move outward.

What problem are we actually solving?

Which component owns it?

What evidence is required before another component can proceed?

Where should a model reason, and where should code enforce an invariant?

What happens when a session disappears halfway through the work?

How do we know the test proves what we think it proves?

Who is allowed to change what?

Which decisions still require a human?

How do we inspect a failure after the agents have stopped running?

How do we replace a model without throwing away the engineering system around it?

Those are architecture questions.

ARF started from a practical need: turn real IDE problems into reproducible engineering work and reviewable fixes. Building it taught me that the impressive part of an AI system is rarely the moment a model writes code.

The interesting work is everything required to make that moment trustworthy and repeatable.

The model is not the system.

Sep 2026PublishedOne Successful Agent Run Proves Almost NothingHow I used parallel runs of the same engineering ticket to expose variance, compare agent behavior, and turn lucky successes into repeatable system behavior.Agentic EngineeringSep 2026PublishedI Was Still the Architect. AI Made Me Faster.How I used a dedicated AI expert to help design, implement, and test changes across a multi-repository agentic engineering system without giving up architectural authority.Agentic Engineering