A personal notebook on technology, work, and curiosity.
← Back to the notebook

Agentic Engineering / August 2026

Agentic AI Primitives (Explained Through ARF)

A guided tour of the experts, models, tools, shared artifacts, and human decisions inside the ARF engineering pipeline.

Specialist agents moving through serial and parallel work around shared evidence and human decision gates.

Most confusion about agentic engineering comes from arguing about models instead of defining the pieces. So instead of a glossary, here's the whole thing explained through one real system: ARF.

What ARF is

ARF — Autonomous Repro & Fix — takes the messy start of a software defect and turns it into something an engineer can actually review. It ends with a merge-ready pull request sitting in front of a human.

Between those two points there's a pipeline of experts, a shared filesystem, a set of triggers, and three places where a human has to decide.

The workflow is bounded to the client side — the Augment VS Code extension and the JetBrains plugin. A problem that needs a backend change gets routed to the owning team instead of being "fixed" against a stale copy.

Triggers

The pipeline doesn't run because somebody remembered to run it. It runs on a trigger.

The first trigger is a support engineer reading a customer ticket and deciding it looks like a good candidate for ARF. That's a human judgment. Most tickets don't qualify. The ones that do get picked up.

But a ticket review isn't the only way in. Any expert in the system can be triggered on demand by a user, or by anything else inside or outside the system that decides a run should happen. A webhook. An API call. A scheduled job. Another expert.

Trigger decides when. The pipeline decides what.

The primitives

Expert / Agent. The core unit. Not a personality wrapped around an LLM — a package with a purpose. A model, a goal, prompts and guidelines, skills, tools and MCP integrations, a session, an execution environment, permissions, access to the VFS, and a defined output.

Model. The reasoning engine inside an expert. Swappable. Not the architecture. Different experts can run different model families, and the fixer and verifier deliberately do.

Skill. A reusable capability, backed by deterministic code where possible. Something the system already knows how to do, so a model doesn't rediscover it every run. If a procedure is known, it shouldn't be a model's problem.

Prompt / Guideline. Instructions that shape an expert's behavior. If changing it changes production behavior, it's software — and it belongs in version control.

MCP / Tool. How an expert reaches outside itself. A filesystem, a database, a browser, an internal API.

Environment. Where an expert runs. A dedicated cloud VM, configured with the tools, permissions, and state it's allowed to touch. ARF runs separate lanes for VS Code on Linux, VS Code on Windows 11, and JetBrains on Linux. macOS reports are admitted but have no native worker.

VFS. A shared virtual filesystem every expert can read and write. In Augment Code this is where experts exchange information — in my case, structured JSONL files. It's how one expert hands evidence to the next without either of them knowing anything about the other's internals.

Session. One execution of an expert, with its own context and history.

Artifact. Structured output an expert leaves behind — evidence, decisions, breadcrumbs. Usually written to the VFS so the next expert or a human can read it. ARF uses typed artifacts like BugBrief, ReproResult, FixResult, VerifyResult, and CoordinatorResult.

Contract. The agreed shape of what one expert hands to the next. Goals in, artifacts out.

Run / Action. A single pass of the pipeline over a real input.

Verifier. An expert with one job: challenge the claim that the work is correct. Different model, different prompt, different objective from whoever built the thing. The builder should not be the only judge of what it built.

Human checkpoint. A decision point that stays human on purpose.

The experts

The experts cover the main intake-to-verification workflow and the operating environment around it.

Intake and investigation

Execution

Operational visibility

Shared improvement

The three human gates

ARF is autonomous between the gates, not across them.

Gate 1 — Support marks the ticket.

A support engineer reads a customer ticket and decides it's a good candidate for ARF. That starts the pipeline. Nothing before this point is automatic, because deciding what deserves engineering attention is a judgment call, not a routing rule.

Gate 2 — Triage and root cause.

The pipeline produces an internal dev ticket. It does the triage and deduplication work, and it arrives with at most three hypotheses about what's actually wrong.

Then a human sits down with the Decision Helper expert, discusses the hypotheses, and decides which one is correct — or supplies a new one the pipeline missed.

This gate is the reason the pipeline doesn't just barrel ahead. The experts can narrow the space. They can't authoritatively pick the root cause. That's still a decision with consequences, and it belongs to a person who understands the codebase.

Once the hypothesis is chosen, the pipeline continues on its own.

Gate 3 — Merge.

Eventually there's a merge-ready PR. Other experts have already reviewed it. CI has run. Verification has done its job — the regression test passes with the fix, fails without it, and the wider checks still succeed.

But the merge itself is a human decision.

The reviewer can sit with a separate expert and "talk" to the PR — ask questions, walk through what changed and why, check the reasoning against the evidence. Then decide whether to merge.

The expert doesn't merge. It helps the human understand what they're about to merge.

Why the gates are where they are

Three places, all of them judgment calls with real cost if wrong:

Everything between those points is fair game for automation. Everything at those points stays human.

The rest of the pipeline — reproduction, fixing, verification, automated review, CI feedback — is where the experts earn their keep. They run in dedicated cloud VMs, exchange artifacts through the VFS, follow contracts, and leave evidence behind. Where the next useful action depends on judgment, a model decides. Where it's a known procedure, deterministic code does it.

The coding model is one component. The system is the point.

Sep 2026PublishedARF series · 4 / 4One Successful Agent Run Proves Almost NothingHow I used parallel runs of the same engineering ticket to expose variance, compare agent behavior, and turn lucky successes into repeatable system behavior.Agentic EngineeringSep 2026PublishedARF series · 3 / 4I Was Still the Architect. AI Made Me Faster.How I used a dedicated AI expert to help design, implement, and test changes across a multi-repository agentic engineering system without giving up architectural authority.Agentic EngineeringAug 2026PublishedARF series · 2 / 4The Model Is Not the SystemWhat building ARF taught me about turning capable coding models into a repeatable engineering system with specialist agents, deterministic skills, independent verification, environments, and versioned architecture.Agentic Engineering