Agentic Engineering / August 2026
Agentic AI Primitives (Explained Through ARF)
A guided tour of the experts, models, tools, shared artifacts, and human decisions inside the ARF engineering pipeline.

Most confusion about agentic engineering comes from arguing about models instead of defining the pieces. So instead of a glossary, here's the whole thing explained through one real system: ARF.
What ARF is
ARF — Autonomous Repro & Fix — takes the messy start of a software defect and turns it into something an engineer can actually review. It ends with a merge-ready pull request sitting in front of a human.
Between those two points there's a pipeline of experts, a shared filesystem, a set of triggers, and three places where a human has to decide.
The workflow is bounded to the client side — the Augment VS Code extension and the JetBrains plugin. A problem that needs a backend change gets routed to the owning team instead of being "fixed" against a stale copy.
Triggers
The pipeline doesn't run because somebody remembered to run it. It runs on a trigger.
The first trigger is a support engineer reading a customer ticket and deciding it looks like a good candidate for ARF. That's a human judgment. Most tickets don't qualify. The ones that do get picked up.
But a ticket review isn't the only way in. Any expert in the system can be triggered on demand by a user, or by anything else inside or outside the system that decides a run should happen. A webhook. An API call. A scheduled job. Another expert.
Trigger decides when. The pipeline decides what.
The primitives
Expert / Agent. The core unit. Not a personality wrapped around an LLM — a package with a purpose. A model, a goal, prompts and guidelines, skills, tools and MCP integrations, a session, an execution environment, permissions, access to the VFS, and a defined output.
Model. The reasoning engine inside an expert. Swappable. Not the architecture. Different experts can run different model families, and the fixer and verifier deliberately do.
Skill. A reusable capability, backed by deterministic code where possible. Something the system already knows how to do, so a model doesn't rediscover it every run. If a procedure is known, it shouldn't be a model's problem.
Prompt / Guideline. Instructions that shape an expert's behavior. If changing it changes production behavior, it's software — and it belongs in version control.
MCP / Tool. How an expert reaches outside itself. A filesystem, a database, a browser, an internal API.
Environment. Where an expert runs. A dedicated cloud VM, configured with the tools, permissions, and state it's allowed to touch. ARF runs separate lanes for VS Code on Linux, VS Code on Windows 11, and JetBrains on Linux. macOS reports are admitted but have no native worker.
VFS. A shared virtual filesystem every expert can read and write. In Augment Code this is where experts exchange information — in my case, structured JSONL files. It's how one expert hands evidence to the next without either of them knowing anything about the other's internals.
Session. One execution of an expert, with its own context and history.
Artifact. Structured output an expert leaves behind — evidence, decisions, breadcrumbs. Usually written to the VFS so the next expert or a human can read it. ARF uses typed artifacts like BugBrief, ReproResult, FixResult, VerifyResult, and CoordinatorResult.
Contract. The agreed shape of what one expert hands to the next. Goals in, artifacts out.
Run / Action. A single pass of the pipeline over a real input.
Verifier. An expert with one job: challenge the claim that the work is correct. Different model, different prompt, different objective from whoever built the thing. The builder should not be the only judge of what it built.
Human checkpoint. A decision point that stays human on purpose.
The experts
The experts cover the main intake-to-verification workflow and the operating environment around it.
Intake and investigation
Slack Intake — investigates a root extension-feedback message and finds or creates the right DevRev ticket. Slack stays read-only.
Intake — admits candidate tickets, separates distinct defects, clusters, and attaches to or creates Linear issues.
Diagnosis — builds and compares candidate root causes from support evidence, engineering timelines, source context, prior issues, and error signals.
Diagnosis Decision Helper — explains the evidence, supports further investigation, and records the operator's approved objective and candidate. It cannot rewrite the candidate set.
Execution
BugBrief — turns issue evidence and decisions into a structured execution brief.
Coordinator — runs the resumable state machine, routes workers, validates results, manages run state, and owns PR promotion. It does not edit source, build, test, approve, or merge.
Repro — reproduces VS Code defects on Linux with a real extension and a Playwright/Electron harness.
Repro Win — reproduces VS Code defects on a real Windows 11 desktop.
Repro JB — reproduces IntelliJ/JetBrains defects on Linux, picking the lowest adequate testing tier.
Fix — fixes the VS Code/Linux defect and handles assigned CI and review work. Reuses Repro's branch and PR.
Fix Win — applies and tests fixes on the native Windows lane.
Fix JB — applies and tests JetBrains fixes.
Verifier — independently rechecks the fix and the failing baseline on a clean checkout.
Operational visibility
Dashboard — collects fleet evidence, builds state, KPIs, and an attention queue, and publishes the portal snapshot. Read-only against the source systems.
IDE Status Reporter — produces a weekly IDE-product status report by joining support, Linear, and ARF data.
Shared improvement
Crumb Curator — turns repeated environment lessons into curated hints or harness-document PRs.
VS Code Test Author — expands the VS Code test suite from a prioritized backlog.
JetBrains Test Author — expands JetBrains integration-test coverage.
The three human gates
ARF is autonomous between the gates, not across them.
Gate 1 — Support marks the ticket.
A support engineer reads a customer ticket and decides it's a good candidate for ARF. That starts the pipeline. Nothing before this point is automatic, because deciding what deserves engineering attention is a judgment call, not a routing rule.
Gate 2 — Triage and root cause.
The pipeline produces an internal dev ticket. It does the triage and deduplication work, and it arrives with at most three hypotheses about what's actually wrong.
Then a human sits down with the Decision Helper expert, discusses the hypotheses, and decides which one is correct — or supplies a new one the pipeline missed.
This gate is the reason the pipeline doesn't just barrel ahead. The experts can narrow the space. They can't authoritatively pick the root cause. That's still a decision with consequences, and it belongs to a person who understands the codebase.
Once the hypothesis is chosen, the pipeline continues on its own.
Gate 3 — Merge.
Eventually there's a merge-ready PR. Other experts have already reviewed it. CI has run. Verification has done its job — the regression test passes with the fix, fails without it, and the wider checks still succeed.
But the merge itself is a human decision.
The reviewer can sit with a separate expert and "talk" to the PR — ask questions, walk through what changed and why, check the reasoning against the evidence. Then decide whether to merge.
The expert doesn't merge. It helps the human understand what they're about to merge.
Why the gates are where they are
Three places, all of them judgment calls with real cost if wrong:
Is this ticket worth engineering time?
Is this the actual root cause?
Should this change land?
Everything between those points is fair game for automation. Everything at those points stays human.
The rest of the pipeline — reproduction, fixing, verification, automated review, CI feedback — is where the experts earn their keep. They run in dedicated cloud VMs, exchange artifacts through the VFS, follow contracts, and leave evidence behind. Where the next useful action depends on judgment, a model decides. Where it's a known procedure, deterministic code does it.
The coding model is one component. The system is the point.


