A personal notebook on technology, work, and curiosity.
← Back to the notebook

Agentic Engineering / October 2026

I Don’t Use One AI: I Build an AI Engineering Team

A practical workflow for separating architecture, implementation, review, verification and learning while keeping human ownership.

When I use AI for engineering, I find it useful to think in responsibilities. Exploring a design, implementing it, challenging it, checking its behavior and teaching me how it works are different jobs.

Calling this a team is a working analogy. The roles do not imply human colleagues, independent judgment by default, or a requirement to run several agents at once. One bounded task may need only an implementation session and a good check. A difficult change may benefit from more separation.

A side project is enough to make these responsibilities useful. I can use AI to explore options and implement a chosen approach, then give review a separate objective: challenge whether the result meets the requirement. The workflow below is a working pattern, not a claim that every task needs every role.

One human owner directs and accepts the work; architecture, build, review, verification and learning are distinct responsibilities.
One owner, several responsibilities. These roles need not mean separate models or simultaneous agents.

Begin with the decision and the evidence

Before assigning roles, I want a clear outcome. What problem are we solving? Who needs it? Which constraints are real? What would show that the work is ready?

A convincing demonstration may still miss an important requirement. I want the brief to identify what matters and what evidence would establish it before implementation makes one approach feel inevitable.

A useful brief gives the agents a problem, an acceptance bar and boundaries. It also records open questions. If the architecture is undecided, investigation should produce alternatives for a decision. Implementation authority follows the scope I actually assigned.

Give each role a job it can finish

A human owner sets the problem and acceptance bar. Architecture, implementation, critical review, verification and teaching return evidence and questions.
Separate the jobs and their handoffs; bring evidence and unanswered questions back to the owner.Scroll the diagram sideways on a narrow screen.
Role Responsibility Handoff
Architecture and reasoning Explore options, constraints, trade-offs and failure modes Proposed approach, alternatives, assumptions and unresolved decisions
Implementation Build the agreed change within its scope Inspectable change, relevant checks and known limitations
Critical review Challenge the design or implementation against requirements and evidence Specific findings, counterexamples and remaining uncertainty
Verification Execute appropriate checks and examine what they establish Results tied to the tested artifact, conditions and claim
Teaching and comprehension Help the human trace behavior and examine weak understanding A walkthrough, examples and questions the human can answer
Human owner Set direction, resolve trade-offs, assign authority and accept the result A reasoned decision with an owner and acknowledged limits

On a narrow screen, scroll the table sideways to read all columns.

These are responsibilities, not a shopping list of models. Verification may mostly be ordinary code and tooling. Teaching may happen in the architecture session. A review may need a fresh context because the builder's reasoning has become too dominant.

Model choice should follow the role and observed performance on relevant work. I do not need a permanent declaration that one model is best at everything.

Separate implementation from acceptance

The implementation agent needs enough context and authority to complete routine work. Repeatedly interrupting it for decisions already settled does not create useful control.

It also needs a clear stopping condition. A material change in scope, an unresolved architectural choice or a consequential action outside the brief should return to the owner. That boundary is easier to respect when it is explicit before the work starts.

At handoff, I want the actual change and its evidence. A summary helps me navigate; it does not replace the artifact. The reviewer needs the requirements, relevant source material, the changed version and any important limitations in the checks.

If work runs in parallel, responsibilities and shared interfaces need to be clear. Several writers changing the same component can create integration work that cancels the benefit. I bound parallelism by what I can supervise, including the review queue.

Another model is not automatic independence

A second model can challenge an assumption that the first accepted. It can also repeat the same mistake. If both receive the same leading question, incomplete evidence and preferred conclusion, different branding has not solved the problem.

For a consequential review, I want a distinct objective: find where the claim fails, identify missing evidence and test whether the proposed result meets the requirement. The reviewer should inspect underlying material, including evidence the builder did not emphasize.

I may ask for an initial assessment from the requirements and artifact before showing the builder's rationale. That is a proposed way to reduce anchoring, not a guarantee of independence. Context cannot simply be removed without cost; the reviewer still needs enough information to judge the work fairly.

I would evaluate review quality through the defects it catches and misses, its false alarms and its cost. Agreement between two models is not a reliability percentage. Disagreement is a prompt to investigate the issue, not a vote that the human has to break by intuition alone.

Choose a gate for the claim

For a reproducible bug, a useful check may demonstrate failure against the unfixed version and success against the fix, while preserving the relevant test conditions. Broader regression checks answer additional questions.

That evidence has a boundary. A synthetic failure can show that a fix handles an injected condition. It does not establish that the same condition caused a particular customer's incident. I need customer evidence for that causal claim.

For an AI feature, I would define evaluation cases around its intended use and inspect the failures. Some assertions can be checked deterministically. Model behavior may require repeated trials, representative examples and review of cases the automated score misses. The evaluation is evidence about those conditions, not proof of correctness everywhere.

For an architectural decision, tests alone cannot choose the right customer trade-off. A technically sound design may still be too expensive, difficult to operate or mismatched to the actual workflow. The human gate needs the requirement, options, evidence and consequences in a form that supports a decision.

Make comprehension part of the workflow

In a side project, the implementation can move ahead of my ability to explain an important choice. I want to notice that gap and pause long enough to understand the reasoning, rather than rely on a polished summary.

A useful learning loop is to ask a concrete question, explain the answer back and check its implications. That helps expose what I still need to understand; it does not by itself establish that the implementation is correct.

For work I am about to accept, I want a stronger version of the same loop. I can ask the teaching role to trace a request through the artifact, identify a trade-off and propose a failure scenario. Then I should explain or predict the behavior myself and check that explanation against the implementation.

This catch-up belongs before the decision that depends on it. A gap about an internal detail may be manageable through a documented contract. A gap about who can access customer data is relevant immediately. The required depth depends on what I am taking responsibility for.

Keep a small, useful handoff

For a significant change, I want a record another engineer can use:

This can be a short document for a small change. It should point to evidence rather than repeat every agent message. It should also make clear when a reviewer is looking at a newer version than the one the checks covered.

A workflow with a way back

An engineering workflow with a way back

Human owner: direction, authority, trade-offs and acceptance.

  1. Brief

    Define the requirement and evidence needed for acceptance.

  2. Architecture and reasoning

    Explore alternatives, assumptions and failure modes.

  3. Human design decisionHuman decision

    Choose the approach and authorize its scope.

  4. Implementation

    Produce an inspectable artifact and supporting checks.

  5. Critical review + Verification

    Challenge the claims and check the behavior. Combine the evidence; these roles need not run in parallel.

    Insufficient evidence → return to Implementation. Changed requirement → return to Brief.

  6. Comprehension check

    Trace, predict and explain the important decisions.

    Gap remains → learn and check again. Design problem → return to Architecture and reasoning.

  7. Human acceptanceHuman decision

    Accept → record the result, evidence and ownership.

    Unresolved trade-off → return to the relevant decision.

Roles can share a session. Different models do not guarantee independent errors.

Separate responsibilities, inspectable handoffs and evidence matched to the decision. A proposed operating pattern, not an audited execution trace.

If evidence is insufficient, return to the relevant work. If the requirement is wrong, revisit the brief. If I cannot explain an important decision, use the artifact and evidence to learn it before acceptance.

The process should get smaller when it can. A new role that merely restates the previous answer adds little. A check that cannot affect a decision may be unnecessary. I want enough structure to expose consequential errors and preserve ownership, with room for the model to reason within its task.

I am comfortable saying that AI did substantial implementation work. My claim is about the problem I framed, the architecture and decisions I owned, the process I directed and the result I can defend. That account is useful only when I can show the evidence and acknowledge its limits.

The six-principle framework puts this workflow beside the habits that keep it manageable. The open questions include the parts I have not settled, especially how to measure review independence and sufficient understanding.

Oct 2026PublishedAI Slop vs AI CraftHow you use AI on a task makes a big difference: in how you reach the solution, and in whether you can own it afterwards.Agentic EngineeringSep 2026PublishedARF series · 4 / 4One Successful Agent Run Proves Almost NothingImproving the pipeline over time — using manual methods and an automatic breadcrumb system.Agentic EngineeringSep 2026PublishedARF series · 3 / 4I Was Still the Architect. AI Made Me Faster.I created an expert to help me work on the ARF pipeline.Agentic Engineering