Agentic Engineering / October 2026
I Don’t Use One AI: I Build an AI Engineering Team
A practical workflow for separating architecture, implementation, review, verification and learning while keeping human ownership.
When I use AI for engineering, I find it useful to think in responsibilities. Exploring a design, implementing it, challenging it, checking its behavior and teaching me how it works are different jobs.
Calling this a team is a working analogy. The roles do not imply human colleagues, independent judgment by default, or a requirement to run several agents at once. One bounded task may need only an implementation session and a good check. A difficult change may benefit from more separation.
A side project is enough to make these responsibilities useful. I can use AI to explore options and implement a chosen approach, then give review a separate objective: challenge whether the result meets the requirement. The workflow below is a working pattern, not a claim that every task needs every role.
Begin with the decision and the evidence
Before assigning roles, I want a clear outcome. What problem are we solving? Who needs it? Which constraints are real? What would show that the work is ready?
A convincing demonstration may still miss an important requirement. I want the brief to identify what matters and what evidence would establish it before implementation makes one approach feel inevitable.
A useful brief gives the agents a problem, an acceptance bar and boundaries. It also records open questions. If the architecture is undecided, investigation should produce alternatives for a decision. Implementation authority follows the scope I actually assigned.
Give each role a job it can finish
| Role | Responsibility | Handoff |
|---|---|---|
| Architecture and reasoning | Explore options, constraints, trade-offs and failure modes | Proposed approach, alternatives, assumptions and unresolved decisions |
| Implementation | Build the agreed change within its scope | Inspectable change, relevant checks and known limitations |
| Critical review | Challenge the design or implementation against requirements and evidence | Specific findings, counterexamples and remaining uncertainty |
| Verification | Execute appropriate checks and examine what they establish | Results tied to the tested artifact, conditions and claim |
| Teaching and comprehension | Help the human trace behavior and examine weak understanding | A walkthrough, examples and questions the human can answer |
| Human owner | Set direction, resolve trade-offs, assign authority and accept the result | A reasoned decision with an owner and acknowledged limits |
On a narrow screen, scroll the table sideways to read all columns.
These are responsibilities, not a shopping list of models. Verification may mostly be ordinary code and tooling. Teaching may happen in the architecture session. A review may need a fresh context because the builder's reasoning has become too dominant.
Model choice should follow the role and observed performance on relevant work. I do not need a permanent declaration that one model is best at everything.
Separate implementation from acceptance
The implementation agent needs enough context and authority to complete routine work. Repeatedly interrupting it for decisions already settled does not create useful control.
It also needs a clear stopping condition. A material change in scope, an unresolved architectural choice or a consequential action outside the brief should return to the owner. That boundary is easier to respect when it is explicit before the work starts.
At handoff, I want the actual change and its evidence. A summary helps me navigate; it does not replace the artifact. The reviewer needs the requirements, relevant source material, the changed version and any important limitations in the checks.
If work runs in parallel, responsibilities and shared interfaces need to be clear. Several writers changing the same component can create integration work that cancels the benefit. I bound parallelism by what I can supervise, including the review queue.
Another model is not automatic independence
A second model can challenge an assumption that the first accepted. It can also repeat the same mistake. If both receive the same leading question, incomplete evidence and preferred conclusion, different branding has not solved the problem.
For a consequential review, I want a distinct objective: find where the claim fails, identify missing evidence and test whether the proposed result meets the requirement. The reviewer should inspect underlying material, including evidence the builder did not emphasize.
I may ask for an initial assessment from the requirements and artifact before showing the builder's rationale. That is a proposed way to reduce anchoring, not a guarantee of independence. Context cannot simply be removed without cost; the reviewer still needs enough information to judge the work fairly.
I would evaluate review quality through the defects it catches and misses, its false alarms and its cost. Agreement between two models is not a reliability percentage. Disagreement is a prompt to investigate the issue, not a vote that the human has to break by intuition alone.
Choose a gate for the claim
For a reproducible bug, a useful check may demonstrate failure against the unfixed version and success against the fix, while preserving the relevant test conditions. Broader regression checks answer additional questions.
That evidence has a boundary. A synthetic failure can show that a fix handles an injected condition. It does not establish that the same condition caused a particular customer's incident. I need customer evidence for that causal claim.
For an AI feature, I would define evaluation cases around its intended use and inspect the failures. Some assertions can be checked deterministically. Model behavior may require repeated trials, representative examples and review of cases the automated score misses. The evaluation is evidence about those conditions, not proof of correctness everywhere.
For an architectural decision, tests alone cannot choose the right customer trade-off. A technically sound design may still be too expensive, difficult to operate or mismatched to the actual workflow. The human gate needs the requirement, options, evidence and consequences in a form that supports a decision.
Make comprehension part of the workflow
In a side project, the implementation can move ahead of my ability to explain an important choice. I want to notice that gap and pause long enough to understand the reasoning, rather than rely on a polished summary.
A useful learning loop is to ask a concrete question, explain the answer back and check its implications. That helps expose what I still need to understand; it does not by itself establish that the implementation is correct.
For work I am about to accept, I want a stronger version of the same loop. I can ask the teaching role to trace a request through the artifact, identify a trade-off and propose a failure scenario. Then I should explain or predict the behavior myself and check that explanation against the implementation.
This catch-up belongs before the decision that depends on it. A gap about an internal detail may be manageable through a documented contract. A gap about who can access customer data is relevant immediately. The required depth depends on what I am taking responsibility for.
Keep a small, useful handoff
For a significant change, I want a record another engineer can use:
- The objective, scope and accepted requirements.
- The artifact or version being reviewed.
- Important design decisions, assumptions and alternatives.
- Checks performed, their results and what they actually support.
- Review findings, their resolution and unresolved limits.
- The human acceptance decision, ownership and relevant recovery approach.
This can be a short document for a small change. It should point to evidence rather than repeat every agent message. It should also make clear when a reviewer is looking at a newer version than the one the checks covered.
A workflow with a way back
An engineering workflow with a way back
Human owner: direction, authority, trade-offs and acceptance.
- Brief
Define the requirement and evidence needed for acceptance.
- Architecture and reasoning
Explore alternatives, assumptions and failure modes.
- Human design decisionHuman decision
Choose the approach and authorize its scope.
- Implementation
Produce an inspectable artifact and supporting checks.
- Critical review + Verification
Challenge the claims and check the behavior. Combine the evidence; these roles need not run in parallel.
Insufficient evidence → return to Implementation. Changed requirement → return to Brief.
- Comprehension check
Trace, predict and explain the important decisions.
Gap remains → learn and check again. Design problem → return to Architecture and reasoning.
- Human acceptanceHuman decision
Accept → record the result, evidence and ownership.
Unresolved trade-off → return to the relevant decision.
Roles can share a session. Different models do not guarantee independent errors.
If evidence is insufficient, return to the relevant work. If the requirement is wrong, revisit the brief. If I cannot explain an important decision, use the artifact and evidence to learn it before acceptance.
The process should get smaller when it can. A new role that merely restates the previous answer adds little. A check that cannot affect a decision may be unnecessary. I want enough structure to expose consequential errors and preserve ownership, with room for the model to reason within its task.
I am comfortable saying that AI did substantial implementation work. My claim is about the problem I framed, the architecture and decisions I owned, the process I directed and the result I can defend. That account is useful only when I can show the evidence and acknowledge its limits.
The six-principle framework puts this workflow beside the habits that keep it manageable. The open questions include the parts I have not settled, especially how to measure review independence and sufficient understanding.

