Agentic Engineering / September 2026
I Was Still the Architect. AI Made Me Faster.
How I used a dedicated AI expert to help design, implement, and test changes across a multi-repository agentic engineering system without giving up architectural authority.

For several weeks, Pipeline Fleet Architect and I worked on ARF almost every day.
Sometimes we made several changes in a single day.
The loop was usually simple:
problem → options → decision → implementation → test → discussion
What changed over time was not who owned the architecture.
I did.
What changed was how much of the work between an architectural decision and a tested implementation I had to do manually.
Pipeline Fleet Architect — PFA — became the tool I used to make that loop much faster.
I did not want another coding agent
ARF is an agentic engineering system I built at Augment Code around a real software-delivery workflow: support intake, diagnosis, reproduction, fixing, independent verification, review, and a mergeable pull request.
As the system grew, changing ARF itself became a substantial engineering task.
A behavior that looked like a prompt problem might actually belong in a shared skill. A failed run might be caused by the execution environment. A repeated workaround might need to become deterministic code. A change could span several repositories.
I could have continued doing all of that manually.
Instead, I built a specialist for the job.
PFA was not intended to decide what ARF should become. It was designed to help me understand the system, explore options, and implement the changes I chose.
That distinction mattered enough that it became part of its prompt.
When I described a problem, the default was not:
Go change the system.
It was:
Investigate. Understand the problem. Give me options.
Sometimes I chose one of those options.
Sometimes I combined two.
Sometimes I added an option PFA had not suggested.
Then I told it what I wanted implemented.
I was the expert.
PFA was the tool that made me more efficient.
Give the tool enough context to do real architecture work
PFA was useful because it did not see ARF as a pile of YAML files.
It had access to the evidence produced by the running system.
That included expert-session transcripts, the shared virtual filesystem, stage deliverables, ledger and state artifacts, JSONL execution traces, Linear issues, GitHub repositories, pull requests, and the code and definitions spread across multiple repositories.
That meant a conversation about a failed run did not have to begin with me reconstructing everything that had happened.
PFA could inspect the evidence itself.
It could compare what an expert was instructed to do with what it actually did. It could look at the artifact handed to the next stage. It could inspect the repository change. It could see the PR. It could follow the resulting state through the pipeline.
For architecture work, that is a very different level of context from pasting an error message into a chat window.
If you want AI to help debug an AI system, give it evidence, not a postmortem.
The first job was understanding where a change belonged
One of the more useful effects of working this way was that every problem did not automatically become a prompt edit.
Suppose an expert needed a tool that was not available in its environment.
A weak solution would be to add more instructions:
If the tool is missing, install it.
Now every future run spends time rediscovering the same setup problem. Different models may solve it differently. One may install the right package. Another may choose a workaround. Another may burn tokens and still fail.
The better question is architectural:
Should this capability already exist in the environment?
If the answer is yes, fix the environment.
The same applies to skills.
If an agent has to repeatedly reason through a procedure we already understand, perhaps that procedure should become a reusable skill.
And if a behavior must be consistent, perhaps the right answer is not a stronger prompt at all. Perhaps it belongs in deterministic code.
Over those weeks, PFA and I made changes across all of those layers.
Some were environment changes.
Some were new skills.
Some changed expert instructions or responsibilities.
Some changed supporting code.
The point was not to maximize how much AI was involved.
The point was to put each behavior in the right place.
Architecture across repositories
ARF did not live in one neat repository.
Different parts of the system had different owners and different deployment paths.
That turned out to be important.
When a problem appeared in one stage, the correct fix was not necessarily in that stage's definition.
A symptom in an expert session could originate in a shared skill.
A missing capability could belong in an environment image.
A workflow problem could belong in orchestration.
A deterministic invariant could belong in supporting code.
So PFA had to understand more than what to change.
It had to understand where the behavior was owned.
Depending on the repository and type of change, it could make a direct change or prepare a pull request for review.
For me, that is one of the clearest differences between using an AI coding assistant and using AI as an engineering tool at the architecture level.
The coding question is:
Which file should I edit?
The architecture question is:
Which component should own this behavior?
Those are not the same problem.
My review moved up a level
I was not spending my time manually editing every prompt, script, or configuration file.
That did not mean I stopped reviewing the work.
The review changed.
I talked with PFA about the solution.
Why did it choose this layer?
What alternatives did it consider?
Why should this become a skill instead of a prompt rule?
Why change the environment instead of letting the expert install the tool?
How would the change affect the rest of the pipeline?
How would we test it?
The implementation mattered, but the conversation was primarily about whether the implementation represented the architecture I wanted.
That is an important distinction.
I was not delegating the architectural decision. I was delegating the work required to explore and implement it.
For a senior engineer, architect, or technical lead, that is a much more interesting use of AI than asking for another code completion.
"Don't implement anything yet"
One behavior became important enough that I made it explicit in PFA's instructions.
Sometimes I wanted investigation only.
No edits.
No PR.
No "I noticed the problem, so I fixed it for you."
Just understand the system and come back with options.
This sounds minor, but it changes the working relationship.
A capable agent can move from observation to implementation extremely quickly. That is often useful. It can also collapse two very different decisions into one:
- What is wrong?
- What should we do about it?
I wanted those separated when the architectural choice mattered.
PFA could investigate broadly, inspect the evidence, trace ownership across repositories, and present possible approaches.
Then I could make the decision.
Only after that did implementation authority follow.
The faster the implementation becomes, the more important that boundary becomes.
The prompt evolved with the working relationship
PFA's own prompt did not arrive fully formed.
It evolved over the weeks we worked together.
That is another thing I think gets lost in discussions about AI agents.
A good expert definition is not just a personality description.
It is accumulated operating experience.
If PFA implemented too quickly when I wanted analysis, that became a clearer rule.
If it needed better access to a particular kind of evidence, we improved that.
If a planning pattern consistently worked well, it became part of how we collaborated.
In that sense, the prompt was partly an interface between my engineering style and the system.
The better the interface became, the less time I spent correcting the mechanics of the collaboration and the more time we spent on the engineering problem itself.
This was not autonomous self-modification
There is an obvious way to tell this story badly:
I built an AI that rewrote its own AI system.
That sounds exciting.
It is also not how I worked.
PFA could change ARF when I instructed it to.
It could inspect the system, propose changes, implement them, run tests, work across repositories, and prepare or make changes depending on the repository and task.
But it did not own the direction.
A failed run was not permission to rewrite the pipeline.
An observation was not a decision.
A proposed improvement was not approval.
I remained the authority deciding what should change.
That boundary was deliberate.
The more capable the tool became, the more useful it was to be explicit about where its authority ended.
What changed for me
The biggest change was not that I typed less code.
It was that the distance between identifying an architectural problem and seeing a tested implementation became much shorter.
I could spend more time on questions such as:
Is this the right abstraction?
Is this really a model problem?
Should this knowledge become a skill?
Is this behavior owned by the correct component?
Are we fixing one run or improving the system?
What evidence will tell us whether the change worked?
PFA handled much of the investigation and implementation work needed to answer those questions in practice.
That is where I think AI becomes especially interesting for experienced engineers.
The value is not only generating code faster.
It is compressing the feedback loop between intent, architecture, implementation, and evidence.
The architect did not disappear
I did not stop engineering ARF.
I stopped needing to be the person who manually made every edit.
Those are very different things.
My role moved further toward defining the problem, evaluating options, choosing the architecture, challenging the solution, and deciding what good looked like.
PFA made that work faster because it could operate on the system directly instead of waiting for me to translate every decision into implementation details.
For me, that is a more useful way to think about AI and senior engineering.
The goal is not to remove the expert from the loop.
It is to give the expert better leverage.
I was still the architect. AI made me faster.

