This is the part of agentic coding I spend the most time thinking about.
The most visible part of agentic coding is the agent writing code. I spend much more time thinking about what surrounds that moment: where the change should live, what the agent needs to understand before making it, and what evidence would convince me that it works. The previous three posts followed that question through skills, evaluations, and the choice between instructions and structure. Bujhchi is where those choices first became practical problems for me.
It was the first properly large software project I had built, around 70,000 lines of code, and I was building much of it through agents. I did not begin with a theory of how to organize this kind of work. Mostly, I kept getting hurt. An agent would make a mistake, so I would change the process. Something would drift, so I would add a check. Two models would agree on something wrong, so I would reconsider what each was allowed to see and decide.
Some of those failures were especially instructive. An agent could build on a broken earlier step without recognizing that the later work had lost its foundation. A document could look authoritative while describing a database that no longer existed in that form. A passing evaluation could be real and still say nothing about the current grading prompt, because the prompt had changed since the evaluation ran.
Each problem forced me to make something explicit that I had previously expected someone to remember. The work gradually became less about getting the agent to produce a patch and more about arranging the conditions under which a patch could be trusted. I was learning that the model is only one component of the coding system.
That did not remove the difficult decisions. Someone still had to decide what the product should do, which tradeoffs were acceptable, where state belonged, and what failures mattered. An agent could help me think through those questions, but it could also implement a bad answer remarkably quickly. Giving it more responsibility made my understanding of the system more consequential.
Peter Naur’s Programming as Theory Building helped me put a name to that understanding. His argument is that programmers develop a theory of the program: how it relates to the problem, why it is organized as it is, and how it should change. Code and documentation do not fully preserve that understanding. Keeping the repository after losing the people who built it does not give a replacement team everything those people knew. Naur’s paper.
I find this convincing, though large systems complicate it. No single person may understand the whole thing. One engineer knows a subsystem deeply; another understands its dependencies. Even experienced people work in unfamiliar areas. They have to construct enough understanding to make a decision, then check whether that decision holds.
Sean Goedecke connects theory-building to coding agents: their investigations can show hypotheses being formed, tested, and revised, while retaining that understanding across runs remains a problem. He also argues elsewhere that partial understanding is an ordinary condition of working on large systems. His essay on agents and theory-building and his defense of partial understanding helped me connect these concerns.
A fresh agent makes the problem hard to ignore. It is like an engineer who joined five seconds ago: capable, quick to read, comfortable with tools, and missing the history of this particular system. It does not automatically know which obvious solution was already rejected or why an awkward abstraction exists. Whatever the previous run understood has to be retained somewhere useful or reconstructed.
That gives me a concrete target for codebase quality. I want an unfamiliar agent to be able to make a good local decision without first understanding everything.
find the relevant area
↓
construct a local understanding
↓
understand the surrounding contracts
↓
make the change
↓
verify that it still fits
This is why boundaries matter to me. If I am changing module A, I should not need to understand every implementation detail of module B. I need to know what they exchange, what each promises, and what would reveal a broken promise. The boundary limits how much unrelated knowledge I must reconstruct before acting.
I think of it like a jigsaw. The edges give me information about how a piece fits even when I cannot see the whole picture. The analogy has limits—software contracts rarely capture everything—but it expresses what I want the architecture to do: make local reasoning useful without pretending it is complete.
Bujhchi’s consumer registry came from a failure of that local view. A public surface could have consequences scattered across the project, and an agent working on one part might not find every place that depended on it. I made an explicit index from important surfaces to their consumers so the agent had somewhere to begin investigating the effect of a change.
That moved the question beyond “where is this code?” to “who relies on it?” Public RPCs, exports, storage paths, and shared schemas all acquire dependencies. Some relationships can be recovered through types, static analysis, or code search. Others are easier to miss because the connection is expressed through a name, a convention, or domain knowledge.
The registry was one way to expose those relationships. I do not think every project needs that exact artifact. I wanted important dependency knowledge to be available before an agent made a change, instead of waiting for a distant failure to reveal it.
Even when dependencies are visible, the agent still needs to decide where new behavior belongs. Consider a repository where these are all plausible destinations:
utils/
helpers/
services/
lib/
common/
core/
shared/
The names alone do not tell it much. It may find an example nearby and follow it. If several conflicting patterns exist, it has to choose among them using whatever context it happens to have.
Sometimes we call the result “AI slop” when the codebase offered seven locally reasonable options and the agent chose number four. That does not excuse the result. It gives me another place to fix the problem.
I want the conventional path to require fewer decisions than the shortcut. A feature should have an intelligible home. Dependency direction should be visible. The relevant tests should be discoverable from the code they check. If two modules communicate, their contract should be more precise than a comment saying that everyone needs to be careful.
Suppose module A produces Foo and module B accepts Foo. I do not want the actual arrangement to be MaybeFooExceptSometimesBarBecauseLegacy, understood only by whoever has been around longest. Types, schemas, and boundary validation can carry parts of that agreement. I can then reason about the relationship between the modules without manually remembering every field-level condition.
Architecture can affect the speed of the work as well as its clarity. Cursor reports that restructuring its experimental browser into self-contained Rust crates reduced compilation waiting and increased throughput severalfold. Cursor’s self-driving-codebase account.
That makes architecture’s role unusually tangible. Time spent resolving structure, compiling unrelated code, or navigating inconsistent patterns is time the agent cannot spend on the change. A more capable model may tolerate more of that friction, but I would rather remove friction that serves no purpose.
The previous post argued for moving hard rules out of prose when possible. Here, that becomes part of the architecture. If service code must not import UI code, the system should have a way to reject that dependency:
ForbiddenDependency
service → ui is not allowed
Move this behavior behind the domain boundary.
See docs/architecture/dependencies.md.
The error does more than say no. It identifies the violated rule and gives the agent a route back to the intended design. OpenAI describes using custom linters and structural tests this way, including remediation instructions in the errors. OpenAI’s harness-engineering account.
Cursor also reports better results from explicit boundaries than general reminders, though its examples are prompt constraints, not mechanical enforcement. That distinction matters: stating a boundary clearly and making violations fail are different levels of protection. Cursor’s discussion of prompting.
For my own work, a repeated review comment is a reason to investigate whether the rule can move into structure. Perhaps lint can detect it. Perhaps the type is too weak. Perhaps an API should not be available from that layer. Perhaps the correct operation needs to be easier to use.
I still have to decide which rule deserves enforcement. A checker will faithfully preserve a bad rule too. But once the rule is clear, asking every fresh agent to remember it is unnecessary work.
There is a similar problem with project knowledge. Bujhchi’s AGENTS.md grew to around 77 KB. Every addition had a reason, which is exactly how it became so large. One lesson mattered, then another, then another model needed something stated differently. Eventually, the document meant to orient the agent had become a substantial amount of context it had to process regardless of the task.
The map from the previous post was my response to that mistake. OpenAI describes a similar move from a large instruction file to a short index into repository knowledge; this also connects with the context pointers I discussed through Matt Pocock’s work. OpenAI’s account.
An index for a project with these concerns might look like this:
AGENTS.md
Architecture → docs/architecture/
Database → docs/database/
Grading → packages/grading/
Evals → packages/grading/evals/
UI → ui/
Research → docs/research/
The paths here illustrate the arrangement. What matters is that the agent can reach the relevant source without first reading the whole project manual.
But discoverability is only half the problem. A beautifully organized document can still be wrong. In Bujhchi, I had documentation describing the database, while the database continued to change. I wrote a check that reads the Supabase migrations, extracts the tables, and compares them with the documented list in both directions. A mismatch makes the check fail.
migrations change
↓
compare with documented tables
↓
disagreement?
↓
CI failure
That check does not certify every claim in the database documentation. It establishes a particular relationship between the document and the migrations. That limited, explicit guarantee is useful. “The documentation exists” tells me much less.
This is the kind of knowledge I increasingly want around an agent: discoverable, specific about its authority, and connected to something that can be checked. Historical reasoning still needs prose. Facts that can be generated or compared against the system should not depend entirely on someone remembering to update a paragraph.
Authority matters when multiple representations exist too. For Bujhchi’s visual work, I made the actual JSX screens and design tokens the visual source of truth. A design document could describe direction, but it did not get to silently compete with the thing that rendered.
Otherwise, an agent could find:
design notes say A
implemented UI says B
and have to guess which one I intended it to follow. Sometimes that disagreement represents a requested change; sometimes it is stale documentation. The environment needs to make the distinction visible. Merely giving the agent both artifacts does not resolve it.
Once the agent knows where to work and which information to trust, it still needs a way to observe the result. This is where I became much more demanding about verification.
For a web application, I want the agent to use the application: open it, navigate, click, inspect browser errors and requests, read logs, and capture screenshots or traces. For an iOS application, that may mean a simulator. In another environment, it may require a debugger, a sidecar process, or a custom inspection command.
The question is what evidence I would use if I were checking the change myself. Can the agent obtain it? If it can only read source and compile, I should not be surprised when its confidence exceeds what it has observed.
OpenAI made its application bootable per worktree and exposed browser control, logs, metrics, and traces to Codex. Those capabilities let the agent investigate runtime behavior itself. OpenAI’s description of the environment.
I want the working loop to reach that evidence:
change
↓
run
↓
observe
↓
check against the intended behavior
↓
wrong?
↓
investigate and change again
That is a stronger loop than “edit, then explain why the edit should work.” It does not guarantee correctness, but it gives the agent an opportunity to encounter a contradiction before handing the work back to me.
The tools from the previous post matter here because obtaining evidence can itself become laborious. If every investigation begins with finding the development server, discovering a port, creating a test account, and searching for the right logs, verification is expensive before the agent reaches the actual question.
A useful interface might offer:
app doctor
app seed
app login test-user
app inspect request-721
app logs --request request-721
app browser snapshot
app trace
These are possible commands, not a claim that every one exists in Bujhchi. They show the level of operation I want available. The skill can guide how evidence is used; the tool can make that evidence practical to obtain.
The SWE-agent researchers call this the Agent-Computer Interface, or ACI. Their work treats interface design as part of agent performance, using features such as bounded file views, concise search results, and editing checks. These are small changes to how the agent encounters the system, but they affect what it has to infer and how quickly it can recover from mistakes. SWE-agent’s ACI documentation.
I like CLIs when they make those operations predictable. Useful flags, examples in help, clear errors, safe reruns, and dry-run behavior all make a command easier to discover and check. Cursor’s cli-for-agent plugin collects guidance along these lines. The official plugin.
Output deserves the same care. This may be enough for a person glancing at a terminal:
Everything looks good!
Database ✓
User ✓
Environment ✓
For a tool another program or agent will consume, I may prefer explicit fields:
{
"database": "connected",
"user": "present",
"environment": "development"
}
And failure should give the next investigation somewhere to begin:
{
"status": "failed",
"reason": "test_user_missing",
"recoverable": true,
"next": "app seed --user test"
}
JSON is not automatically the best format. Anthropic’s tool-design guidance emphasizes returning useful context and choosing representations for the task; it also describes consolidating repeated retrieval into higher-level operations such as get_customer_context. Anthropic’s guidance.
The point for me is to avoid making the model infer a fact the interface could state directly. If an ordinary program can collect and filter the evidence, the agent need not carry every intermediate log dump into its reasoning.
Cheap verification also changes how I sequence implementation. In Bujhchi, I broke work into blocks with their own validation. If a block failed, dependent work stopped. The workflow acquired resumable state, limited repair rounds, separate reviews, and terminal checks.
At the time, I was not thinking in terms of an elegant engineering principle. I was thinking:
Why the hell would I let block six continue when block three is already false?
If later work assumes an earlier result is correct, delayed verification makes the eventual failure harder to locate and more expensive to repair. I wanted a check at the point where that assumption began to matter.
known-good state
↓
implement one block
↓
validate
↓
green?
yes → continue
no → repair or stop
PStack’s Sequence Work into Verifiable Units gave me clearer language for this: each unit should end in a check before the next dependent unit proceeds. The principle.
This was a recurring experience while building Bujhchi. The practical pain came first; the name for the principle came later. Reading other people’s work helped me distinguish a useful general idea from a mechanism I had built for one particular failure.
A passing check introduced another problem, though. It could be valid evidence for an earlier version of the system.
Bujhchi’s grading evaluations made this concrete. If I changed the grading prompt and reused an old passing report, the report did not establish anything about the new prompt. The result had not become false. I was applying it to the wrong thing.
So the evaluation report became tied to a fingerprint of the prompt it evaluated. A changed prompt meant the old report no longer satisfied that gate.
prompt evaluated → fingerprint A
current prompt → fingerprint B
A ≠ B
The old report does not certify the current prompt.
A fingerprint does not make the evaluation comprehensive, and the prompt is not the only thing that can affect behavior. But it closes a particular way stale evidence could be mistaken for current evidence.
The broader question applies to any automated result. Which commit, model, configuration, dataset version, environment, or migration state produced it? What exactly does this check establish? The more readily agents generate reports, the more I care about keeping those answers attached.
I also became uncomfortable with one model planning, implementing, and serving as the only reviewer of its work. For riskier tasks, I separated the roles. One version of the process looked like this:
contract writer
↓
independent specification audit
↓
hidden-test writer
↓
locked contract
↓
implementer, without the hidden tests
↓
contract review
↓
safety review
↓
terminal checks
↓
me
The exact model assignments changed. The important part was who could see what and which decisions they owned. The implementer worked from the contract without seeing the hidden tests. Review was separate from implementation. A positive model review did not cancel an objective failing check.
This does not make the reviewers independent in every sense. Models can share blind spots, and several models agreeing is not proof. Separating roles gives me different opportunities to find an error; it does not let me stop checking the evidence.
Nor does every change deserve the full process. A tiny edit should not pay the same coordination cost as a risky change to grading or a shared contract. I ended up scaling the workflow with the consequences of being wrong.
Cursor’s experiments reinforce that these arrangements are provisional. Its team tried shared coordination, changed agent roles, and removed bottlenecks as observed failures demanded. It also accepted some temporary errors to preserve throughput—a different tradeoff from my dependent-block gate. Cursor’s account.
I do not want to copy an organization chart and treat it as the answer. I want to understand which failure a role prevents, what overhead it introduces, and whether a simpler arrangement now works. The workflow has to earn its place just as a sentence in a skill does.
That includes checking the workflow itself. My research processes acquired canonical paper identifiers, deduplication, restrictions on which models could perform particular jobs, protection for canonical documents, and isolation when resolving disputes. Those mechanisms were responses to errors in the process around the work.
The orchestration loop is software. The verifier is software. The documentation checker is software. Each can fail. Moving a responsibility out of the model gives it a different implementation, not immunity from mistakes.
Running several agents adds another source of uncertainty: they may change the world each other is inspecting. Two agents editing the same worktree can invalidate one another’s assumptions. Shared test data can make one agent’s passing test depend on another agent’s setup. Mixed logs can conceal which run produced the evidence.
Where independent work is useful, I want the corresponding environments to be independent too:
agent A
worktree A
runtime A
logs A
test state A
agent B
worktree B
runtime B
logs B
test state B
A separate checkout alone is not enough if both applications still modify the same database. The isolation has to extend to the state the task depends on. Then each agent can change, observe, and verify its own environment before the results are integrated.
This connects back to local understanding. I can only reason usefully about a slice of the system if that slice remains coherent long enough to investigate it. Isolation limits the number of unexplained changes an agent has to account for.
Even after integration, the work is not finished in quite the way it used to be. The repository becomes an example for the next agent. A workaround has a second cost: it can be copied.
One agent finds an awkward implementation and follows it. Another finds two examples and interprets them as a convention. A pattern that began as an exception gradually becomes the easiest available precedent.
OpenAI describes recurring cleanup guided by shared principles to reduce this accumulation. Its discussion of codebase maintenance.
That is why I think agent-heavy repositories need gardening. The question is not only whether a piece of code works today. It is also:
What will the next agent learn by copying this?
A larger context window does not settle that question. More available information may help, but the information can still be stale, contradictory, or misleading. The agent needs a way to distinguish the pattern it should follow from the accident it happened to find.
This brings me back to my 77 KB instruction file. I had accumulated useful knowledge without organizing it well enough for use. The same mistake can happen in code, documentation, tool output, and workflow. Adding more of any of them is easier than deciding which part should be authoritative and which part should disappear.
What I now look for in a codebase is whether those pieces support one another. The map should lead to the relevant boundary. The boundary should expose the contract. The tools should make the result observable. The checks should establish specific properties of the version under review. Parallel work should preserve enough isolation for the evidence to mean something.
Bujhchi was where I began learning to arrange those relationships. Some of what I built helped. Some went too far. I do not see the elaborate workflows or the size of the documentation as achievements by themselves. Their value is whether they let an unfamiliar agent make a useful change with less supervision and give me better grounds for accepting it.
There is still a version of agentic coding I do not buy: understand as little as possible and let the model handle everything. I can delegate much more implementation than I used to, but that makes it more important to understand the decisions the implementation serves.
I do not need to hold every field check or function body in my head. Types, schemas, tests, and tools have always let programmers work above some details. Agents move that boundary again. I want to use the freed attention to understand state, ownership, data flow, failure, and the evidence required at important boundaries.
My understanding can become less detailed in one place and more useful in another. What I want to avoid is accepting an implementation because I no longer know enough to question it.
That is also where the connection to research becomes more interesting to me. You described code where understanding one strange line of Python can require several papers, the mathematics, the experiment configuration, the active assumptions, and the history of an implementation choice.
In that setting, the relevant system extends beyond the source files:
code
papers
experiment history
researcher memory
An agent arriving at the repository has to find a way into those relationships. A larger skill might help with part of the problem, but it cannot substitute for making the research itself navigable.
A map for a method such as QAOA could connect the implementation to its assumptions, literature, experiments, and checks:
QAOA
implementation → src/qaoa/
assumptions → docs/qaoa/assumptions.md
papers → docs/qaoa/papers.md
experiments → experiments/qaoa/
verification → tools/qaoa-check
This is an illustrative layout. You would know which distinctions matter in the actual project. The useful property is that a new person or agent could follow a line from the code to the reasons for it, rather than trying to infer the whole method from the implementation.
The same environment could expose operations at the level researchers use:
experiment inspect run-742
experiment reproduce run-742
experiment compare run-742 run-891
experiment check-invariants run-742
Inspection could gather the relevant configuration and results. Reproduction would need the recorded code, data, arguments, and environment. Comparison could expose the differences between runs. An invariant check would need the assumptions and tolerances under which the property is expected to hold.
If a particular state should remain normalized, for example, a check could report:
InvariantViolation
expected norm: 1.000000
observed: 0.973141
outside the configured tolerance
That only makes sense where normalization is the appropriate invariant. Deciding that is part of the expertise being encoded. Once the condition is established, the system can detect a violation without depending on the agent to remember the check.
Your judgment would still matter in explaining why it failed. Perhaps the result points toward a numerical issue, an invalid assumption, or a misunderstanding of what the implementation represents. I would want a skill to guide that investigation where you can explain the procedure, while the environment carries the recorded state and repeatable checks.
The software lessons become especially useful when these pieces meet. A check should identify which run it examined. A comparison should expose which parameters changed. A reproduced result should be tied to the version that produced it. An agent investigating one experiment should not silently modify the state another investigation depends on.
I would start with one workflow where you already know how to judge the work. Watch the agent attempt it. When it gets lost, expose the missing relationship. When it repeats a mechanical sequence, consider a tool. When it misses a check, see whether that check can be made executable. When it reaches the wrong conclusion despite having the evidence, examine the reasoning and the procedure.
real task
↓
observe where the agent struggles
↓
change one part of the environment or procedure
↓
run again
↓
compare what happened
That brings the first three posts together. Skills carry a procedure. Evaluations help establish whether a change helps. Maps, tools, and constraints carry responsibilities that should not depend entirely on a fresh conversation. The codebase and workflow determine whether those pieces can function as one system.
Cursor reports a research run peaking near 1,000 commits an hour across about ten million tool calls in a week, while still crediting humans with taste, judgment, and direction. Its conclusion.
I do not take that as a permanent division of labor. Some judgments will become checks; some procedures will become tools. At any given point, though, someone has to decide what can be carried reliably by the system and what still needs closer attention.
That is the question I want to explore with you. What would a research environment look like if an unfamiliar agent could find the relevant assumptions, work within a clear boundary, run the necessary checks, and leave evidence you could inspect? Which parts of your attention would that free, and which parts would become more important?
The aim I keep returning to is simple:
Make good decisions easy to make, bad decisions hard to make, and correctness cheap to check.
Less politely: make the codebase so good that even an agent having a bad day has a decent chance of doing the right thing—and a clear way to discover when it has not.