For my cousin. I spent the last two posts talking about skills. This one is about when I think you should stop writing them.

Once you understand skills, everything starts looking like a skill issue. The agent cannot find something, so you write a skill. It keeps debugging badly, so you add a procedure. It violates an architectural rule, so you add another instruction. It cannot inspect some awkward internal state, so another section appears in SKILL.md.

Eventually, the file looks like this:

IMPORTANT: do this.

VERY IMPORTANT: never do that.

If this happens, inspect X.

Then find Y.

Run Z.

If that fails, read this file...

Then...

Every line probably has a reason. Some agent made a mistake, and you added a sentence to prevent it from happening again. In the previous post, I asked how to test whether those sentences help. But even an instruction that passes that test leaves another question open: is an instruction the best place for this lesson to live?

Suppose the agent follows an eight-step procedure correctly every time. That sounds like a successful skill. Yet if those eight steps are always the same mechanical work, I may have taught the model to do something a small program could handle more reliably. The evaluation tells me the instruction helps. It does not tell me I should keep solving the problem through instructions.

That is the question I now ask when a skill keeps growing:

Why am I making a probabilistic model remember all of this?

Sometimes the agent needs a better procedure. Sometimes it cannot find the information, perform the operation, or detect the mistake with the environment I have given it. Those can all look like failures to follow instructions, especially when adding instructions is the first remedy I reach for.

I was grouping too many of these problems under the word skill. If I want the agent to approach a class of problems in a particular way, a skill makes sense:

Reproduce the failure before changing code.

Gather evidence before committing to an explanation.

Check the limiting case before trusting the result.

These instructions guide how the agent investigates and decides. But “the authentication code starts here” does a different job. So does a command that creates a test account. So does a check that rejects a forbidden dependency.

I find it useful to separate four needs:

The agent needs… Give it…
A procedure or a way to exercise judgment A skill
A way to find relevant knowledge A map
A repeatable action or observation A tool
A rule the system must enforce A constraint

These can work together. A skill may point to a map, call a tool, and explain what to do when a constraint fails. The distinction helps me decide which part should carry the work.

Take the simplest case: the agent cannot find something. My instinct used to be to put the missing information somewhere it would always see it. If authentication matters, explain authentication in AGENTS.md. Then payments matter. Then experiments. Eventually, the document contains plenty of useful knowledge, most of which the current task does not need.

OpenAI describes trying a large AGENTS.md and then moving toward a short index into a structured knowledge base. Its advice is to give Codex a map rather than a thousand-page manual. OpenAI’s harness-engineering account.

A map might look like this:

AGENTS.md

Architecture → docs/architecture/
Auth         → src/auth/ + docs/auth.md
Payments     → src/payments/
Frontend     → docs/frontend/
Experiments  → docs/experiments/

The agent still has to read and understand the relevant material. But it no longer has to discover where that material lives before it can begin. I already knew where authentication starts. There is little value in making each fresh run reconstruct that fact through a sequence of searches.

Some of what it needs will be harder to locate because it was never recorded properly. A strange boundary may make sense only if you remember a production incident. An obvious approach may have been tried and rejected in a conversation six months ago. A pointer cannot recover a reason that exists only in someone’s memory; there needs to be something useful behind it.

That does not mean documenting every thought anyone had while building the system. I want to preserve the reasons that would change a future decision and make them discoverable from the place where that decision arises. The map’s job is to get the agent to enough context to work, without requiring it to carry the whole project at once.

Finding the right place, however, does not solve the next problem. Once the agent gets there, it may still have to repeat the same mechanical sequence every time.

Imagine an Electron app where each debugging session begins like this:

start the dev server
find the Chrome debugging port
connect through CDP
find the correct page
create a test user
navigate to settings
find the relevant logs
capture a trace

I could write a careful skill explaining all eight steps. Every new agent would then get the privilege of performing the ritual again. If the setup changes, I would revise the recipe and hope the next agent interprets it correctly.

At some point, I want to ask:

Do I have eight instructions, or one missing tool?

This is where PStack’s Build the Lever principle helps me. It recommends doing an initial unit manually to understand the work, then building a script, generator, or other artifact that can perform or verify it again. The tool itself becomes something a reviewer can inspect and rerun. PStack’s “Build the Lever.”

For the debugging example, I might expose commands like these:

app doctor
app seed
app login test-user
app snapshot --json
app logs --since 30s --json
app screenshot
app trace start
app trace stop

These are illustrative commands, but the division of work is what interests me. The CLI handles starting the environment, preparing test data, and retrieving evidence. The skill can concentrate on the investigation:

Reproduce the problem before editing.

Gather evidence for the proposed cause.

Do not commit to the first explanation.

After the fix, repeat the original reproduction
and check whether the failure is gone.

The agent still has to decide which evidence matters and what it supports. It no longer has to reconstruct every operation required to obtain it. The debugging procedure becomes smaller because some of its work now has an implementation.

This applies to observation as much as action. Suppose I want the agent to establish the current database state. I can tell it to find the migrations, determine which ones ran, locate the configuration, connect, query several tables, and combine the results. Or I can give it:

app db status --json

If the operation is well defined, ordinary code can perform it. The model can then reason about the result instead of spending its context on discovering and coordinating the steps.

Anthropic makes a similar argument in its tool-design guidance. Tools need not correspond one-for-one with low-level API calls. Its customer example combines profile information, transactions, and notes into a single get_customer_context operation. Anthropic’s tool-design guidance.

For a recurring task that needs all three, the difference is easy to see:

get_customer()
get_transactions()
get_notes()

becomes:

get_customer_context()

The computer does the retrieval and combination. The agent receives the view it needs for the task. I would not combine unrelated operations simply to reduce the number of tools, but when the same sequence keeps appearing, it is worth asking whether the interface stops one level too low.

A good tool can also keep irrelevant material out of context. Without one, the agent might find a process, read a configuration, dump hundreds of log lines, locate a request ID, search again, and query the database. Every intermediate result becomes more material to interpret.

A purpose-built operation could instead return:

{
  "request_id": "r_721",
  "status": "failed",
  "failure": "test_user_missing",
  "relevant_logs": [
    "AuthError: no matching test user"
  ],
  "next": "app seed --user test"
}

That is what I mean when I think of tools as context compression. The tool performs the mechanical work before the output reaches the model. It can retain a route to fuller evidence when needed, while making the ordinary case much easier to inspect.

But a reliable operation still leaves a question: will the agent remember to use it? If a check protects a rule that must hold every time, making it available as a command may not be enough.

Suppose the skill says:

Always make sure API responses match this schema.

That instruction may help. Yet if the schema can be checked at the boundary, I can make validation part of the operation itself. The agent then encounters a failure when the response is invalid, rather than needing to remember to look for one.

Or suppose the rule is:

Never import persistence code directly into the UI layer.

If I keep repeating that in reviews, I want to know whether a dependency check can reject the import. OpenAI describes using custom linters and structural tests to enforce architectural boundaries, with error messages that explain how to repair violations. OpenAI’s account of architectural enforcement.

An error in my own project might look like this:

ForbiddenDependency

UI → persistence is not allowed.

Use the service boundary.
See docs/architecture.md#layers.

Now the rule has a way to detect a violation, reject it, and point toward the permitted route. Prose can still explain why the boundary exists. It no longer carries the entire responsibility for keeping the boundary intact.

PStack calls this Encode Lessons in Structure: when a correction keeps recurring, consider whether it should become a lint rule, runtime check, metadata flag, or script. PStack’s principle index.

That gives me the rule tying these cases together:

The more deterministic something is, the less I want it living in prose.

The qualification matters. I cannot turn every review judgment into a linter, and an executable check only enforces what I managed to express. I still need to decide whether the rule is right, where it applies, and what counts as an acceptable exception. But once those decisions are clear enough to implement, repeatedly reminding an agent is a weak way to preserve them.

Looking back at the debugging example, one failure could require all four pieces. The skill tells the agent to investigate before editing. The map points to the subsystem and its history. The tool gathers the relevant state. The constraint rejects a change that crosses a forbidden boundary.

If the agent fails again, I can ask where that arrangement broke down. Did it choose the wrong hypothesis despite having the evidence? That may require better reasoning or a better procedure. Did it never find the relevant document? Improve the route to it. Did gathering the evidence require reconstructing a known sequence? Consider a tool. Did it make a change the system should always reject? Add enforcement where the rule can actually be checked.

This is why I no longer want to respond to every failure by making SKILL.md more emphatic. “IMPORTANT” cannot provide an operation the environment does not support. “NEVER” cannot enforce a boundary by itself. And a longer explanation of good debugging will not help much if the agent cannot observe the state it is supposed to debug.

There are larger design questions behind this: whether a feature has an obvious home, whether a module can be changed without understanding every other module, and whether the agent can verify the running product. Those are where this approach leads me next. For this decision, though, I can start smaller: identify one repeated failure and ask what kind of help is missing.

That also changes the question I want to ask about research code. You described how understanding a tiny piece of Python can require several papers, the mathematics, the experiment setup, active assumptions, and the history behind a strange implementation choice. Putting all of that into one skill would be my old response.

Some of it probably does belong in a skill. If a suspicious result makes you consider three particular hypotheses, I would want to learn how you make that judgment. If an approximation is trustworthy under one condition and dangerous under another, the procedure needs to preserve that distinction.

But suppose part of the work is simply opening several files and matching configuration IDs. Or remembering the command, commit, and environment required to reproduce an old run. Those are candidates for tooling. A command such as:

experiment reproduce run-742

could carry the recorded setup, provided the underlying code, data, and environment are available. The agent would still need to interpret what happened when the experiment ran.

Likewise, if a mathematical property can be checked under clearly stated assumptions and tolerances, perhaps that check can become executable. You would supply the conditions under which it is valid. The system could then notice a violation without depending on the agent to remember the instruction, while your expertise would still matter in explaining the failure.

So I want to extend the question from the first post. Which parts of your work could become skills—and which parts would be better carried somewhere else? When you correct someone, are you teaching a judgment, pointing them toward missing context, walking them through a mechanical operation, or reminding them of a rule that could be checked?

I started this series by asking what my agents should remember. Then I asked how to tell whether what I taught them helped. Now I want to know what I can change so they do not need to remember it at all.

Sometimes it really is a skill issue. But before I spend another hour rewriting the instructions, I want to check whether the codebase is missing a map, a tool, or a constraint. In the next post, I’ll bring these choices together through Bujhchi, where I first had to make them: how I organized the code, gave agents ways to check their work, and revised the process when it failed.