For my cousin. I wanted to explain how I build software with agents, and ask where the same approach might be useful in physics research.
If I have to give an agent the same correction every time, then part of the work never really gets finished.
The code may be saved. The bug may be fixed. But the lesson still lives in my head. When the same problem appears a week later, I have to notice it again, remember what went wrong last time, and explain the whole thing from the beginning.
A skill gives that lesson somewhere else to live.
By a skill, I mean a set of instructions and supporting material an agent can consult for a particular kind of work.
Instead of repeatedly saying, “When this happens, do it this way,” I can turn the procedure into something reusable.
At first I thought of skills mainly as a way to make my own agents better. Then I began to see a more interesting possibility. Skills may also be a way for people to share how they work.
Suppose I have a teammate with a good idea, but they don't know how to do one part of it. I do. If I were solving the problem myself, I would check one thing first, inspect something else next, take a different route if a certain condition appeared, and verify three things before finishing.
Normally I would have to explain that whole process to them. But I could instead write down the map of how I would solve it. They could give an agent both their task and my process.
This obviously wouldn't transfer all my judgment or experience. But it would transfer part of how I work, not merely what I know.
That could be useful across a team. One person may be especially good at reviewing frontend work. Another may have a strong process for debugging databases. Someone else may know how to check a deployment properly. If even part of each process can live in a skill, the rest of the team can use it without first becoming experts in every area.
Once I started thinking of skills this way, the writing mattered more. A vague prompt might be good enough for a task I am supervising once. A shared skill has to preserve the useful parts of a procedure after the conversation that produced it has disappeared.
A lot of how I think about this comes from Matt Pocock's writing-for-agents work. His ideas about context pointers, context and cognitive load, information hierarchy, leading words, completion criteria, pruning, sediment, and no-ops gave me names for things I had already been running into.
The first problem sounds almost too simple: before a skill can help, the agent has to find it.
A skill the agent never finds
Imagine I have this skill:
---
name: frontend-review
description: Review frontend changes for visual consistency,
accessibility, responsive behavior, and interaction quality.
---
When reviewing frontend work:
1. inspect the rendered page
2. check spacing and hierarchy
3. test responsive states
...
Sometimes I want to invoke a skill myself. If I already know I want a frontend review, I can say:
/frontend-review
In that case I carry the responsibility. I have to remember the skill exists, recognize the kind of problem in front of me, and call it.
Other times I want the model to notice that a skill is relevant and load it automatically. Depending on the agent system, both options may be available. So I don't think the useful distinction is that there are exactly two kinds of skill. The useful question is:
Who carries the responsibility for invocation—the human, the model, or both?
If I carry it, I pay more cognitive load. I have another tool to remember.
If the model carries it, the model needs some information in its context telling it that the skill exists and when it applies. That creates context load.
The model doesn't need the entire review procedure in its head all the time. It needs only enough information to know that useful knowledge exists over there and when it should go and get it.
Pocock calls the description a context pointer. It points toward a larger body of information that can remain outside the working context until it becomes relevant.
This matters because context is finite. I don't want every procedure, architecture note, debugging lesson, API quirk, and style guide loaded on every turn. I want small signs that lead the agent to the right material.
But a context pointer isn't a normal function call. It is probabilistic. The model still has to interpret the description. Change the wording and you may change whether it decides to load the skill.
That means a skill can fail in two different ways:
- The agent may fail to retrieve a useful skill.
- It may retrieve the skill and then follow it badly.
Those failures need different fixes. A perfect procedure is useless if the model never finds it. Later I want to write separately about evaluating skills, including trigger and non-trigger cases. For now, the important thing is that the description isn't merely a label. It is part of the mechanism.
Solving that problem created the next one for me. Once the model finds the skill, what should it find inside?
The danger of giving it everything
The obvious answer is everything it might need. That answer doesn't survive contact with a real project.
A project may have years of conventions, examples, architecture notes, API details, and remembered mistakes. If I put all of that into one skill, the agent will find the knowledge. But it may lose the procedure inside it.
Pocock makes a useful distinction here between steps and reference.
Steps tell the agent what to do. A debugging procedure might be:
1. reproduce the bug
2. inspect the relevant code
3. form a hypothesis
4. test the hypothesis
5. implement the smallest fix
6. rerun the relevant tests
7. inspect the final diff
Reference is information the agent may need while doing those steps:
- project conventions
- architecture notes
- examples
- domain knowledge
- API documentation
- known edge cases
- checklists
- previous lessons
- design rules
Both are useful. The harder question is where each should live.
Some information belongs in the main procedure because every run needs it. Other information can sit nearby as reference. Material needed only in one particular branch can go into another file behind a context pointer.
A frontend-review skill, for example, might look like this:
frontend-review/
SKILL.md
accessibility.md
animation.md
responsive.md
The main skill could then say:
If the change affects keyboard navigation or semantic HTML, read
accessibility.md.
If the change introduces motion or transitions, read
animation.md.
Now the agent loads the extra knowledge only when it enters that branch. It doesn't have to carry every possible branch from the beginning.
This is the part of information hierarchy I find most useful. Show the common path first. Reveal the next piece when it becomes relevant. I think of it as branching context.
That led me to one of my strongest rules for skills: smaller is usually better.
Not because short writing is automatically good, but because every unnecessary sentence competes for attention. Smaller skills tend to be:
- easier to audit
- easier to update
- harder to contradict
- cheaper in tokens
- easier for the agent to follow
But “smaller” becomes another bad rule if it is followed blindly. A short skill that omits a necessary instruction is merely incomplete. The aim isn't the shortest possible skill. It is the smallest skill that still changes behavior in the way I want.
This last phrase matters, because having the right information and doing the right work are not the same thing.
The agent can know and still rush
For a while I thought that if a skill contained the right information, the agent would naturally behave the right way. But an agent can know everything I told it and still begin implementing before it understands the problem. It can accept its first explanation, read too little of the codebase, or reach the requested output through a process I wouldn't trust.
At that point, adding more reference material won't help. The problem is no longer what the agent knows. It is how the agent behaves.
One way I try to steer that behavior is with leading words: compact concepts that already carry useful meaning for the model. Instead of expressing the same idea through twenty unrelated instructions, I give the procedure a center of gravity.
Suppose a debugging agent keeps jumping to conclusions. I might organize the skill around one word:
evidence.
Start from evidence, not assumptions.
Before changing code, gather evidence for the failure.
For each proposed cause, identify the supporting evidence.
If two explanations remain possible, gather more evidence.
After the fix, collect evidence that the original failure is gone.
No one sentence does all the work. The same idea follows the agent through the procedure:
evidence → evidence → evidence
This reminds me of a leitwort, a recurring word that helps unify a piece of writing. Here the purpose is procedural. The word keeps bringing the agent back to the kind of work I want it to do.
Depending on the task, the leading word might be:
- verify
- trace
- minimal
- falsify
- reproduce
- exhaustive
The word isn't magic. It needs a concrete meaning in the instructions, and I still need to test whether it changes the behavior I care about. But it can steer more cleanly than a pile of disconnected commandments.
Leading words help when the agent is reasoning in the wrong way. A harder problem appears when it technically follows the procedure but does too little of it.
I think of this as the legwork problem.
Planning is a good example. A simple workflow might say:
- Ask clarifying questions.
- Create a plan.
That sounds correct. But the agent can satisfy the first step by asking two obvious questions and then rush toward the thing it really wants to produce: the plan.
I don't particularly like Claude Code's built-in plan mode for this. In my experience, Codex is better at it. But there is a more interesting pattern than choosing between the two.
Pocock has skills built around grilling. Instead of immediately producing the artifact, the agent keeps investigating the problem. He also has a grill-with-docs workflow, where that interrogation happens alongside the relevant project documentation.
The useful part isn't the name of either skill. It is the staging. The agent inspects the repository, asks questions, checks the documentation, and follows contradictions before it is allowed to produce a specification or implementation.
If the agent can already see “Step 2: produce the PRD” while it is doing Step 1, the later step exerts a kind of gravitational pull. The model knows where it is going, so it may decide it has “clarified enough.”
One response is to hide the later stage until the current one has met its completion criteria. “Ask questions” is easy to satisfy. “Resolve every branch that would change the implementation” demands more legwork.
An agent can also fail in the opposite direction: it keeps planning after the important uncertainty is gone, adding branches and caveats while the task never moves. This resembles the intuition behind ReAct, which interleaves reasoning with action so new observations can update the plan. Anthropic's guidance on effective agents is more explicit about feedback loops and stopping conditions. Once the next step is clear, take the smallest reversible action and let the result teach you what to plan next.
Not every workflow needs separate stages. They become useful when an agent repeatedly rushes ahead because it can already see what comes next.
And then, if the skill works, it develops a different kind of problem.
Every fix leaves a layer
An agent makes a mistake, so I add an instruction. Another mistake happens, so I add another. Someone else works on the project and adds three more. Six months later, the skill is 2,000 lines long and nobody knows which rules still matter.
Pocock has a good word for this: sediment.
A team encounters a problem and adds a rule. Months later, someone finds the rule but doesn't know whether it still matters. Deleting it feels dangerous—surely it was added for a reason—so they leave it and add their own rule underneath.
The document develops geological layers. Adding feels safe. Removing feels risky. Knowledge accumulates faster than it disappears.
So maintaining a skill can't mean merely keeping every correct instruction. I have to ask harder questions:
- Does this instruction belong in this skill?
- Does it apply to every run or only one branch?
- Does another sentence already produce the same behavior?
- Can the agent discover this directly from the environment?
- Is this reference material rather than a step?
- Was it useful once but stale now?
Depending on the answer, I can move it behind a pointer, merge it, or delete it.
There is an important difference between repeating a leading word and repeating an instruction. Repeating evidence can connect distinct stages of a procedure. Repeating the same paragraph three ways creates noise and three places to maintain the same rule.
Some instructions don't even need time to become sediment. They do nothing from the day they are written.
Consider this one:
Carefully understand the user's request before proceeding.
Would the agent have deliberately misunderstood the request without that sentence? What behavior did I buy with those tokens?
Pocock calls instructions like these no-ops: text that sounds sensible but doesn't meaningfully alter what the model would otherwise do.
That leaves me with a hard question for every line in a skill:
If I remove this instruction, does the agent behave differently?
I can't always answer that by reading the file. Two people can disagree about whether a sentence matters because they disagree about what the model would do without it. At some point, I have to run the skill and compare the behavior.
That is the subject of the next article. For now, it is enough to notice that the real artifact isn't the Markdown file. It is the behavior produced when a model reads it.
Four questions I keep coming back to
The path through all of this now seems simple, though it didn't begin that way.
First the agent has to find the skill. Once it does, the skill has to reveal the right amount of information. That information has to steer behavior, not merely describe good intentions. And as the skill changes, someone has to remove the instructions that no longer earn their place.
Taken together, Pocock's framework leaves me with four practical checks:
Can the agent find it?
- How does the agent know this knowledge exists?
- Who decides when to invoke it?
What should it load?
- What belongs on the common path?
- What can remain behind a context pointer?
Does it change behavior?
- What behavior am I trying to induce?
- Do the completion criteria demand enough legwork?
What can be removed?
- What is duplicated, stale, or a no-op?
- What can disappear without changing behavior?
These terms—context pointers, the two loads, information hierarchy, leading words, sediment, no-ops, and completion criteria—come from Pocock's writing-for-agents work. I am using his framework here to think more clearly about the skills I build.
The more I work this way, the less I think of skills as prompts.
They feel more like externalized procedures that survive individual conversations.
In the next article, I will explain how I turn real agent failures into evals, compare old and new skills across repeated runs and models, and use agents to keep that loop moving.
This is the part I actually wanted to ask you about
I think skills become most useful when you have a clear way to check a particular part of the work.
You don't necessarily need the full answer. But you need some way to check whether the reasoning is going in the right direction.
And physics seems full of cases like that.
You already told me there are things you automatically check when looking at a result.
Maybe:
does this mathematical condition hold?
does the limiting case make sense?
should this quantity be conserved?
does the scaling behave correctly?
does this agree with a known result?
if this assumption changes, what should happen?
You probably have a lot of these checks that feel obvious to you now because you have done them for years.
That is exactly the kind of thing I mean by a skill.
Not:
“Agent, be a good physicist.”
More like:
When you reach this kind of result, check these three things before trusting it.
If this condition fails, investigate these possibilities.
Before moving on, verify this mathematical property.
Try this limiting case.
Compare against this known behavior.
Basically, take some small part of the way you would approach the problem and try to encode that behavior into the agent.
Then eval it.
Does the agent actually perform the check?
Does it catch mistakes that the base model misses?
If I remove one of the instructions, does performance drop?
Does the behavior still hold across Claude, Codex, Gemini, whatever model I am using?
If it fails, improve the skill and run the eval again.
So here is what I am really curious to ask you:
Which parts of your own research already have this property?
Which parts make you think:
“Whenever I do this, I always check X, Y and Z.”
Or:
“If the result is correct, this mathematical thing should always happen.”
Or even:
“A junior researcher often misses this, but I know to look for it immediately.”
Those seem like the best candidates.
Because once you can clearly check this part of the work, you can encode the process, eval the behavior, and judge how much supervision it still needs.
Not all of physics.
Just the pieces where your own verification process is clear enough to teach and measure.
That is the part I want your opinion on.