I first heard this line from Charlie Guo of OpenAI:
Every word must justify its existence.
It may be the best rule I have found for writing agent skills, because adding instructions is so much easier than knowing whether they help. An agent does something stupid, so I add a sentence. Then it does something else stupid, so I add another. A month later, the file looks like this:
Do X.
ALWAYS do Y.
IMPORTANT: never forget Z.
Before doing anything, remember A.
CRITICAL: make sure B.
Every line has a story behind it. None looks obviously wrong. Yet I may no longer know which ones matter. The current model might already handle some of these cases. One instruction might help Claude and do nothing for Codex. Another might have prevented one failure while introducing two others. A sentence might simply sound wise.
So I keep coming back to a harsher question:
What happens when I remove you?
If nothing changes, the sentence probably should not be there. The annoying part is that I cannot establish that by rereading the Markdown and thinking really hard. I need an experiment. That is where evals started making sense to me.
A skill makes a claim about behavior, even when I have not written the claim explicitly. Suppose an agent keeps jumping into code before reproducing a bug. I give it a debugging skill that tells it to reproduce the failure first. I am now claiming two things: these instructions will change the order in which the agent works, and that change will make its debugging more reliable.
Both claims can be tested:
- Did the skill change the behavior I wanted to change?
- Did that behavior improve the result?
I need both questions because it is easy to teach an agent a beautiful ritual that does not help. The debugging skill might produce careful reproduction steps, five hypotheses, neat traces, and a professional-looking investigation. But if it solves fewer bugs, I have taught the model to perform my preferred ceremony while making it worse at the task.
What I care about is the behavioral delta: what changes after the agent reads the instructions. The document is how I deliver that intervention. Usually, the outcome tells me whether it helped, and the trace—the record of what the agent did—helps me understand why.
To measure a change, I first need a case worth changing. Hamel Husain’s emphasis on error analysis is useful here: start by watching the agent work and finding failures that matter, rather than inventing twenty abstract properties of a good agent.
Suppose the agent edits a database migration without checking how migrations work in the repository. I can preserve:
- The task.
- The starting repository state.
- The trace.
- The failure I care about.
Now I have something to return to before permanently adding another sentence to SKILL.md. The usual response is:
agent fails
↓
add rule
I want one more step before the rule:
agent fails
↓
save failure as eval
↓
change skill
↓
run the failure again
That gives me a working rule:
Every interesting failure becomes an eval before it becomes a permanent instruction.
Otherwise, skills gradually turn into folklore. “This sentence is here because six months ago some agent did something bad.” Fine, but does it still prevent that failure? Does the current model already handle it? Does it help one model and hurt another? Without the saved case, the instruction survives partly because deleting it feels dangerous.
An eval gives me a way to revisit that decision. But running the new skill and watching it succeed is only the beginning. The model might have succeeded without my change.
This is what I like about the baseline comparisons in Anthropic’s skill-creator. For a new skill, the comparison is:
A = no skill
B = skill
For a revision:
A = old skill
B = new skill
For pruning:
A = current skill
B = current skill minus the thing I want to delete
I want the same task, repository, starting state, tools, and model, with fresh context for each run. As far as possible, the skill should be the thing that changes.
Then “B succeeded” becomes a more useful question: what changed between A and B? If both solve the task equally well, I have not yet shown that the revision helps. If B prevents the target failure but introduces two new ones, I may not want it. The skill cannot claim credit for a task the model already handles without it.
A successful demo shows that the system can work. The comparison helps me decide whether my intervention mattered.
For that comparison to mean anything, I also need to avoid helping the agent through the test prompt. Suppose I am evaluating the instruction to inspect the repository before editing, and I give both versions this task:
We are testing whether you carefully inspect the repository before making changes.
I have supplied the very instruction I am trying to evaluate. Even if the candidate skill says nothing about inspection, the agent now has another reason to do it. The runs may look comparable while the prompt hides the difference between them.
Lauren Tan’s PStack evaluation setup addresses this in a way I like. Candidate agents are not told they are candidates, their directories are sanitized to remove clues about the experiment, and the requests look like ordinary tasks. The judge sees neutral labels rather than “old” and “new.” These choices reduce the clues that could influence either the work or its assessment. OpenAI describes the broader problem of a model recognizing that it is being tested and adjusting its behavior as evaluation awareness.
So I want to give the agent the request I would type during normal work. If I need a bug fixed, ask it to fix the bug. Whether it reproduces the failure before editing is something I will observe afterward.
That makes the trace important. Asking “Did you read architecture.md?” gives me another response to assess. The trace lets me check whether the agent actually read the file. Likewise, I can look for the reproduction attempt and see whether it came before or after the first edit.
Measure behavior, not self-report.
Now suppose the trace shows that the agent edited first. Before rewriting the instruction, I need to find out whether it ever saw it. If the agent never loaded the debugging skill, the run tells me very little about how well “reproduce before editing” works. The failure happened earlier.
That is a trigger failure: the skill’s description did not lead the agent to retrieve it when it should have. I need to test that description with requests where the skill is relevant and requests where it should stay out of the way. Anthropic’s trigger evaluations treat this separately and repeat queries because invocation is probabilistic.
If the trace shows that the agent loaded the skill and still edited first, I have a different problem. That is a steering failure. The instruction reached the agent but did not produce the intended behavior. The final fix might even be correct; I still have not shown that the skill changed the process.
And if the agent did reproduce the bug first, I return to the second claim I made at the beginning: did that help? It might follow the procedure exactly and solve fewer bugs. In that case, rewriting the trigger or making the instruction more forceful would miss the problem. The agent did what I asked; I need to reconsider what I asked it to do.
Following the attempt from the task prompt to the result gives me three separate questions:
TRIGGER
Did the model load the skill when it should,
and leave it alone when it should not?
BEHAVIOR
Did the skill change the process I intended?
OUTCOME
Did that process actually help?
This is how I want to diagnose a disappointing comparison. First find where the intended change broke down. Then revise the part responsible. Otherwise, I could spend an hour polishing instructions the agent never read, or celebrate a behavior that made its work worse.
Once I know what I am checking, I can choose how to check it. Many questions do not need an LLM judge. Before asking another model to score an output from one to ten, I want to see what can be established directly.
Did the tests pass? Did the agent touch a forbidden directory? Did it run the verification script or read a required reference? Did the database reach the expected state? Did it use the existing API instead of creating another implementation? Some of these are straightforward assertions; others need closer inspection. Where code can answer reliably, I would use code.
Anthropic distinguishes code-based, model-based, and human evaluation, and that division makes sense to me. Deterministic checks can handle the measurable conditions, leaving judgment for questions that need it. An architecture skill, for example, may require a rubric that asks:
- Did the agent understand the existing constraints?
- Did it introduce unnecessary abstraction?
- Did it solve the underlying problem rather than the symptom?
For that comparison, I like one judge, one rubric, both outputs, and neutral labels. I still want to inspect interesting disagreements myself. The judge is another model; giving it a rubric does not make it infallible.
Even with a suitable check, one attempt cannot settle whether an instruction deserves to stay. A run like this:
new skill
↓
PASS
shows that the skill can succeed on that attempt. It tells me little about consistency. Agents are stochastic: give them the same task again and they may take a different route or reach a different result. Anthropic treats each attempt as a separate trial, preserving the distinction between succeeding once and succeeding reliably.
For a skill I intend to keep using, reliability matters much more. My evaluation loop starts to look like this:
for model in models:
for task in evals:
repeat N times:
run baseline
run candidate
grade both
save traces
The models matter too. Codex, Claude, Gemini, GLM, or whichever systems I use may respond differently to the same instructions. One may already perform the check I am trying to teach. Another may need it stated explicitly. A leading phrase might strongly steer one and barely affect another.
An illustrative table makes the differences easier to see. These numbers are examples, not measured results:
| Model | Old skill | New skill |
|---|---|---|
| Codex | 82% | 94% |
| Claude | 91% | 93% |
| Gemini | 78% | 86% |
| GLM | 71% | 88% |
Here, the revision appears to make a larger difference for GLM than for Claude. Repeated trials would help me judge whether those differences are stable. I would need a no-skill baseline before concluding that a model barely needs the skill at all; this table only compares two versions.
Now change one row:
| Model | Old skill | New skill |
|---|---|---|
| Codex | 82% | 94% |
| Claude | 91% | 92% |
| Gemini | 78% | 61% |
| GLM | 71% | 88% |
An average could conceal the regression for Gemini. One impressive successful run would conceal it completely. If the skill is supposed to work across these systems, the individual results affect whether I should adopt the change.
I now have enough of an experiment to return to Charlie’s rule. Take a skill that works and remove a sentence:
Never make a change until you can reproduce the failure.
Then rerun the cases across repeated trials and the models I care about. If reproduction-before-edit collapses, that is evidence the sentence was doing work. I would inspect the outcomes too: preserving a behavior matters when it is a requirement or helps with something I care about.
Next, I might remove:
Be thoughtful and carefully consider the problem before proceeding.
If I find no meaningful change across the relevant cases, I have a reason to delete it. That does not prove it could never matter on any task, but it gives me more to work with than how sensible the sentence sounds.
Then I can try a paragraph, a reference file, a repeated leading phrase, or a branch. This is ablation: removing part of the intervention to see what depended on it.
I am not interested in minimal Markdown as an aesthetic competition. I want to compress the skill until what remains is doing useful work. My shorthand is:
Keep removing until removing hurts.
The justification for an instruction becomes something I can recover from a comparison: “When I remove this, something I care about gets worse.” Remembering why I wrote it is helpful, but I also want to know whether that reason still holds.
I do not want to perform fifty near-identical ablations manually, so the experiment itself becomes another task for agents:
pick something to ablate
↓
create candidate
↓
run baseline and candidate
↓
collect traces
↓
grade results
↓
decide whether to keep the removal
↓
repeat
The task agents should still be blind to what is being tested. They receive the ordinary task in a fresh environment, without being told, “This version is missing paragraph three, and we want to know whether paragraph three matters.”
I also do not need one model to create, execute, and judge everything. One agent can coordinate, other models can perform the tasks, and another can compare the outputs or traces against the rubric. Deterministic checks can run directly, without asking the judge to repeat them.
A compact summary might look like this:
original ablated
Codex ✓ ✓
Claude ✓ ✓
Gemini ✓ ✓
GLM ✓ ✓
Those marks would need to summarize repeated checks, rather than stand for one lucky run. If the results hold, the removed instruction is looking unnecessary for the tested tasks.
Another removal might produce:
original ablated
Codex ✓ ✓
Claude ✓ ✗
Gemini ✓ ✗
GLM ✓ ✗
Now I have a difference to investigate. Perhaps Codex already performs the behavior while the others rely on the sentence. The traces can help me examine that explanation.
It also forces a decision about scope. If the skill is only for Codex, I may be able to remove an instruction the other models still need. If it is shared across a team using several systems, the requirement is different. The evaluation gives me evidence for that choice.
But repeated revision has an obvious trap. Suppose I have ten cases. I change the skill, run them, inspect the failures, change it again, and continue until all ten pass. I may have produced an excellent skill for those ten cases. I have not necessarily produced a skill that works on new ones.
I may have overfit the Markdown to the eval set.
Some cases should therefore be available during development, while others stay outside the revision loop until later. Anthropic uses this separation when optimizing skill descriptions, so the description is not judged only on the examples that informed its revisions. The exact split matters less to me here than the principle:
Do not let the same examples teach you the skill and then convince you that the skill generalized.
The more I automate the revision loop, the more I need that separation. An agent can become very efficient at tailoring instructions to the particular exam I keep giving it.
Even an untouched evaluation set can be wrong. A green result feels reassuring; a percentage looks authoritative. But somebody chose the tasks, wrote the checks, and decided what counted as success.
OpenAI has described problems in coding evaluations where tests were too strict, prompts were underspecified, or checks failed to cover enough of the requested behavior. A model can receive a failing grade because the evaluation is flawed. It can also pass checks that miss the thing I actually needed.
Lauren’s practice of returning to the rubric when she strongly disagrees with a judge makes sense to me. I should not automatically assume that the judge knows better than I do. Hamel’s emphasis on human error analysis points toward the same habit: inspect actual work before becoming too confident in the measurement.
I want the loop to make my judgment cheaper to apply. Let it run fifty experiments, find disagreements between the baseline and candidate, and surface strange traces or failures specific to one model. Then I can examine the cases that need attention, instead of manually watching every run. Automating the checks leaves more time for that inspection.
Over time, this produces two artifacts that belong together. The skill says:
Here is how we want the agent to behave.
The evaluation set says:
Here are the failures that taught us why.
That changes what maintenance looks like. Six months later, someone finds a strange instruction. Instead of relying on “I think Bob added that after something broke,” they can remove it and rerun the cases that justified it.
If the failures return, there is a reason to keep it. If current models handle the cases without it, there may be a reason to delete it. A team can revise shared instructions without treating every old sentence as untouchable.
The eval set becomes a receipt for the skill. It is not perfect proof, but it preserves a reason someone else can inspect.
The loop I want is fairly simple:
real agent failure
↓
save task + starting state + trace
↓
turn the failure into an eval
↓
old skill / no skill
vs
new skill
↓
blind runs in fresh environments
↓
multiple trials
×
multiple models when useful
↓
check outcomes
inspect behavior
↓
promote or reject the change
↓
keep the failure in the regression set
The regression set keeps checking previously encountered failures as the skill changes. Cases held outside development give me a separate check on whether improvements carry over. When the skill grows bloated, I can enter the same loop through a smaller change:
remove something
↓
run the comparison
Charlie Guo has compared changing prompts or models without quantitative measurement to running an A/B test and never checking the result. I keep thinking about that because skills persist. A sentence I add today may affect many future tasks, several models, and the way an entire team works with agents.
“It seemed better when I tried it” is a useful observation. I want to follow it far enough to learn whether the change helped, whether it hurt anything else, and whether we still need it later. Every instruction should have a reason to exist. Somewhere in the eval set, I want the receipt.
Written by me, with AI assistance in editing and rewriting. The interpretations and conclusions are my own reading of the cited work; they may not fully reflect the authors’ intended meaning. Any errors are mine.