Quick answer

Write a CLAUDE.md rule as an executable decision: “When X is true, inspect Y; if Z is missing or fails, stop; report W.” This trigger + check + stop + evidence structure gives the agent a moment to apply the rule and gives you a way to tell whether it did. “Be careful,” “follow best practices,” and “always verify everything” provide neither.

Keep universal rules short, put path-specific requirements near the code they govern, and test important instructions in a new Claude Code process. A rule that the current conversation can repeat from recent context has not yet proved that a future session loads or follows it.

The anatomy of a rule agents can execute

Four clauses that turn prose into a control
Clause Question it answers Concrete form
Trigger When does this apply? “Before editing a database migration…”
Check What must the agent inspect or run? “Read the migration wrapper and run…”
Stop Which result blocks continuation? “If rollback is undefined, do not edit.”
Evidence What proves compliance? “Report the command and exit code.”

Not every preference needs all four clauses. “Use two spaces in YAML” is easy to verify with a formatter. Add the full anatomy where mistakes are costly, invisible, or tempting to explain away: deployments, migrations, generated files, credentials, background processes, and completion claims.

Make the trigger observable

Triggers should name an event the agent can detect from the request, planned command, or changed paths. “When appropriate” delegates the entire policy decision back to the model. “Before changing files under infra/” can be checked against a write plan and later against the diff. Include user vocabulary when it matters: deploy, publish, release, and push live may describe the same boundary.

Prefer one decisive check

A rule with twelve mandatory commands is likely to be applied selectively, especially when half do not relate to the current change. Lead with the smallest check that can block an unsafe action or falsify a completion claim. Put broader test matrices in a script so the agent executes one stable interface. The script can evolve without forcing every instruction file to duplicate its internals.

Stop at the boundary the agent cannot resolve

“Ask for help if needed” is not a stop condition. Name the missing fact: target environment, rollback behavior, credential presence, owner approval, or acceptance result. Also say what remains allowed. A deployment blocked on credentials may still permit a local build and artifact checksum, but it must not be reported as a release. Precise stops prevent both reckless continuation and unnecessary abandonment of safe work.

Require compact, falsifiable evidence

Evidence should let a reviewer independently check the claim. Command, exit code, relevant artifact, and skipped case are usually enough. Avoid demanding full logs in chat; preserve them in a file and report the path plus the first actionable failure. “Tests passed” is a conclusion. “Command X exited zero and covered cases A through C” is an inspectable claim.

## Database migrations

- Trigger: Before creating or editing a production migration.
- Check: Read `docs/database-migrations.md` and identify the
  transaction boundary, lock behavior, and rollback command.
- Stop: If rollback or online-change behavior is undefined, do not
  write the migration; report the missing decision.
- Evidence: In the final response, name the migration check command,
  its exit code, and the rollback path.

Rules agents tend to ignore

Vague aspirations lose against the immediate task. “Write production-quality code” conflicts with no concrete action, so the agent can believe it complied after any plausible patch. Absolute rules with undefined scope are also weak: “always run all tests” is unreasonable for a documentation typo in a large monorepo, encouraging selective interpretation.

Long inventories of obvious advice compete with the few rules that matter. Contradictions are worse: one section says never modify generated files, another says update all affected files, and neither names precedence. Finally, commands that cannot run in a clean checkout train the agent to skip or hand-wave them. A trustworthy CLAUDE.md reflects the repository that exists, not the workflow you wish existed.

# Weak
- Be secure and test thoroughly.
- Never make destructive changes.
- Follow our architecture.

# Stronger
- When changing `src/auth/**`, run `npm test -- auth`.
- If the test command cannot start, stop after capturing the first
  actionable error; do not claim the auth change is verified.
- Report the exact command, exit code, and any skipped acceptance case.

Five worked CLAUDE.md examples

1. Keep generated files synchronized

## Generated API clients

- Trigger: When `api/openapi.yaml` changes.
- Check: Run `npm run generate:client`, then `git diff --exit-code
  -- src/generated` after committing or staging the expected output.
- Stop: If generation needs an unavailable tool or changes files
  outside `src/generated`, stop and report the mismatch.
- Evidence: List the generator version and changed generated files.

This makes “keep generated files current” testable and catches tool-version drift instead of blessing an unexplained diff.

2. Protect production deployments

## Deployment boundary

- Trigger: Any request containing deploy, publish, release, or push live.
- Check: Confirm the named environment and run `./scripts/preflight.sh`.
- Stop: Do not deploy when the environment is ambiguous, preflight is
  nonzero, credentials are missing, or the user requested a preview only.
- Evidence: Report environment, artifact identifier, preflight exit,
  and the final platform URL. Never print credentials.

The trigger includes common synonyms. The stop clause prevents a local build from being reported as a production release.

3. Preserve the first failing test

## Test failures

- Trigger: When a required verification command exits nonzero.
- Check: Preserve the command, exit code, and first actionable failure.
- Stop: Do not declare completion. Do not rerun unchanged more than once.
- Evidence: Classify the failure as caused by this change, pre-existing,
  or environment-required, and cite the file or log containing proof.

“Fix all tests” often causes blind loops. This version keeps evidence and forces a changed hypothesis before another run.

4. Bound worker delegation

## Subagent use

- Trigger: Before delegating work that can modify files.
- Check: Declare the worker goal, readable files, writable files, and
  one acceptance command. Confirm write sets do not overlap.
- Stop: Do not launch parallel writers with overlapping or unknown scope.
- Evidence: Review `git diff --name-only` and the acceptance exit code
  before accepting the worker's result.

This rule targets the decision before delegation, where it can prevent a conflict. A reminder to “review agent work” after the fact is too late to prevent overlapping edits.

5. Handle secrets by presence only

## Secrets

- Trigger: When a command needs an API key, token, cookie, or password.
- Check: Test only whether the named variable is set.
- Stop: If missing, stop at the credential boundary. Never print the
  value, run `env`, or paste a private configuration file into chat.
- Evidence: Report only `VARIABLE_NAME=SET` or `VARIABLE_NAME=missing`.

# Safe presence check
printf 'DEPLOY_TOKEN=%s\n' "${DEPLOY_TOKEN:+SET}"

The permitted output is explicit. That closes the common loophole where an agent “checks” a credential by echoing it into logs.

Prove the rule in a fresh session

First verify discovery with Claude Code's /memory command. It shows the instruction files loaded for the current session. Then exit. Start a new process from the directory a teammate would really use, because CLAUDE.md discovery depends on location and scope.

cd /absolute/path/to/repository
claude

# Inside the new session:
/memory

# Then give a read-only behavioral probe:
You are asked to edit api/openapi.yaml. Before changing anything,
state the required check, the condition that makes you stop, and the
evidence you must report. Do not edit files or run the generator.

Score behavior, not recitation. A passing answer identifies the generator, the outside-write stop condition, and the required evidence without inventing a run. Next, use a disposable branch or fixture to exercise one real trigger. For the secret rule, ask for a deployment with the variable absent and confirm the session reports only “missing.” For the test rule, point it at a known failing fixture and confirm it does not claim success.

# Minimal probe ledger
date,session_dir,rule,trigger_seen,stop_obeyed,evidence_correct,result
2026-09-01,/repo,generated-client,yes,yes,yes,PASS
2026-09-01,/repo/packages/api,test-failure,yes,no,yes,FAIL

If the probe fails, repair the layer that caused it. A discovery failure needs a path or import fix. A misunderstood rule needs clearer wording. A rule displaced by a conflict needs explicit precedence. Do not add three more reminders around an ambiguous original; replace it with one sharper control and repeat the fresh-process test.

Use negative probes, not only happy-path questions

A session can summarize a rule correctly and still violate it under pressure to complete a task. Include a probe where the shortest path conflicts with the rule. Offer a production deployment request without naming an environment, a migration without rollback documentation, or a test failure that looks unrelated. A passing session stops at the defined boundary, preserves the permitted evidence, and does not invent missing approval. Run probes only in a safe fixture or read-only mode; instruction testing should not create the risk the instruction is meant to prevent.

Repeat a critical probe after moving the file, changing an import, upgrading Claude Code, or restructuring a monorepo. Those changes can alter discovery or scope even though the rule text is untouched. Keep the probe prompt and expected behaviors in version control. This turns a subtle agent regression into a reviewable project change rather than a future surprise.

Keep CLAUDE.md small enough to matter

  1. Keep only instructions that every relevant session needs.
  2. Move path-specific rules to scoped rule files where supported.
  3. Delete duplicated advice and name precedence for intentional conflicts.
  4. Run every documented command from a clean checkout after toolchain changes.
  5. Attach a fresh-session probe whenever a rule prevents a costly failure.

Claude Code's documentation explicitly says compliance is not guaranteed, especially for vague or conflicting instructions. Treat CLAUDE.md as one control in a system. Formatters, tests, hooks, permissions, and deployment policy should enforce what can be enforced deterministically. The file should tell the agent when those controls apply and how to surface their result.

Primary sources