Quick answer

Claude Code usage is not determined only by how many prompts you type. Long repository reads, repeated context, tool output, implementation turns, retries, and large pasted logs all create more work for the model. The exact allowance and reset behavior depend on your current plan and product rules, so there is no honest universal “tokens per week” number to promise.

The durable strategy is to keep a strong commander focused on decisions and reviews, send bulk reading and bounded implementation to a permitted worker such as Codex or GLM, and record enough evidence to see whether delegation reduces Claude work or merely adds coordination overhead.

What consumes quota in practice?

Think in terms of context and turns, not just messages. A short question can be expensive when it carries a large repository snapshot. A longer prompt can be efficient when it gives a worker a narrow file set, acceptance criteria, and a clear stop condition.

Repository reads and repeated context

Asking an agent to rediscover the same tree, configuration, and test output in every turn is avoidable overhead. Large command output is also context. Prefer targeted searches, summaries, and explicit file boundaries. Keep a compact working note for decisions so a new turn does not need to reconstruct the entire history.

Tool calls and output

Searches, tests, diffs, logs, and generated patches become inputs to later reasoning. A failing test with 2,000 lines of noisy output is less useful than the first failure plus the relevant source excerpt. Truncate routine output and preserve the decisive lines.

Implementation and retries

Code generation is not the only cost. A vague task produces a draft, a correction, a second correction, and a review. The total work rises when the first request lacks scope or acceptance criteria. A small, testable packet often saves more usage than a clever prompt.

# Prefer focused evidence over an entire repository dump
rg -n "quota|usage|model" README.md docs src
git diff -- path/to/changed-file
npm test -- --runInBand 2>&1 | tail -80

These commands are examples of an evidence-first workflow, not a guarantee about how any provider meters usage. Check the plan-specific usage display and official limits for your account.

Use a commander/worker split

The commander owns intent, decomposition, prioritization, risk decisions, and final acceptance. Workers handle bounded tasks that can be checked independently. This split is useful when the commander’s scarce context is more valuable for judgment than for reading every file or typing every mechanical edit.

Keep the commander’s packet small

Task: add a parser test
Goal: cover quoted commas in CSV fields
Read: src/csv/parse.ts, test/csv.test.ts
Write: test/csv.test.ts only
Accept: test passes; no production-file changes
Verify: npm test -- csv.test.ts

A good packet names the goal, readable files, writable files, acceptance criteria, and verification command. It prevents a worker from wandering into unrelated discovery and gives the commander a compact result to review.

Offload bulk reads and implementation

Codex or GLM can be appropriate for repository inventory, repetitive transformations, fixture creation, focused tests, and a first implementation when the environment and policy allow it. Do not send secrets, private credentials, unrelated personal data, or an entire repository by default. Give the worker only the context needed for its declared task.

Delegation is not automatically cheaper. If a worker needs many clarifications, or if the commander must rewrite its patch, Claude usage may increase. Measure the whole path: commander turns, worker turns, review turns, and rework.

Commander: decide scope, assign task, review diff
Worker: inspect named files, edit declared files, run focused test
Commander: reject out-of-scope changes, run final verification
Result: merge only evidence-backed work

Measure where the tokens and turns go

Start a simple per-task ledger. You do not need a perfect token counter to find waste. Count the visible work that matters: prompt turns, large reads, tool-output volume, retries, review passes, and whether the task reached acceptance.

task,owner,claude_turns,worker_turns,retries,large_reads,accepted
parser-test,commander+glm,2,3,0,1,yes
auth-refactor,commander,7,0,3,4,no
docs-index,commander+codex,1,2,0,0,yes

After a few tasks, look for patterns. High Claude turns with repeated large reads suggests poor context boundaries. High worker turns with many retries suggests the packet or acceptance test is weak. High review time suggests the worker is touching too much scope.

Use a before-and-after experiment

  1. Choose two similar tasks, or compare the same workflow over two periods.
  2. Record Claude turns, worker turns, retries, large reads, and acceptance.
  3. Change one variable, such as assigning fixture work to a worker.
  4. Keep the change only if total effort falls without lowering review quality.

Do not convert a dashboard number into a false precision claim. Provider counters may use plan-specific accounting that is not visible in your local ledger. Your ledger is for workflow decisions; the provider’s usage display is authoritative for your allowance.

Adopt a model tier policy

Choose a model by the consequence of being wrong, the amount of context required, and the cost of rework. Do not use the strongest model for every grep, formatting pass, or repetitive fixture. Do not use a cheaper or external worker for a decision that needs deep repository context, security judgment, or final approval.

Example policy for assigning work
WorkPreferred ownerReason
Goal, architecture, risk, final reviewCommanderHigh judgment and accountability
Repository inventory and bulk readsCodex or GLM workerBounded, evidence-producing work
Mechanical implementation with testsWorker firstProtects commander context
Ambiguous cross-cutting changeCommanderClarify before delegation
Secrets, credentials, private dataNever externalizeKeep sensitive material in its approved environment

Policies must also cover fallback behavior. If the selected worker is unavailable, either do the small task locally or stop and re-scope it; silently switching to an unapproved model can violate your organization’s rules. Keep exact model and endpoint requirements in the project instructions, and verify the selected model before a run.

Make the policy executable

Write the rule where the team will see it, then attach a check to each handoff. For example: “workers may edit only declared files,” “the commander reviews every diff,” and “a failed verification returns to the owner instead of triggering an unbounded retry.” These are practical guardrails because they limit both accidental scope and repeated context.

Use a stop condition as well. If two attempts produce no accepted result, pause and inspect the task packet, test, and model assignment. More prompts are not automatically more progress. A changed hypothesis—smaller scope, better fixture, different owner, or clearer evidence—is the responsible next experiment.

Common questions

Does saving prompts make quota last longer?

Shorter prompts can help, but saved prompts do not remove the context and tool work required by the task. Reuse concise instructions and provide only current evidence.

Should I delegate every task?

No. Delegation has coordination cost. Use it for bounded work with a clear acceptance test, not for ambiguous decisions or sensitive material.

Can I predict exactly when my weekly limit resets?

Use the plan’s current usage interface and official documentation. Do not rely on a remembered reset time or a number copied from another plan.

What is the best way to save quota today?

Stop repeated rediscovery first: narrow file reads, cap noisy output, define writable scope, and review a worker’s diff once. Those changes reduce avoidable turns without pretending to change provider accounting.

Review the ledger at the end of the week, not after every message. The useful question is whether a changed assignment produced an accepted result with fewer expensive commander turns and less rework. If not, change the task boundary before changing the model.