Quick answer

A weekly limit lasts longer when the expensive conversation is reserved for work that benefits from continuity and judgment: choosing the goal, resolving ambiguity, defining acceptance criteria, reviewing risky diffs, and deciding whether evidence is sufficient. Repetitive search, log reduction, mechanical edits, fixture generation, and first-pass test execution belong in a bounded worker lane.

Native Claude Code subagents protect the main context, but they still consume Claude usage. External CLI workers may move work to another provider or local model, but they add orchestration, security, and verification costs. The right metric is not “agents spawned.” It is commander usage per accepted change, including failed handoffs and review rework.

Give the commander judgment, not chores

The commander owns the state that must remain coherent across the whole task. It understands the user request, repository constraints, architectural tradeoffs, and current evidence. Workers receive only one independently testable slice. This is a responsibility boundary, not a model hierarchy: a capable worker can write most of a patch, while the commander remains accountable for accepting it.

A practical commander/worker split
Commander keeps Worker receives
Goal, non-goals, risk, and final decision One explicit deliverable with a stop condition
Architecture and cross-module tradeoffs Local implementation inside declared files
Conflict resolution and integration Focused tests, logs, or a patch artifact
Security-sensitive or destructive approval Read-only analysis with redacted inputs

A useful test is: can you tell whether the worker succeeded by inspecting files and running one command? If not, the packet is probably too broad. “Improve the backend” forces the commander to reconstruct intent later. “Add validation to these two routes; these four tests must pass; do not change the database” creates reviewable evidence.

What not to spend commander tokens on

Do not use the main conversation as a transcript sink. Large test logs, complete directory listings, generated files, dependency install output, and dozens of nearly identical search hits dilute the state that makes the commander useful. Reduce them before they enter the main context.

# Bad handoff: paste 18,000 lines into the commander.
npm test > full-test.log 2>&1

# Better: preserve the source, return a focused digest and its location.
npm test > full-test.log 2>&1
rc=$?
{
  printf 'exit=%s\n' "$rc"
  rg -n "FAIL|Error:|AssertionError" full-test.log | head -80
  printf 'full_log=%s\n' "$PWD/full-test.log"
} > test-summary.txt
exit "$rc"

Also avoid commander turns for formatting-only changes, bulk renames with a deterministic mapping, boilerplate fixtures, initial repository maps, and repeated reruns of an unchanged failure. A worker should return the smallest decision-changing evidence. The commander should not ask “did it work?” when a test exit code, diff, or generated report can answer directly.

Native subagents versus external CLI workers

Claude Code subagents start in isolated contexts and return a summary to the caller. That makes them excellent for noisy repository exploration, focused review, or a task needing restricted tools. They reduce pollution of the main context; they do not create free Claude usage. A verbose subagent result can even spend more because the worker reads context and the commander then reads its report.

External CLI workers are separate processes: another coding agent, a local model, or a deterministic program. They are most useful when they have a separate allocation or are cheap at the delegated task. They should run in a dedicated worktree or own disjoint files, receive no secrets, and leave durable output. Their independence is also their risk: they may not inherit the same rules, permissions, or understanding as Claude Code.

# Generic external-worker envelope. Replace WORKER_COMMAND explicitly.
run_id="tests-$(date -u +%Y%m%dT%H%M%SZ)"
run_dir=".agent-runs/$run_id"
mkdir -p "$run_dir"

set +e
WORKER_COMMAND \
  --prompt "$run_dir/task.md" \
  >"$run_dir/worker.log" 2>&1
rc=$?
set -e

printf '%s\n' "$rc" >"$run_dir/worker.exit.tmp"
mv "$run_dir/worker.exit.tmp" "$run_dir/worker.exit"
git diff --stat >"$run_dir/diff-stat.txt"
exit "$rc"

Prefer a native subagent when context isolation is the main benefit and Claude-specific tools or project memory matter. Prefer an external worker when the task is provider-neutral, the external allocation is intentional, and its output can be verified without trusting its narrative. Work directly when the task is a tiny edit or every phase shares the same context; delegation overhead can exceed the saved usage.

Use a task packet that prevents expensive rework

A cheap worker with a vague prompt is expensive. It explores the wrong files, produces an oversized patch, and consumes a commander review cycle. Put the contract in a file so the exact packet and result can be audited.

# Task: validate webhook timestamps

Goal:
- Reject requests older than 300 seconds.

Read:
- src/webhooks/verify.ts
- test/webhooks/verify.test.ts

Write:
- src/webhooks/verify.ts
- test/webhooks/verify.test.ts

Do not:
- Change dependencies, public response bodies, or unrelated formatting.

Acceptance:
- npm test -- test/webhooks/verify.test.ts passes.
- A timestamp at exactly 300 seconds is accepted.
- A timestamp at 301 seconds is rejected.

Return:
- Changed files, test exit code, and unresolved risks.

Stop:
- Stop after one changed retry if the focused test still fails.

The boundary case in the acceptance criteria is deliberate. It prevents the worker from choosing semantics silently. The write list limits collisions. The stop rule prevents an autonomous loop from burning both worker capacity and commander attention.

Measure where usage actually goes

Claude Code's /usage command shows plan usage for subscribers and token details for API users. Record a baseline before a substantial task and another value after acceptance. Do not convert subscription bars into invented token counts; store the units the interface provides.

date,task,commander_before,commander_after,worker,attempts,accepted,rework_min
2026-09-01,webhook-age,42%,47%,external-cli,1,yes,6
2026-09-02,repo-map,47%,49%,native-explore,1,yes,2
2026-09-03,rename-api,49%,55%,none,0,yes,0

Add outcome columns: accepted change, failed handoff, minutes spent reviewing, and whether the task was rolled back. After a week, compare commander delta per accepted result. If native exploration routinely returns pages the commander re-reads, require a tighter report. If external workers need extensive repair, move that task family back to Claude or strengthen the acceptance command. A lower usage number paired with bad diffs is not efficiency.

Attribute usage to task phases, not personalities

A single before-and-after number tells you whether the whole task was expensive, but not why. Split substantial work into framing, exploration, implementation, review, and recovery. Record a usage observation at each boundary when the interface makes that practical. The goal is not precise accounting down to a token. It is to identify the phase that repeatedly expands and to change the routing rule for the next comparable task.

For example, an implementation may be cheap while review costs dominate because the worker returns a broad, undocumented diff. Moving more implementation away from the commander would make the wrong phase cheaper. The improvement is a narrower write set, a diff summary tied to acceptance criteria, and a focused verification command. Conversely, if framing consumes several turns because the task is ambiguous, delegation should wait; every worker would otherwise receive a different guess.

Count discarded runs. A worker attempt that produces no acceptable artifact still consumed task-writing, launch, monitoring, and review attention. Record it as a failed handoff, not as zero commander cost. Also count context reloads: starting a fresh worker for a tightly coupled second phase may require it to reread the same architecture and files. Keeping the work in the commander's current context can be cheaper even when the commander's model is nominally more expensive.

After several similar tasks, turn the observations into a routing table. “Use a read-only native subagent for repository maps over ten files” is actionable. “Delegate more” is not. Revisit the table when model behavior, plan limits, repository size, or worker tooling changes. These strategies are operating hypotheses, not permanent facts about every Claude Code plan.

Run a weekly allocation loop

  1. Check /usage before committing to a large task.
  2. Reserve commander capacity for integration, review, and urgent defects late in the period.
  3. Delegate only independent tasks with objective acceptance evidence.
  4. Cap every worker at one initial attempt and one materially changed retry.
  5. Review the real diff and run the focused command locally before acceptance.
  6. At week's end, remove any delegation pattern whose rework exceeds its saved usage.

When Claude Code reports that the weekly limit is reached, it blocks further requests until the displayed reset time unless eligible extra usage is enabled. No shell wrapper can bypass an account limit. Good orchestration changes which tasks consume the allocation before that point; it does not alter the plan's quota or guarantee a particular number of hours.

Primary sources