Quick answer

Give each background AI worker its own directory containing worker.log, worker.pid, and worker.exit. Write the PID immediately after launch. Write the exit code to a temporary file and rename it after the agent command returns. The watcher reads the exit file before it considers the PID or log age.

Report DONE only for an exit value of zero, FAILED for a nonzero exit or a vanished process with no exit evidence, RUNNING for a live process with recent output, and STALL for a live process whose log has been idle longer than a task-specific threshold. STALL is a prompt to inspect; it is not proof that a model call is dead.

Build a protocol around terminal evidence

Background jobs fail in awkward places: before the agent starts, during a tool call, after the agent finishes but before the wrapper records status, or when the parent shell disappears. No single file covers all states. The three-file protocol makes each signal narrow and auditable.

What each worker file proves
File Meaning Does not prove
worker.log Observable output and its last change time Success or current process identity
worker.pid The wrapper PID captured at launch That the same process is still alive later
worker.exit The wrapper reached terminal status That the produced code is correct
.agent-runs/
├── api-review/
│   ├── task.md
│   ├── worker.log
│   ├── worker.pid
│   └── worker.exit
├── test-fix/
│   ├── task.md
│   ├── worker.log
│   ├── worker.pid
│   └── worker.exit
└── docs-pass/
    ├── task.md
    ├── worker.log
    ├── worker.pid
    └── worker.exit

Store the wrapper PID, not an arbitrary descendant discovered later with pgrep. The wrapper owns exit capture and should forward termination to its child. A production-grade supervisor may use process groups, cgroups, or a service manager; the file protocol remains useful as the human-readable run receipt.

Use one marker namespace per worker

A shared DONE file creates a race. The first completed worker makes the whole pool appear finished. A shared log is almost as bad: interleaved output makes the last line belong to whichever process wrote last, not the worker you are diagnosing.

# Broken: three workers can overwrite the same terminal marker.
worker api   && touch .agent-runs/DONE &
worker tests && touch .agent-runs/DONE &
worker docs  && touch .agent-runs/DONE &

# Correct shape: terminal evidence lives under the worker ID.
for id in api tests docs; do
  mkdir -p ".agent-runs/$id"
  ./run-agent.sh "$id" &
  printf '%s\n' "$!" >".agent-runs/$id/worker.pid"
done

Worker IDs must also be safe path components. Reject empty IDs, slashes, and traversal sequences instead of letting a prompt or task title choose a filesystem path.

id=${1-}
if [[ ! "$id" =~ ^[a-z0-9][a-z0-9._-]{0,63}$ ]]; then
  printf 'invalid worker id: %q\n' "$id" >&2
  exit 64
fi

Launch the agent and capture exit atomically

Save the following as run-agent.sh. Replace AGENT_COMMAND with an explicitly configured non-interactive CLI. Do not use eval to assemble a command from task text.

#!/usr/bin/env bash
set -u

id=${1-}
root=${RUN_ROOT:-"$PWD/.agent-runs"}
dir="$root/$id"
log="$dir/worker.log"
exit_file="$dir/worker.exit"

mkdir -p "$dir"
rm -f "$exit_file" "$exit_file.tmp"

{
  printf '__AGENT_START__=%s\n' "$(date -u +%FT%TZ)"
  set +e
  AGENT_COMMAND --prompt-file "$dir/task.md"
  rc=$?
  set -e
  printf '__AGENT_EXIT__=%s\n' "$rc"
  printf '%s\n' "$rc" >"$exit_file.tmp"
  mv "$exit_file.tmp" "$exit_file"
  exit "$rc"
} >"$log" 2>&1

The rename prevents the watcher from reading an empty or partially written exit file. It is atomic when both paths are on the same filesystem. Remove stale exit files before a new run, or better, use a new immutable run ID rather than reusing the directory. Capture $! immediately after backgrounding the wrapper.

id="api-review-$(date -u +%Y%m%dT%H%M%SZ)"
mkdir -p ".agent-runs/$id"
cp prompts/api-review.md ".agent-runs/$id/task.md"

nohup ./run-agent.sh "$id" >/dev/null 2>&1 &
pid=$!
printf '%s\n' "$pid" >".agent-runs/$id/worker.pid"
printf 'started id=%s pid=%s\n' "$id" "$pid"

Walk through a portable watcher

#!/usr/bin/env bash
set -u

root=${RUN_ROOT:-"$PWD/.agent-runs"}
stall_after=${STALL_AFTER_SECONDS:-900}
now=$(date +%s)

mtime() {
  if stat -f %m "$1" >/dev/null 2>&1; then
    stat -f %m "$1"          # macOS and BSD
  else
    stat -c %Y "$1"          # GNU/Linux
  fi
}

for dir in "$root"/*; do
  [[ -d "$dir" ]] || continue
  id=${dir##*/}
  pid_file="$dir/worker.pid"
  log_file="$dir/worker.log"
  exit_file="$dir/worker.exit"

  if [[ -s "$exit_file" ]]; then
    read -r rc <"$exit_file"
    if [[ "$rc" =~ ^[0-9]+$ ]] && (( rc == 0 )); then
      printf '%-28s DONE    exit=0\n' "$id"
    elif [[ "$rc" =~ ^[0-9]+$ ]]; then
      printf '%-28s FAILED  exit=%s\n' "$id" "$rc"
    else
      printf '%-28s FAILED  malformed exit file\n' "$id"
    fi
    continue
  fi

  if [[ ! -s "$pid_file" ]]; then
    printf '%-28s FAILED  missing pid and exit\n' "$id"
    continue
  fi

  read -r pid <"$pid_file"
  if [[ ! "$pid" =~ ^[0-9]+$ ]] || ! kill -0 "$pid" 2>/dev/null; then
    printf '%-28s FAILED  wrapper vanished\n' "$id"
    continue
  fi

  if [[ ! -e "$log_file" ]]; then
    printf '%-28s RUNNING live, no log yet\n' "$id"
    continue
  fi

  changed=$(mtime "$log_file")
  idle=$((now - changed))
  if (( idle >= stall_after )); then
    printf '%-28s STALL   live, idle=%ss\n' "$id" "$idle"
  else
    printf '%-28s RUNNING live, idle=%ss\n' "$id" "$idle"
  fi
done

The order is the important part. An exit file wins even if the PID was recycled. A missing process without terminal evidence is FAILED, because the supervisor cannot prove a clean result. Only a live process reaches the inactivity calculation. The watcher handles both BSD and GNU stat syntaxes and rejects a malformed exit file rather than treating arbitrary text as zero.

Choose the threshold from observed task latency. A repository scan might be suspicious after five quiet minutes; a large test suite or remote model request may legitimately remain silent longer. Add a hard runtime ceiling separately. Inactivity and total elapsed time answer different questions.

Separate activity, progress, and completion

A changing log proves activity, not progress. An agent can emit retries, animated status lines, or the same recoverable error indefinitely while the modification time stays fresh. If the CLI offers structured events, record a second timestamp for the last meaningful event: tool completion, file change, test result, or phase transition. Keep the raw log timestamp too; the difference reveals a process that is noisy but stationary.

Do not manufacture activity by touching the log from the watcher. That makes the monitor reset its own inactivity clock. If a wrapper needs a heartbeat during a silent remote request, write it to a separate worker.heartbeat file. The wrapper or supervised child—not the watcher—must own that heartbeat, and its interval and meaning should be documented. A heartbeat proves the wrapper loop is scheduled; it still does not prove the remote operation is making progress.

Completion remains a distinct state. Neither a fresh heartbeat nor a newly appended log line can replace worker.exit. If the worker reports a semantic final message but the process does not exit, the watcher should continue to show RUNNING or STALL. Fix the CLI invocation or shutdown path instead of teaching the watcher to scrape prose for a success phrase.

Add a separate maximum runtime

Some failures never become inactive: a fast retry loop can append errors forever. Record the start epoch in a metadata file or derive it from an immutable run directory, then compare total elapsed time with a hard ceiling. Call the result TIMEOUT or include it as a reason under STALL; do not confuse it with a nonzero agent exit that was actually captured. The operator can then decide whether to signal the process based on task policy.

Thresholds should be per task family. A five-minute ceiling may suit a formatter worker and be destructive for an integration test suite. Start from observed successful durations, add a deliberate margin, and review outliers. Automatic termination is safest only when the wrapper forwards signals, preserves partial output, and records that the supervisor—not the agent— caused the final interruption.

Avoid the log-tail echo trap

A common monitoring loop prints a label, then tails a log. When the log is empty or missing, the label becomes the last visible line. Another script greps that combined output and mistakes its own word “DONE” or “FAILED” for worker evidence.

# Broken: monitor text and agent evidence share one stream.
echo "DONE? checking $id"
tail -n 20 ".agent-runs/$id/worker.log" 2>/dev/null

# Also broken: a prompt or model response can contain the marker.
if tail -n 20 "$log" | grep -q 'DONE'; then
  echo "$id DONE"
fi

# Correct: classify from the exit file, then show a clearly delimited tail.
printf '%s\n' "--- $id diagnostic tail (not status) ---" >&2
tail -n 20 "$log" >&2 || true

Logs are untrusted text. They may quote the task packet, print source code containing “DONE,” repeat a previous error, or end before buffered output is flushed. Use log modification time as an activity hint and the atomic exit file as terminal state. If you emit structured events, keep them on a separate channel or validate their schema and worker identity.

Respond to a stall without erasing the cause

  1. Record the current watcher output, log tail, elapsed time, and process command.
  2. Inspect whether a child process is waiting for input, network, a lock, or a permission prompt.
  3. Review partial repository changes before sending a signal.
  4. Try a normal termination first and wait for the wrapper to record an exit.
  5. For a retry, create a new run directory and change one material condition.
dir=.agent-runs/api-review-20260901T090000Z
pid=$(cat "$dir/worker.pid")

ps -o pid=,ppid=,lstart=,etime=,state=,command= -p "$pid"
tail -n 100 "$dir/worker.log"
git status --short

# Confirm the command is still the intended wrapper before signaling.
kill -TERM "$pid"

PID reuse is real. Before sending a signal, compare the command and start time with the run. If the wrapper is gone and the PID belongs to another process, do not kill it. Classify the agent run as FAILED because its exit evidence was lost, preserve the directory, and fix the wrapper or supervisor before retrying.

Finally, DONE is process evidence, not acceptance evidence. A zero exit means the CLI and wrapper completed. The coordinator must still inspect the diff, run the declared acceptance command, and decide whether the work satisfies the task. Keep those review results next to the run rather than overwriting the original log.

Primary sources