Quick answer
Do not treat a background PID as proof that a Codex task is healthy. Start each codex exec invocation through a wrapper that writes worker.log, worker.pid, and worker.exit. Write the exit code through a temporary file and rename it only after Codex returns, so readers never observe a half-written result.
Classify a worker as DONE when its exit file contains zero, FAILED when the exit code is nonzero or the PID vanished before an exit file appeared, and STALL when the process is alive but its log has not changed for a chosen threshold. STALL is a warning for inspection, not proof that the model is dead.
Choose explicit non-interactive flags
codex exec runs without the terminal UI. The current CLI streams progress to standard error and prints the final agent message to standard output. For a worker that may edit its repository, make the working root and sandbox explicit instead of inheriting whichever directory and configuration the parent shell happens to have.
codex exec \
--cd "/absolute/path/to/repository" \
--sandbox workspace-write \
--color never \
--output-last-message "/absolute/path/to/run/final.md" \
"Run the focused task, verify it, and report changed files."
Use --sandbox read-only for analysis workers and --sandbox workspace-write only for workers that must edit. Avoid danger-full-access unless an external container or disposable runner supplies the real security boundary. --ephemeral is useful when saved session rollout files are unnecessary.
If another program needs event-level output, add --json and keep stdout as an unmodified JSONL file. Do not merge stderr into that file; one diagnostic line would make it invalid JSONL. The simpler launcher below deliberately combines both streams into a human log, while -o preserves the final response separately.
Launch workers with durable evidence
Save this as run-worker.sh. It validates the repository and prompt before starting, records unambiguous start and exit markers, and writes the exit status atomically. Each worker gets a separate directory, so concurrent processes never append to the same files.
#!/usr/bin/env bash
set -u
id=$1
repo=$2
prompt_file=$3
run_root=${RUN_ROOT:-"$PWD/.codex-runs"}
run_dir="$run_root/$id"
log_file="$run_dir/worker.log"
exit_file="$run_dir/worker.exit"
mkdir -p "$run_dir"
if ! git -C "$repo" rev-parse --is-inside-work-tree >/dev/null 2>&1; then
printf 'Not a Git worktree: %s\n' "$repo" >"$log_file"
printf '64\n' >"$exit_file"
exit 64
fi
if [[ ! -s "$prompt_file" ]]; then
printf 'Prompt is missing or empty: %s\n' "$prompt_file" >"$log_file"
printf '66\n' >"$exit_file"
exit 66
fi
{
printf '__CODEX_START__=%s\n' "$(date -u +%FT%TZ)"
set +e
codex exec \
--cd "$repo" \
--sandbox workspace-write \
--color never \
--output-last-message "$run_dir/final.md" \
- <"$prompt_file"
rc=$?
set -e
printf '__CODEX_EXIT__=%s\n' "$rc"
printf '%s\n' "$rc" >"$exit_file.tmp"
mv "$exit_file.tmp" "$exit_file"
exit "$rc"
} >"$log_file" 2>&1
Launch it with nohup when the workers must survive a closed terminal. Capture $! immediately; after another background command starts, it refers to the newer process.
mkdir -p prompts .codex-runs/api .codex-runs/tests .codex-runs/docs
nohup ./run-worker.sh api /srv/app prompts/api.txt >/dev/null 2>&1 &
printf '%s\n' "$!" > .codex-runs/api/worker.pid
nohup ./run-worker.sh tests /srv/app prompts/tests.txt >/dev/null 2>&1 &
printf '%s\n' "$!" > .codex-runs/tests/worker.pid
nohup ./run-worker.sh docs /srv/app prompts/docs.txt >/dev/null 2>&1 &
printf '%s\n' "$!" > .codex-runs/docs/worker.pid
Parallel writers should have disjoint file ownership. If the API and test workers can both edit the same module, give one worker read-only review responsibility or place them in separate Git worktrees. Concurrency does not resolve merge conflicts; it only makes them arrive sooner.
Distinguish DONE, FAILED, and STALL
The watcher below is portable across macOS and GNU/Linux. It checks terminal evidence first, then process liveness, then log inactivity. Set the threshold longer than normal silent model or tool calls; fifteen minutes is a practical starting point for small repository tasks.
#!/usr/bin/env bash
set -u
run_root=${RUN_ROOT:-"$PWD/.codex-runs"}
stall_after=${STALL_AFTER_SECONDS:-900}
now=$(date +%s)
mtime() {
if stat -f %m "$1" >/dev/null 2>&1; then
stat -f %m "$1" # macOS/BSD
else
stat -c %Y "$1" # GNU/Linux
fi
}
for dir in "$run_root"/*; do
[[ -d "$dir" ]] || continue
id=${dir##*/}
pid_file="$dir/worker.pid"
log_file="$dir/worker.log"
exit_file="$dir/worker.exit"
if [[ -s "$exit_file" ]]; then
read -r rc <"$exit_file"
if [[ "$rc" == 0 ]]; then
printf '%-16s DONE exit=0\n' "$id"
else
printf '%-16s FAILED exit=%s\n' "$id" "$rc"
fi
continue
fi
if [[ ! -s "$pid_file" ]]; then
printf '%-16s FAILED missing pid and exit evidence\n' "$id"
continue
fi
read -r pid <"$pid_file"
if ! kill -0 "$pid" 2>/dev/null; then
printf '%-16s FAILED process vanished before exit capture\n' "$id"
continue
fi
changed=$(mtime "$log_file" 2>/dev/null || printf '%s' "$now")
idle=$((now - changed))
if (( idle >= stall_after )); then
printf '%-16s STALL alive, log idle=%ss\n' "$id" "$idle"
else
printf '%-16s RUNNING alive, log idle=%ss\n' "$id" "$idle"
fi
done
Run watch -n 10 ./watch-workers.sh on systems with watch, or call the script from a short shell loop. Before killing a STALL, inspect its last log lines and process state:
tail -n 80 .codex-runs/api/worker.log
pid=$(cat .codex-runs/api/worker.pid)
ps -o pid=,ppid=,etime=,state=,command= -p "$pid"
A long tool call can legitimately produce no output. Escalate only when the inactivity exceeds the task's normal latency and the log shows no pending operation worth preserving.
Common traps that look like worker failures
The repository or trust check fails
Codex expects to run inside a Git repository unless --skip-git-repo-check is used. Prefer fixing --cd and the checkout path; skipping the check can hide a worker launched in the wrong directory. Project-scoped .codex/config.toml is loaded only for a trusted project. Review and trust the intended repository interactively before unattended runs that depend on that local configuration.
The sandbox has no network
Workspace write access does not imply outbound network access. A worker may edit files successfully and then appear stuck while a package download, API request, or remote Git operation is blocked. Prefer installing dependencies before the worker. If network is genuinely required, scope it explicitly in Codex configuration instead of switching off the sandbox:
codex exec \
--sandbox workspace-write \
-c sandbox_workspace_write.network_access=true \
--cd "$repo" \
- <"$prompt_file"
The success marker lies
This pattern emits a marker only after success, so a missing marker cannot distinguish failure, termination, or lost output:
# Fragile: echo never runs when codex exits nonzero
codex exec "task" && echo DONE
# Correct: capture immediately, then emit the actual status
set +e
codex exec "task"
rc=$?
printf '__CODEX_EXIT__=%s\n' "$rc"
exit "$rc"
A marker in a log is helpful for humans, but the separate atomic exit file is the machine-readable source of truth.
Operate the worker pool safely
- Give each worker a bounded goal, allowed files, and a verification command.
- Use unique run IDs and never reuse a directory until its evidence is archived.
- Cap concurrency according to API limits, CPU, memory, and repository overlap.
- Inspect the diff and tests locally; an exit code of zero proves the process completed, not that the change is correct.
- On interruption, signal the recorded PID, wait, and only then force termination if the process ignores a normal stop.
pid=$(cat .codex-runs/api/worker.pid)
kill -TERM "$pid"
# After a reasonable grace period, confirm it is still the intended process.
ps -p "$pid" -o pid=,command=
Finally, treat STALL as an operational decision point. The useful question is not “is the PID alive?” but “has this worker produced new evidence, and will waiting change the outcome?” That framing prevents both premature kills and background jobs that consume time indefinitely.
Review a stall before retrying
A retry is useful only when it changes the condition that made the first attempt stop producing evidence. Before restarting, preserve the run directory and write down four facts: the last meaningful log line, the process elapsed time, the files already changed, and the external operation that may still be pending. Those facts distinguish a slow worker from a repeatable defect. Deleting the log and launching the same prompt removes the best evidence while spending another full run on the same uncertainty.
Confirm that the PID still names the expected wrapper. Process identifiers can be reused after a process exits, especially on a busy CI host. The atomic exit file normally closes that race, but a killed wrapper may leave only the old PID. Compare the command, parent, and start time with the recorded run before sending a signal. If they do not match, classify the run as FAILED because its terminal evidence was lost; do not signal the unrelated process that inherited the number.
pid=$(cat .codex-runs/api/worker.pid)
ps -o pid=,ppid=,lstart=,etime=,command= -p "$pid"
git -C /srv/app status --short
tail -n 120 .codex-runs/api/worker.log
Next, classify the likely wait. Network-denied package installs need a dependency or sandbox change. Two workers editing the same lockfile need serialization or separate worktrees. A test process waiting for input needs a non-interactive flag. A model call that is merely slow may need more time, not a different prompt. Change exactly one material variable for the retry and record it in the new run directory; otherwise the second attempt cannot teach you why the first one failed.
Do not resume directly into a repository containing unreviewed partial edits unless continuity is intentional. First inspect git diff and generated files. Either accept that state as the new baseline and tell the retry what already exists, or restore it through a reviewed, recoverable process. A fresh Git worktree is often safer because it preserves the stalled worker's evidence while giving the retry a known commit and an independent write surface.
Finally, cap retry count. One changed retry is usually enough to test a diagnosis. If it stalls at the same operation, stop the pool and fix the shared cause—authentication, network policy, repository lock, prompt ambiguity, or an interactive command. More parallel workers amplify a shared blocker; they do not route around it. The operational goal is a trustworthy final state, not the largest number of processes that can remain alive.