Quick answer

Verification is a chain of evidence: the intended command ran, its real exit status survived logging and cleanup, the CI system did not downgrade the failure, and the final patch satisfies the task. Break any link and a failed test can appear green. This is common when an agent runs test | tail, prints reassuring output, and reports the pipeline's status—the status of tail—instead of the test.

For Bash verification scripts, start with set -Eeuo pipefail, capture statuses immediately, and make cleanup return the original code. In GitHub Actions, inspect both the step's pre-exception outcome and its final conclusion when continue-on-error is present.

Why | tail hides failures

Without Bash's pipefail option, a pipeline returns the status of its last command. Logging tools such as tail, tee, and sed often succeed after the program on their left fails. The log contains the failure, but the shell presents zero to the caller.

# The test exits 1. tail reads its input and exits 0.
npm test 2>&1 | tail -n 80
printf 'reported status=%s\n' "$?"   # commonly 0

# This is also unsafe without pipefail.
pytest -q 2>&1 | tee test.log
echo "verification passed"

The bug is not cosmetic. A wrapper, agent supervisor, or CI runner receives zero and marks the step successful. Later agents use that false success as their baseline. Always decide which command owns success before adding output filters.

#!/usr/bin/env bash
set -o pipefail

npm test 2>&1 | tail -n 80
rc=$?
printf 'test pipeline status=%s\n' "$rc"
exit "$rc"

With pipefail, the pipeline returns the rightmost nonzero command status, or zero only when every component succeeds. That is usually the desired rule for verification logs. It still does not tell you which component failed when several fail; use PIPESTATUS for that detail.

Use set -euo pipefail as a baseline, not magic

Strict mode reduces accidental success. The -e option exits on many unhandled nonzero statuses, -u rejects unset variables, and pipefail exposes failures inside pipelines. Add -E when an ERR trap should be inherited by shell functions, command substitutions, and subshell contexts.

#!/usr/bin/env bash
set -Eeuo pipefail

verify() {
  npm run lint
  npm test 2>&1 | tee artifacts/test.log
  npm run build
}

mkdir -p artifacts
verify

Bash has documented exceptions to errexit. Failures used as an if condition, in most && or || lists, after !, and in non-final pipeline commands without pipefail may not terminate the script. Use explicit branches when failure is expected, and never assume -e replaces status handling.

# Honest expected-failure handling
if ! npm test -- --runInBand; then
  printf 'Focused tests failed; rejecting agent output.\n' >&2
  exit 1
fi

# Dangerous: hides every failure and continues
npm test || true

Capture PIPESTATUS[0] immediately

Bash stores the status of every command in the most recent foreground pipeline in the PIPESTATUS array. Index zero is the first command. Copy the array immediately: almost any later command, including a diagnostic echo, replaces it.

set +e
npm test 2>&1 | tee artifacts/test.log | tail -n 120
statuses=("${PIPESTATUS[@]}")
set -e

test_rc=${statuses[0]}
tee_rc=${statuses[1]}
tail_rc=${statuses[2]}

printf 'test=%s tee=%s tail=%s\n' \
  "$test_rc" "$tee_rc" "$tail_rc"

(( test_rc == 0 )) || exit "$test_rc"
(( tee_rc == 0 )) || exit "$tee_rc"
(( tail_rc == 0 )) || exit "$tail_rc"

Temporarily disabling -e lets the script inspect all statuses instead of exiting before it can write evidence. The important part is restoring strict mode and explicitly exiting nonzero. If only the producer matters, capture rc=${PIPESTATUS[0]} immediately.

This array is Bash-specific. If a script declares #!/bin/sh, do not use PIPESTATUS or assume pipefail exists. Either require Bash explicitly or avoid a pipeline by sending command output to a file and displaying it afterward.

#!/bin/sh
log=artifacts/test.log
if npm test >"$log" 2>&1; then
  tail -n 120 "$log"
else
  rc=$?
  tail -n 120 "$log"
  exit "$rc"
fi

Do not let cleanup traps overwrite the exit code

Traps are valuable for deleting temporary files and writing run receipts. They become dangerous when the trap runs another command before saving $?, or explicitly exits zero. Capture the incoming status as the first action, disable recursive traps if necessary, perform cleanup, then return that original status.

# Broken: explicit exit 0 turns failure into success.
cleanup_bad() {
  rm -f "$tmp_file"
  exit 0
}
trap cleanup_bad EXIT

# Correct: preserve the status that triggered EXIT.
cleanup() {
  rc=$?
  trap - EXIT
  rm -f -- "$tmp_file" || true
  printf 'verification_exit=%s\n' "$rc" >artifacts/exit.txt
  exit "$rc"
}
trap cleanup EXIT

An ERR trap has the same evidence rule. Save the status before formatting messages. Quote optional variables safely when nounset is enabled, and keep cleanup errors separate from the primary verification failure.

on_error() {
  rc=$?
  line=${BASH_LINENO[0]:-unknown}
  printf 'command failed: rc=%s line=%s\n' "$rc" "$line" >&2
  return "$rc"
}
trap on_error ERR

Audit GitHub Actions continue-on-error

GitHub Actions can deliberately allow a failed step or job to continue. For a step with continue-on-error: true, the original outcome can be failure while its final conclusion is success. That may be correct for an experimental matrix entry, but it must not authorize accepting an agent's patch.

- name: Verify agent patch
  id: verify_agent
  continue-on-error: true
  shell: bash
  run: |
    set -Eeuo pipefail
    ./scripts/verify-agent-patch.sh 2>&1 | tee verify.log

- name: Upload verification log
  if: always()
  uses: actions/upload-artifact@v4
  with:
    name: agent-verification
    path: verify.log

- name: Enforce verification result
  if: always()
  shell: bash
  env:
    VERIFY_OUTCOME: ${{ steps.verify_agent.outcome }}
  run: |
    if [[ "$VERIFY_OUTCOME" != success ]]; then
      printf 'Agent verification outcome: %s\n' "$VERIFY_OUTCOME" >&2
      exit 1
    fi

Search the workflow and reusable actions for continue-on-error, || true, unconditional success markers, and verification steps guarded by conditions that can skip them. Also inspect matrix fail-fast settings: they control cancellation behavior, not whether a failed experimental job is acceptable evidence for the required job.

Verification checklist before accepting agent output

  1. Read the task, acceptance criteria, allowed files, and claimed verification commands.
  2. Inspect git status --short and the complete diff, including generated and untracked files.
  3. Run the smallest focused check yourself from a known directory and clean enough state.
  4. Confirm the invoked shell matches the script: Bash features require Bash.
  5. Audit pipelines for pipefail or immediate PIPESTATUS capture.
  6. Audit traps and wrappers for preserved nonzero exit codes.
  7. Search CI for downgraded failures, skipped gates, and unconditional success output.
  8. Run the broader suite needed to catch interactions outside the focused change.
  9. Review behavior and assertions; exit zero proves execution, not test quality.
  10. Accept only when the patch, commands, statuses, and CI policy all support the same conclusion.
git status --short
git diff --check
git diff --stat
git diff

rg -n 'continue-on-error|\|\| true|set \+e|trap .*EXIT' \
  .github scripts package.json

bash -n scripts/*.sh
./scripts/verify-agent-patch.sh
printf 'independent verification exit=%s\n' "$?"

Keep the raw command, environment-relevant version information, exit status, and log artifact together. Do not accept a pasted excerpt that omits the start of the command or the final status. If a check fails, classify the failure before retrying: product defect, flaky test, environment mismatch, missing credential, or broken verifier. Changing the verifier merely to obtain green is not a fix.

The safest mindset is adversarial but fair. Assume the agent may have made an honest verification mistake, and make success mechanically difficult to misreport. A trustworthy green result is one you can reproduce without relying on the agent's summary.

Test the verifier with a known failure

A verification script needs its own negative control. Run a tiny command that is guaranteed to fail through the same logging, trap, wrapper, and CI path used by real tests. The outer process must finish nonzero, the log must remain available, and any final receipt must record the same failure. This catches silent-green wiring before it evaluates an agent patch.

#!/usr/bin/env bash
set -Eeuo pipefail

run_logged() {
  local log=$1
  shift
  "$@" 2>&1 | tee "$log"
}

set +e
run_logged artifacts/negative-control.log \
  bash -c 'printf "EXPECTED_FAILURE\n"; exit 42'
rc=$?
set -e

if [[ "$rc" -ne 42 ]]; then
  printf 'verifier corrupted exit status: got %s\n' "$rc" >&2
  exit 70
fi
rg -q '^EXPECTED_FAILURE$' artifacts/negative-control.log
printf 'negative control preserved exit 42\n'

Notice that the harness does not exit 42 after confirming the control. The known failure is test data; the harness succeeds only when it observes that failure accurately. In production verification, the tested command's nonzero status must instead reject the patch. Keep those two intentions explicit so a helper designed for negative controls is not reused as a permissive production wrapper.

Also test interruption. Conventional shell statuses above 128 often indicate termination by a signal, but do not infer the cause from the number alone. Preserve runner logs and cancellation metadata. A timeout, operator cancellation, out-of-memory kill, and deliberate termination all mean verification was not completed; none should be converted into a pass merely because the product assertions did not print a failure.

Finally, run the verifier against an unchanged known-good commit and a deliberately broken fixture. Those controls establish both directions: it can accept a valid baseline and reject an invalid one. If either control disagrees, fix the verification surface before asking another agent to modify product code.

Primary sources