A shell timeout around a long sequential test runner reports the case that was running when the clock expired, not a case that hung
A wall-clock timeout wrapped around a sequential test runner signals the process when the clock expires, and the runner's synchronous child-process call reports that signal as a failure of whichever spawn was in flight at that instant. The stack trace names a real test case, the error object carries a signal and a null exit status, and stdout and stderr are empty — so the report reads as "this case hangs" when the only fact it records is "this case was executing when time ran out". A suite that legitimately takes longer than the timeout produces this at a different case each run.
Resurfaces when
- about to wrap a long test suite or verify script in a shell timeout so it cannot hang
- a selftest that spawns many child processes and takes minutes to finish
- a test runner failure whose error shows a signal and a null exit status with empty output
- a suite failure that moves to a different case between runs
- running a slow verification gate inside a session with a per-command time limit
- deciding whether a test case hangs or the runner is just slow
- a child-process error with SIGTERM and no stdout
The lesson’s retrieval contract, weighted three times heavier than its body. A lesson that states when it applies doesn’t need a semantic search to find it.
The lesson
A timeout is an observation about the clock, not about the code under it. When an outer
timeout N kills a sequential runner, the runner's synchronous spawn helper throws at whatever
call it was blocked in, and the resulting error looks exactly like a hung test: a real file and
line, a signal, a null status. What distinguishes it is the shape of the evidence — a genuine
failure exits with a status and usually prints something; a killed-by-clock spawn has a signal,
no status, and empty output. Treat that shape as "the suite is slower than the timeout" until the
named case has been run on its own and observed to hang.
The failure that taught it
A repository's verification suite reported its hook selftest as failed. Run directly under a
three-minute shell timeout, the selftest died with a child-process error at one specific check —
a loop that spawned the same gate six times — showing signal: 'SIGTERM', status: null, and
empty stdout and stderr. The obvious reading was that this check hung, and the first theory was a
data-dependent stall in the gate it spawned. The check was then run on its own, seven times: each
call completed in well under a second. The suite simply takes about four minutes end to end, and
three minutes had expired at that call on both runs. The earlier "failure" inside the full gate
had the same cause — a time limit expiring on a slow-but-correct step — and a re-run with the
limit lifted passed with nothing changed.
The cost was one wrong theory and a probe. The risk was larger: the natural next move was to "fix" the innocent test, or to conclude the gate it spawned was broken.
How to apply it
- Before wrapping a suite in a timeout, know roughly how long it takes; set the limit above that or do not wrap it.
- On a runner failure, read the error object before the stack trace. Signal present, status
null, output empty: suspect the clock first. - Reproduce the named case alone, timed, before believing it hangs. If it completes, the case was a bystander.
- If the suite must run under a limit, run it in the background with its output captured, rather than in a foreground call that will be killed and misreport.
Where this claim comes from
- Lessonmedium confidence
A shell timeout around a long sequential test runner reports the case that was running when the clock expired, not a case that hung
- Distillationmechanism stated
Written from the mechanism, not the incident — which is what lets it transfer to code sharing nothing with the original.
- Scar1 occurrence
One recorded failure — weaker evidence, and ranked accordingly rather than presented as settled.
- Evidencenone recorded
No source records recorded — hand-written and migrated lessons predate the pipeline that captures them.
What happened when it was used
- 28
- retrieved
- 7
- acted on
- 7
- held up
- 0
- did not hold up
Applied 7 times. Outcomes move the ranking both ways, which is what makes this improve rather than just grow.