A start-time sweep of a shared scratch root deletes a concurrent sibling's live fixtures, so a suite that passes in every process alone fails only when sharded
A test suite keeps its sandboxes under one fixed directory and, at module load, recursively removes that directory so a previously killed run leaves nothing behind. That is correct while the suite is the only process using the directory. The moment the suite is sharded across processes, every shard's load-time sweep deletes the fixtures its siblings have already created and are mid-assertion on; the victim reads a missing file and fails, and which shard is the victim depends on start order, so the failure moves from run to run. Every shard passes alone and the serial run passes, so the only failing configuration is the one that was just introduced, and it reads as flakiness rather than as a deterministic race on the sweep.
Resurfaces when
- sharding or parallelising a test suite across processes
- adding a start-time cleanup that recursively removes a fixed scratch directory in a test file
- writing a comment that says a test assumes it is not running concurrently with itself
- one assertion fails only when shards run together and passes in every shard run alone
- a test failure that moves between shards or workers from run to run
- a parallel test run reports a missing fixture file that the serial run never misses
The lesson’s retrieval contract, weighted three times heavier than its body. A lesson that states when it applies doesn’t need a semantic search to find it.
The lesson
A cleanup step that runs when a test file loads has an unstated precondition: that nothing else is using what it cleans. A suite that sweeps a fixed scratch root at start-up is safe for exactly as long as it is the only process that will ever touch that root — and the day it is sharded across processes, that precondition breaks in every process at once. Each shard's sweep is correct from its own point of view and destructive from every sibling's.
The failure is deterministic and looks random. The sweep always runs; which sibling has already created a fixture by then depends on start order, so a different assertion fails on each run, and it fails only when the shards run together. Every shard passes alone, the serial run passes, and the only failing configuration is the one that was just introduced. That distribution reads as flakiness, and a flaky safety test is worse than a slow one: it gets ignored, and then the thing it guards is unguarded.
The property the start-time sweep exists for — cleaning up after a run that was killed, which is the run that leaves a mess — must survive the fix. Deleting the sweep trades a race for leftovers. The shape that keeps both is a per-process subdirectory and a sweep that removes only directories whose owning process is gone.
The failure that taught it
A verification gate's slowest step was a hook selftest: hundreds of process spawns, each a real invocation of the code under test, spread evenly across dozens of suites with no hot spot to fix. Sharding it across eight processes was the only lever. The sharded run was fast, the assertion set merged to exactly the serial set, and one assertion failed — a different one on the next run.
Each shard passed when run by itself. Three sequential shard runs were clean. Only concurrent execution failed. The suite kept its sandboxes under one fixed directory beside the test file and removed that directory recursively at module load, so the eighth shard to start deleted whatever the first seven had built. The comment above that line had already said it: the sweep assumed the suite never ran concurrently with itself, and named the per-process subdirectory as the fix if that ever changed. The fix was that. The sweep now skips any directory named for a process that still answers a signal-zero probe, and removes the rest. Three consecutive concurrent runs afterwards: zero failed shards.
The first attempt at the sharding had asserted independence from the fact that every suite used its own session id and its own temporary directory. That rules out ordering collisions and nothing else; a plant that runs shards sequentially cannot find a race between them. The assertion that was wrong had been written with confidence, minutes earlier, about code the same session had just read.
How to apply it
- Before parallelising any test suite, grep it for a recursive remove of a fixed path at module scope — a constant built from the file's own directory, swept before anything runs. That line is a shared-mutable-state declaration wearing cleanup's clothes.
- Read the comment beside it. If it says the suite assumes it is not running concurrently with itself, that assumption is about to become false, and the comment usually names the fix.
- Give each process its own subdirectory under the shared root, keyed by pid, and make the
start-time sweep remove only directories whose pid no longer answers
kill(pid, 0): a no-such-process error means dead, a permission error means alive and someone else's, and a name that is not a pid at all is stale by definition. That keeps the killed-run cleanup while leaving a live sibling alone. - Test both directions of that sweep, not just the fix: plant a dead-pid directory and a non-pid directory and watch both go; start a real process, plant a directory named for it, and watch it survive.
- Then prove independence the only way it can be proved — run the shards concurrently, several times, and check exit codes, not just assertion counts. A merged count that matches the serial run says every suite executed; it says nothing about whether any of them failed.
Carries a runnable check
It names a grep-able signature — the shape this failure takes on sight. Prose doesn’t prevent recurrence; executable checks do. This one runs today, against every edit, as signature-scan.
Where this claim comes from
- Lesson
A start-time sweep of a shared scratch root deletes a concurrent sibling's live fixtures, so a suite that passes in every process alone fails only when sharded
- Distillationmechanism stated
Written from the mechanism, not the incident — which is what lets it transfer to code sharing nothing with the original.
- Scar1 occurrence
One recorded failure — weaker evidence, and ranked accordingly rather than presented as settled.
- Evidencenone recorded
No source records recorded — hand-written and migrated lessons predate the pipeline that captures them.
What happened when it was used
- 15
- retrieved
- 2
- acted on
- 2
- held up
- 0
- did not hold up
Applied 2 times. Outcomes move the ranking both ways, which is what makes this improve rather than just grow.