A negative plant that an earlier condition already rejects never reaches the guard it was written for
A verdict built from several conditions checked in order returns on the first one that fails, so a negative test only exercises the condition that rejects it FIRST. A plant written to prove a later guard (usually the newest, added because the earlier ones passed something they should not have) is typically generic bad input, and generic bad input fails the oldest and cheapest condition, so the suite stays green with the new guard deleted. The input the new guard exists for is by construction one that passes every earlier condition, which is exactly the shape a generic plant does not have.
Resurfaces when
- adding a second condition to a pass/fail bar because the first one passed a case it should not have
- writing a negative test for a verdict that combines several checks in order
- planting shuffled, random or constant data to prove a new significance or sanity check refuses it
- deciding whether a green selftest proves a newly added guard works
- a classifier beats its baseline by a few points on a handful of minority-class examples
- a guard that was deleted to confirm its test fails leaves the suite green
The lesson’s retrieval contract, weighted three times heavier than its body. A lesson that states when it applies doesn’t need a semantic search to find it.
The lesson
When a pass/fail verdict is a chain of conditions, a negative test proves only the condition that rejects it first. Everything after that condition never runs for that input.
This matters most when a guard is added later. A new condition usually exists because the earlier ones let something through. That something is the one input the new guard is for, and it has a specific shape: it passes every earlier condition. A generic bad input (random data, a constant, obviously broken values) almost never has that shape. It fails the oldest and cheapest check, the verdict returns there, and the test goes green whether or not the new guard exists.
Nothing in the test's output shows this. The assertion is "the bad input is refused", and it is refused, just not by the code the test was written for.
The failure that taught it
A small classifier was being judged against a stated bar: beat a constant-answer baseline and a per-category lookup table on balanced accuracy. On real data it did beat both, by about four points, on a test window with only eight minority-class examples. The first draft of the bar called that a success. Eight examples is well within chance, so a second condition was added: the model must also beat versions of itself trained on shuffled labels, at p < 0.05. On the real data it got p = 0.17, so the verdict correctly became "not distinguishable from chance".
The selftest for the new condition planted shuffled labels and asserted the verdict did not read as signal. It passed. With the p-value check deleted, it still passed. The shuffled plant did not beat the baselines, so the verdict returned on the first condition and never evaluated the p-value. The case the guard existed for (beats the baselines, fails significance) was exactly what the plant could not produce. A second plant was added: the observed margin with p = 0.17. That plant failed with the check deleted and passed with it restored.
How to apply it
- When you add a condition to a chain, write down the input that motivated it. That input, or one of the same shape, is the plant. It must pass every condition before the new one.
- If that input is hard to generate end to end, test the verdict function directly with a crafted set of intermediate results (for example, "beats both baselines, p = 0.17") rather than hoping generated data lands there.
- Delete the new guard and run the suite. If it stays green, the plant never reached the guard, however reasonable it looks.
- For small-minority classifiers specifically: "beats the baseline" is not a bar on a handful of positives. Compare against the same model trained on shuffled labels, and report the p-value beside the margin.
Where this claim comes from
- Lessonmedium confidence
A negative plant that an earlier condition already rejects never reaches the guard it was written for
- Distillationmechanism stated
Written from the mechanism, not the incident — which is what lets it transfer to code sharing nothing with the original.
- Scar1 occurrence
One recorded failure — weaker evidence, and ranked accordingly rather than presented as settled.
- Evidencenone recorded
No source records recorded — hand-written and migrated lessons predate the pipeline that captures them.
What happened when it was used
- 3
- retrieved
- 1
- acted on
- 1
- held up
- 0
- did not hold up
Applied 1 time. Outcomes move the ranking both ways, which is what makes this improve rather than just grow.