A set-level gate that reads the top-ranked candidate before the per-item filters lets a tie-break decide whether any answer exists at all
A filter chain holds two rejections of different arity — a set-level one that can return nothing at all, and a per-item one that removes individual candidates — and the set-level test runs first, reading a property of the ranked list's head; so the head is whatever won the score race, including candidates the per-item filter is about to delete, and any tie-break-sized difference in score silently decides whether the caller is told the collection has an answer or has nothing.
Resurfaces when
- adding a check that can reject an entire result set rather than individual entries
- writing a filter chain where one rule rejects the whole output and another rejects single entries
- deciding where in a filter chain to place a test that can empty the output
- about to read index 0 of a sorted list to decide something about the whole list
- a filter chain returns nothing even though a valid entry was present and near the front
- an output that used to be populated becomes empty after the input collection grew
- a small change to a weighting flips an output between populated and empty
- two checks in one path reject at different arities and nothing records why they run in that order
The lesson’s retrieval contract, weighted three times heavier than its body. A lesson that states when it applies doesn’t need a semantic search to find it.
The lesson
When a filter chain can reject at two different arities — one rule that returns nothing at all, and another that removes individual candidates — the order between them is not a refactoring detail. It decides which candidates the set-level rule is allowed to look at.
Put the set-level rule first and it reads the ranked head, which is whatever won the score race. That head may be an item the per-item filter is about to delete. The whole answer is then discarded on behalf of a candidate that was never going to be returned.
The deeper error is letting rank decide existence. Ranking answers "in what order", and its inputs are tuned for ordering — priors, recency, usage multipliers, tie-breaks. None of those are evidence about whether the collection contains an answer. Wire them into an existence decision and a difference far below the resolution anyone reasons about becomes the difference between "here is your answer" and "there is nothing here."
The failure that taught it
A document-lookup layer had a floor that returned an empty result set when a query's vocabulary looked unanswerable by the corpus, with an escape hatch: if the top hit was decisively the answer, return it anyway. The escape read the top-ranked candidate.
A document's own declared trigger phrase — the corpus's own statement of what that document is for — stopped retrieving it. Scored on the live index, the wanted document was ranked second at a score of 28.254, behind 28.356. A gap of 0.36%, opened entirely by a usage-feedback multiplier that has nothing to do with relevance. The escape hatch inspected the leader, found it undistinguished, returned false at its first line, and the floor emptied the answer.
The leader was itself below the per-item coverage floor. It was never going to be shown to anyone. So a rejection that could delete the entire answer consulted a candidate that the very next statement deletes, and the layer that would have cleaned this up correctly ran one step too late to be asked.
Two things made it hard to see:
The regression had no author. Nothing in the lookup code changed. Enough new documents landed to shift a multiplier, and a threshold calibrated against the old corpus quietly stopped holding. A floor tuned against a corpus is a claim about that corpus, and it expires without anyone editing it.
The first fix passed the tests and was wrong. Widening the escape to scan the top three candidates instead of only the first made the failing check green and every negative case still return nothing — so it looked correct. A different suite caught it: a fixture living outside the main negatives found an unrelated document sitting at rank two with a perfect intent score, because that query had only two scorable terms and matching both is a coincidence over a two-term vocabulary, not a declared contract. Separating the good case from the bad one would have meant fitting a constant between one negative at 0.79 and one positive at 0.9964 — tuning to two points. Reordering the two rejections needed no constant at all, and the coincidental match still correctly yields nothing, because it is not the leader and the leader is not decisive.
How to apply it
- Inventory your rejections by arity. Which can return nothing at all, and which remove single items? If both exist in one path, the ordering between them is a decision — make it deliberately and write down why.
- Run the per-item filters first, then ask the set-level question of the survivors. Whether an answer exists should be decided from the candidates that could actually be delivered. Asking afterwards costs nothing and removes the entire class.
- Never let a ranking input reach an existence decision. Priors, usage counts, recency boosts and tie-breaks answer "in what order", not "is there anything here". If a 1% score change can flip your output between populated and empty, that is the defect, not the threshold's value.
- When a set-level check regresses after the data grew, suspect the ordering before the constant. Retuning a threshold against the new data reproduces the same expiry later; the ordering fix has no constant to go stale.
- A fix that turns the failing check green has not been shown to be right. Sabotage it: break the mechanism deliberately and confirm the cases it claims to guard go dark. Then check the suites that do not share fixtures with the one you were fixing — the first wrong fix here was green on every case the failing suite could see.
Carries a runnable check
It names a grep-able signature — the shape this failure takes on sight. Prose doesn’t prevent recurrence; executable checks do. This one runs today, against every edit, as signature-scan.
Where this claim comes from
- Lessonhigh confidence
A set-level gate that reads the top-ranked candidate before the per-item filters lets a tie-break decide whether any answer exists at all
- Distillationmechanism stated
Written from the mechanism, not the incident — which is what lets it transfer to code sharing nothing with the original.
- Scar1 occurrence
One recorded failure — weaker evidence, and ranked accordingly rather than presented as settled.
- Evidencenone recorded
No source records recorded — hand-written and migrated lessons predate the pipeline that captures them.
What happened when it was used
- 10
- retrieved
- 1
- acted on
- 1
- held up
- 0
- did not hold up
Applied 1 time. Outcomes move the ranking both ways, which is what makes this improve rather than just grow.