A code-read verdict is a hypothesis, not a result
- Evidence
- 17 recorded incidents
- Derived from
- 12 narrated failures
- Status
- active
- In force since
- Retrieved
- 417 times by an agent
- Outcome
- confirmed· 127 applications held up in use
17 independent failures forced this. Each is narrated below rather than summarised away: the incident is what makes the rule credible to the next person, and to an agent deciding whether to apply it.
Why this rule exists
- Rule17 scars
A code-read verdict is a hypothesis, not a result
- Promotion
A mechanism that keeps recurring has proved that writing it down didn’t prevent it. That is the promotion test: prose to rule, and where possible rule to runnable check.
- Incidents12 narrated
Starting with: "ABSENT" did not survive contact with the device.
- Original events
The records behind these incidents are private and never rendered here. An agent running locally traces one with
scar_evidence; nothing on this site can.
The incidents behind it
- 01"ABSENT" did not survive contact with the device
A code read concluded a feature was missing. Testing on a real device found it present and working — the read had traced the wrong code path. One in fourteen such verdicts failed this way, which is exactly the rate that makes code-reads feel reliable enough to…
- 02Two reported bugs were refuted, and in both the *report's* read was wrong, not the app
Both had the same failure shape: the feature was looked for in the place the report assumed it would be, rather than the place the reference implementation actually put it. Checking the reference *first* would have caught both, cheaper than a device pass did.
- 03A hit counter caught a test aimed at the wrong target
A fault-injection harness was armed against one endpoint while the screen under test used a differently-named one. The app behaved perfectly, which reads exactly like "not reproducible" — a false negative that would have closed a real bug. Only the harness's o…
- 04Absence of evidence in a log is not evidence of absence
Missing request lines in a console log read cleanly as "this call is never issued" — a far larger bug than the real one, and it was one step from being written up. The instrumentation simply wasn't logging that request type. The hit counter disproved it in a s…
- 05A checker failed silently, in the flattering direction
A status classifier used a substring match that matched the word it was meant to exclude — it treated "unblocked" items as "blocked". Freshly-resolved work stayed counted as outstanding. Substring tests over human-written text fail quietly and, characteristica…
- 06Four disproven hypotheses in a row meant the missing input was observation, not thought
A bug survived four consecutive static-reasoning passes. Each pass produced a confident causal story, and each was disproven the moment it was checked live. The fifth attempt succeeded only because the user supplied screenshots. The lesson is a stopping rule r…
- 07Five hypotheses died before anyone asked the app what it could see
A selection step stopped appearing after a client was pointed at a different backend. Five causal stories were produced and killed in turn — a redirect race, a parameter lost on the first render, a stale bundle, the wrong code path, a mis-set flag — each one p…
- 08A linter reported clean on the one branch it had been exempted from checking
The frontmatter linter's YAML-strictness audit skipped flow collections wholesale, on the reasonable- sounding assumption that brackets make a value unambiguous. Brackets make the *collection* unambiguous; every item inside is still a scalar with YAML's ordina…
- 09The planted violation was planted in a shape the real system never sends
The hypothesis gate's disarm-on-resolve shipped with a suite that armed it, resolved it, and asserted six clean edits afterwards — the discipline of scars 4, 6 and 10, applied properly. It was still inert in production for its entire life. An MCP tool result d…
- 10The planted probe was planted in dead code
A session following this rule — proving a freshly wired verify step could fail before trusting it — appended its `assert(false)` probe to the end of the selftest with a shell `>>`, which put it *after* the script's `process.exit()`. The step ran, verify passed…
- 11Verification was applied to code all session, and never once to the prose describing it
An end-of-session audit checked every claim in a session's own written comments against source: 87 true, 11 false. The distribution is the finding, not the ratio. Every citation into a file the session had not written was correct — thirteen for thirteen. Every…
- 12Every plant proved the mechanism fired; none proved the property still held
A day of changes to this repo's own gates was verified throughout by planting real violations and watching them fail — the discipline scars 4, 10, 11 and 12 all ask for. Two hostile review passes afterwards found five defects anyway, and every one sat in code …
Summarised. The full narrative for each is in the rule below.
The rule
Reading code tells you what it should do. Only running it tells you what it does. When a conclusion matters — a bug is "confirmed", a feature is "missing", a fix "works" — the read is where the investigation starts, not where it ends.
And when you do instrument: verify the instrument fired. A check that has never caught a planted violation is unverified, no matter how clean its output looks.
The scars
All the same shape from different angles: a confident conclusion that reading produced, and running
destroyed. The count lives in scars: above, where frontmatter-lint verifies it against the
incidents below — it deliberately is not restated here, because it read "Seven" while the body held
thirteen, in the one rule whose entire subject is claims that were never checked.
1. "ABSENT" did not survive contact with the device. A code read concluded a feature was missing. Testing on a real device found it present and working — the read had traced the wrong code path. One in fourteen such verdicts failed this way, which is exactly the rate that makes code-reads feel reliable enough to trust and occasionally very wrong.
2 & 3. Two reported bugs were refuted, and in both the report's read was wrong, not the app. Both had the same failure shape: the feature was looked for in the place the report assumed it would be, rather than the place the reference implementation actually put it. Checking the reference first would have caught both, cheaper than a device pass did.
4. A hit counter caught a test aimed at the wrong target. A fault-injection harness was armed
against one endpoint while the screen under test used a differently-named one. The app behaved
perfectly, which reads exactly like "not reproducible" — a false negative that would have closed
a real bug. Only the harness's own hit counter (hits: 0) revealed the fault had never fired.
5. Absence of evidence in a log is not evidence of absence. Missing request lines in a console log read cleanly as "this call is never issued" — a far larger bug than the real one, and it was one step from being written up. The instrumentation simply wasn't logging that request type. The hit counter disproved it in a single run.
6. A checker failed silently, in the flattering direction. A status classifier used a substring match that matched the word it was meant to exclude — it treated "unblocked" items as "blocked". Freshly-resolved work stayed counted as outstanding. Substring tests over human-written text fail quietly and, characteristically, in whichever direction makes the numbers look more normal.
7. Four disproven hypotheses in a row meant the missing input was observation, not thought. A bug survived four consecutive static-reasoning passes. Each pass produced a confident causal story, and each was disproven the moment it was checked live. The fifth attempt succeeded only because the user supplied screenshots. The lesson is a stopping rule rather than another way to read code: once N hypotheses have been generated and killed, the binding constraint is no longer the quality of the reasoning — more of it will keep producing plausible stories at the same rate. It is the absence of observation, and the cheapest next move is to go get some.
8. Five hypotheses died before anyone asked the app what it could see. A selection step stopped appearing after a client was pointed at a different backend. Five causal stories were produced and killed in turn — a redirect race, a parameter lost on the first render, a stale bundle, the wrong code path, a mis-set flag — each one plausible from the source alone. Every round consumed a full log capture. Nobody logged the single value the branch actually read: how many options the screen had been given. The reasoning was never the constraint; the observation was one line away the whole time. Same shape as scar 7, and it recurred anyway, which is why the stopping rule is written as a count rather than as a judgement call.
9. The stopping rule was delivered by the mechanism built for it, and the tool was still never
called. A debugging session was handed this rule's protocol note explicitly, in the
scar_recall response, at the top of the task. It then called scar_hypothesis zero times
and burned four disproven theories before an observation — a pasted server log — found the real
cause: a schema migration that had been written but never run. This is the third recorded instance
of the rule being read and not followed, and the first where the enforcement was also present and
also skipped, which is why the answer was a gate that fires uninvited (tools/hooks/ hypothesis-gate.mjs) rather than a fourth rewrite of the words. It recurred on 2026-09-17 with that gate armed: tuned that
morning to count edits, it watched 108 observation calls go by with none — so it now counts a run of
observations too (OBSERVATION_STREAK; decision 170).
10. A linter reported clean on the one branch it had been exempted from checking. The
frontmatter linter's YAML-strictness audit skipped flow collections wholesale, on the reasonable-
sounding assumption that brackets make a value unambiguous. Brackets make the collection
unambiguous; every item inside is still a scalar with YAML's ordinary rules. A lesson wrote a
quoted phrase followed by more text inside a [...] list, which the permissive in-house parser
accepted and a real YAML parser rejects — so the full gate passed, repeatedly, on a corpus the
viewer could not parse at all. The build failure that surfaced it named an unrelated page and
printed a red-herring warning beside it; the cause appeared only when the viewer's own parser was
run directly against the corpus. Same shape as scars 4 and 6 — a check that reports success
because it never looked — and the fix was verified by planting a fresh violation, because a gate
that has only ever seen the bug it was written for has not been tested.
11. The planted violation was planted in a shape the real system never sends. The hypothesis
gate's disarm-on-resolve shipped with a suite that armed it, resolved it, and asserted six clean
edits afterwards — the discipline of scars 4, 6 and 10, applied properly. It was still inert in
production for its entire life. An MCP tool result does not arrive as the object the handler
returned; it arrives as content blocks with that object already serialised into a string field, so
the gate's JSON.stringify escaped the inner quotes and its "action":"resolve" test matched
nothing. The fixture passed a flat {action:'resolve'} — a shape the harness has never once sent —
under a comment reading "driven exactly as the harness drives it". Every consequence pointed away
from the cause: the gate over-fired, which reads as too-aggressive arming, not as a disarm that
never ran, and the suite was green throughout. It surfaced only from reading a real transcript's
payload beside the fixture. A planted violation proves the check fires; it proves nothing about
whether the input resembles production. And the correction had the same failure a second time
one layer down — the first replacement fixture spent both denials before the resolve, so
MAX_DENIALS silenced the checks and it passed against the unfixed code too. A test asserting
silence has to rule out every other route to silence.
12. The planted probe was planted in dead code. A session following this rule — proving a
freshly wired verify step could fail before trusting it — appended its assert(false) probe to
the end of the selftest with a shell >>, which put it after the script's process.exit().
The step ran, verify passed 13/13, and the "planted failure" was unreachable code the whole time.
Nothing in the output distinguished "wired and proven" from "wired and unproven" except one
number: the check count still read 19 where the probe should have made it 20. Re-planted before
the exit, verify failed loud and the wiring was actually proven. The addition to scar 11: a
planted violation proves nothing by being planted — only by being observed to fire. The proof
is the failure output, never the plant; a probe you did not watch fail has the evidentiary value
of no probe, and dead code delivers that nothing in the flattering direction scar 5 warns about.
13. Verification was applied to code all session, and never once to the prose describing it. An end-of-session audit checked every claim in a session's own written comments against source: 87 true, 11 false. The distribution is the finding, not the ratio. Every citation into a file the session had not written was correct — thirteen for thirteen. Every claim about a third-party library's internals, read out of its source, was correct. All eleven false claims described code the session had written itself, minutes earlier: one stated the opposite of the code it sat above and contradicted two other comments in the same change; one invented a plausible-sounding concurrency rationale that the code's actual ordering makes impossible; one quoted a measurement computed under a default that had since changed, and was internally inconsistent on its own terms. This rule's checklist had been followed all session — a reviewer's claims were re-verified against source repeatedly, and it caught a real error. It was never once aimed at the session's own prose. Effort had tracked how DANGEROUS a claim felt, not how likely it was to be wrong: a line number in an unfamiliar file felt risky and got checked; a comment written twenty minutes earlier felt like memory rather than a claim, and got none. The felt-safe half was the wrong half.
14. Every plant proved the mechanism fired; none proved the property still held. A day of changes to this repo's own gates was verified throughout by planting real violations and watching them fail — the discipline scars 4, 10, 11 and 12 all ask for. Two hostile review passes afterwards found five defects anyway, and every one sat in code that already had a passing planted test. The plants were uniformly of the form does this fire when I break it; the defects were uniformly of the form the thing being protected is no longer protected in some state the plant never created. Three states did it. A missing dependency: pooling the token budgets made a pool skip whenever any member was absent, and one member is a derived file that does not exist before the first build — so on a fresh clone every budgeted file was enforced by nothing, measured at 394% of cap with a clean exit. A changed population: a corpus-health ratio was fixed once for a numerator/denominator mismatch, declared correct, and still was not invariant, because the population it divides by grows as a backlog clears — the metric travels 16.9% to 21.0% with nothing written. A thrown exception: a selftest wrote a fixture into the real corpus and removed it only on the happy path, harmless while the runner stopped at the first failure and a three-way false alarm once it stopped stopping. The plant and the property are different propositions, and passing the first is routinely mistaken for establishing the second — note that this happened in a session that was applying scars 4 and 12 deliberately and citing them while doing it.
15. A cheap check that answers a nearby question was read as the answer to the one asked — three times in one day, in three dialects, with the matching lesson already on screen. A build session made nine recalls and was handed the right lessons every time. One was quoted aloud to the user — "a section existing in the reference UI is not proof the payload you fetch can fill it" — and the next edit built the section on the field whose name matched. It returned empty; typecheck and lint were green; only a device run showed the section never rendering. The same session validated an image route with a headers-only probe, saw a redirect, and built on it; the redirect went to the 404 page, and the lesson about exactly that was in the corpus with the session's own recurrence filed against it. One shift later, a DOM text match was reported to the user as a verified finding that an element rendered, and the user corrected it. Each dialect — status-for-existence, presence-for-rendered, name-for-contents — had or now has its own lesson, and none fires for the others, because the instance vocabulary is disjoint. The shared shape is the scar: a check that is cheap to run answers whichever question it is cheap to answer, and that question sits one step short of the one you asked. It was not classified as verification at the time, in every case, because it did not feel like a verdict — it felt like a quick look on the way to the real work. That is why retrieval and delivery both worked and behaviour did not move: the moment the lesson applied was a moment nobody labelled risky, so nothing was checked against it. Before building on any probe, name the question the probe actually answered and the question you needed answered, and if they differ by one step — exists vs. resolves, present vs. rendered, named vs. populated — take the step.
16. A chained grep returned nothing, and "the reference does not split it" was said aloud on
that nothing. As recorded in the usage report
feedback/2026-09-14-a-narrow-search-tool-produced-a-confidently-wrong-claim-abou.md: re-checking
a defect register, a session needed to know whether a reference web client splits a comma-joined
identifier, and ran a line-based chain — a grep for the field name piped into a grep for
split. Empty. The verdict "web does not split it in source either" went to the user, followed by
a second claim about the server that happened to be true, so the pair read as settled parity. The
reference did split it, in a file the chain had read and discarded: the .split( sat four
lines below the field name, and a line filter can only see a pattern that fits on the line it is
looking at. This is scar 5 in a new dialect — absence of evidence read as evidence of absence —
with one difference that made it worse: the session had verified, by its own account, and this
rule had fired usefully twice earlier that day. The check ran; its shape could not observe the
thing being checked, and a negative from a probe that cannot see the target is indistinguishable
from a true absence. It was caught only by an unprompted multi-line scan before a written record
went to the teammate whose code it described. Before speaking a "does not" that rests on an empty
search result, ask what the search could not have seen — an expression spanning lines, a name
built by concatenation, a file the walk skipped — and re-run with a probe whose shape covers it.
17. Two independent silencers, and the how-to-apply step that cannot be followed through a dead channel. A mobile session chasing a missing feature added three diagnostic probes and got nothing back from any of them, which read as "this code path never runs". It was wrong twice over: production builds of that framework do not forward console output to the system log at all, and the device in use suppresses application logs for non-debuggable builds. Two independent reasons for the same silence, and the observation they produce is identical to the one a genuinely dead code path produces. Two build cycles — minutes each — went into probes that could not have printed under any circumstances. What settled it was abandoning the log channel entirely and routing the value through something the platform would render regardless, then reading it back with a system command.
This is scar 5 in a new dialect, and it recurred with this rule already delivered to that very session at its first recall, which makes it the fourth recorded delivery failure alongside scars 9 and 15. But it also names something the rule did not previously resolve. How-to-apply #4 says "always verify the instrument fired" — and when the transport itself is what is dead, there is no way to verify firing through that transport. The step is circular exactly when it is needed most. The addition: before spending a cycle on a probe, emit something unconditional and confirm you can see it. If you cannot, the channel is the finding — switch to one the platform cannot suppress (a value rendered into the UI, a field some system command will dump back) rather than adding a fourth probe to a channel that has never carried anything.
Recorded as evidence, not as the fix. Per core/rules/enforce-dont-declare, a fourth rewriting of
prose is precisely the move this rule's own history says does not work; what this scar argues for
is a mechanism that fires when a session reports an empty probe result, and it is written down so
that the mechanism has something to be built from.
(And one more, from this framework's own construction: its frontmatter linter silently treated its own generated index files as malformed entries. Found only because a violation was deliberately planted rather than trusting a clean run.)
This rule is now enforced, because reading it demonstrably was not enough
Scars 7 and 8 are the same failure, and scar 8 happened in a session that had read this file at
the start. Five theories, five full log captures, and nobody printed the one value the failing
branch actually read. Three rewrites of the prose would not have changed that: guidance delivered
once at session start is ambient, and ambient guidance loses to whatever is in working memory
during a debugging spiral — which is, by construction, the next theory. enforce-dont-declare
already names this class: a discipline not wired to a mechanism is theater. This rule was theater.
So the stopping rule now runs as code. tools/recall/lib/protocol.mjs, exposed as the
scar_hypothesis MCP tool, keeps the count of disproven theories outside the agent's
working memory — where a spiral cannot quietly reset it — and at 3 disproven it stops
returning advice and returns a refusal to continue reasoning, naming the cheapest observations
available. scar_recall attaches the protocol to any debugging-shaped task, so the instruction
arrives while the plan is being formed rather than five rounds in.
Scar 9 says that was still not enough, and the distinction matters. The protocol note was
delivered, on time, by the mechanism above — and the tool was never called. Attaching guidance to a
response is still guidance; it only looks like enforcement because code produced it. So the count
now runs as a hook that fires without being invited (tools/hooks/hypothesis-gate.mjs): it arms
when a recall carries this rule's note and denies further edits after three tool calls with no
scar_hypothesis. Note the direction of travel across scars 7, 8 and 9 — prose, then a tool the
agent may call, then a gate the agent cannot skip. Each step was taken only after the cheaper one
was observed failing, which is the order this rule would ask for.
Threshold at 3, not 4 or 5: scar 7 took four theories and scar 8 took five, so firing at the observed failure count would fire exactly when it is already too late.
The honest caveat: this depends on the agent calling the tool. It converts "remember a paragraph under pressure" into "make one cheap call per dead theory", which is a much easier thing to comply with — but it is not a hard gate, and if scar 9 arrives with the counter unused, the next move is a hook that fires without being invited, not a fourth rewrite of this section.
How to apply
- While debugging, call
scar_hypothesisas each theory dies. One line, at the moment you rule it out — not batched afterwards, which defeats the point.action: "resolve"when the cause is known, so a stale count doesn't fire on the next unrelated bug. - Treat every code-read conclusion as a hypothesis with a confidence level, and say so. "The source suggests X" is honest; "X is broken" after only reading is not.
- Check the reference implementation before believing any "this is missing" claim — most missing-feature reports are misdirected searches, not absences.
- When reading isn't conclusive, build the instrument. A small harness that manufactures the condition beats another hour of speculation.
- Always verify the instrument fired before believing a negative result. Hit counters, planted violations, deliberate failures — a clean run from an unproven check means nothing. Then check the violation you planted is shaped like production — capture one real payload and assert against that, not against the object your own code constructs. And run the new assertion against the unfixed code: if it does not fail there, it is not testing what you think.
- Be suspicious of results that make things look normal. Failures that flatter are the ones that survive review.
- Remember that a harness which manufactures a condition also manufactures the symptom's apparent severity — a bug that only appears under a 25-second artificial delay should not be written up in language that implies users see it constantly.
- Count your disproven hypotheses. After three or four have died, stop generating a fifth and go acquire an observation instead — a screenshot, a log line, a live run. Reasoning that keeps producing plausible-and-wrong answers is not short of effort; it is short of input.
- Before calling a change done, re-verify the comments and commit messages you wrote about it — and check those FIRST, not last. Scar 13: every citation into unfamiliar code got checked and was correct; every false claim described code written minutes earlier and felt like memory rather than an assertion, so it went unchecked. Felt risk runs backwards from actual error rate. Treat a file:line citation, a claim that something is "the only" instance of a shape, or any number presented as measured — in a comment, a commit message, or a knowledge-base entry — as a claim needing the same check a reviewer's claim gets, whether or not you wrote the code it describes minutes ago.
This rule can be wrong
A hypothesis with 17 confirmations, not a law. If an agent applies it and still fails, that is recorded against the rule. Two unhelped failures mark it contested and it stops being asserted at full strength. A knowledge base that cannot demote its own claims only grows.