A discipline that isn't wired to a mechanism is theater
- Evidence
- 8 recorded incidents
- Derived from
- 8 narrated failures
- Status
- active
- In force since
- Retrieved
- 147 times by an agent
- Outcome
- confirmed· 16 applications held up in use
8 independent failures forced this. Each is narrated below rather than summarised away: the incident is what makes the rule credible to the next person, and to an agent deciding whether to apply it.
Why this rule exists
- Rule8 scars
A discipline that isn't wired to a mechanism is theater
- Promotion
A mechanism that keeps recurring has proved that writing it down didn’t prevent it. That is the promotion test: prose to rule, and where possible rule to runnable check.
- Incidents8 narrated
Starting with: The loading discipline that was never wired to the loader.
- Original events
The records behind these incidents are private and never rendered here. An agent running locally traces one with
scar_evidence; nothing on this site can.
The incidents behind it
- 01The loading discipline that was never wired to the loader
The framework spent 267 lines arguing for layered, progressive context disclosure — system prompt, then role, then task, then context, with an explicit rule that deeper levels are never injected unless requested. Its actual config force-loaded the entire ~4,60…
- 02A status lifecycle that never moved
Three separate status vocabularies were defined (proposed/accepted/deprecated/superseded, and equivalents). All ten decision records read `Accepted`, all dated the same day. A feature spec still read `Proposed` long after the feature shipped. The statuses live…
- 03Access boundaries enforced socially
Seven agent personas were each given explicit CAN-access / CANNOT-access boundaries, plus a bolded "agents never share context." None of it was enforceable. The *only* real boundary in the project was a glob allowlist in config with a default-deny catch-all — …
- 04A one-off script that declared its one-off-ness in a header comment
A migration script carried a comment saying it was to be run once. Nothing stopped a second run: it renamed filename collisions to `-2.md` and threw on the third, so re-running it as a sync would have duplicated ~756 entries before failing halfway. The comment…
- 05The most-scarred rule in the system, declared and not enforced
`core/rules/verification` had eight scars, sat at the top of the always-loaded kernel, and was read at the start of a session that then committed the exact failure it describes — five theories and five log captures on a bug whose answer was one printed value a…
- 06The rule read from the other side: a mechanism that retains nothing is unfalsifiable
Scars 1–5 are all *declared, not enforced*. This one is *enforced, and unauditable* — and it is the same error, because in both cases the claim outruns what can be checked. In one 2026-08-16 review, three mechanisms in this system were each mistaken for a data…
- 07The one loop step still declared, measured at 14%
This system's consultation loop has three steps in the standing kernel: recall, say what came back, report the outcome with `scar_learn`. Two of them acquired mechanisms after failing (`recall-gate`, and the capture half in scar 5's family). The write-back ste…
- 08Two documents describing the same handoff, and no code between them
Two loop skills in this system were designed as a pair: the read-only one ends with a ledger of findings and states that fixes belong to the other; the other opens by surveying the repo, and was documented as the thing the first one feeds. Both sentences were …
Summarised. The full narrative for each is in the rule below.
The rule
If a convention matters, enforce it in code — a script, a hook, a config allowlist, a linter. Documentation states intent; only mechanism produces behavior. Prose rules are for humans deciding what to do, never for guaranteeing something happens.
The scars
The first three from auditing a prior documentation framework that was thoughtfully designed and almost entirely unenforced.
1. The loading discipline that was never wired to the loader. The framework spent 267 lines arguing for layered, progressive context disclosure — system prompt, then role, then task, then context, with an explicit rule that deeper levels are never injected unless requested. Its actual config force-loaded the entire ~4,600-line documentation corpus into every session. The discipline was real, well-argued, and completely inoperative, because nothing connected it to the thing that loads files.
2. A status lifecycle that never moved. Three separate status vocabularies were defined
(proposed/accepted/deprecated/superseded, and equivalents). All ten decision records read
Accepted, all dated the same day. A feature spec still read Proposed long after the feature
shipped. The statuses lived in prose inside document bodies — not queryable, not linted, with no
defined transitions and nothing to flag a stale one. A status nothing can filter on is a label,
not a lifecycle.
3. Access boundaries enforced socially. Seven agent personas were each given explicit CAN-access / CANNOT-access boundaries, plus a bolded "agents never share context." None of it was enforceable. The only real boundary in the project was a glob allowlist in config with a default-deny catch-all — which worked. The prose was theater; the twelve lines of config were the actual policy.
4. A one-off script that declared its one-off-ness in a header comment. A migration script
carried a comment saying it was to be run once. Nothing stopped a second run: it renamed filename
collisions to -2.md and threw on the third, so re-running it as a sync would have duplicated
~756 entries before failing halfway. The comment was read by the person who wrote it and by nobody
under time pressure. The fix was to make the script idempotent, which is enforcement — but the
class is what matters: "do not run this twice" belongs in the script's control flow, not its
header. Every one-off in a toolchain is a re-run waiting for someone who did not write it.
5. The most-scarred rule in the system, declared and not enforced. core/rules/verification
had eight scars, sat at the top of the always-loaded kernel, and was read at the start of a
session that then committed the exact failure it describes — five theories and five log captures
on a bug whose answer was one printed value away. Prose delivered once at session start is
ambient, and ambient guidance loses to whatever is in working memory under pressure. It is now
wired to scar_hypothesis, which counts disproven theories outside the agent's memory and
refuses to continue at three. Note the shape: the rule that predicts this class was itself the
thing being declared instead of enforced, and it took two scars on the other rule to notice.
6. The rule read from the other side: a mechanism that retains nothing is unfalsifiable. Scars
1–5 are all declared, not enforced. This one is enforced, and unauditable — and it is the same
error, because in both cases the claim outruns what can be checked. In one 2026-08-16 review, three
mechanisms in this system were each mistaken for a dataset: the hypothesis channel (assumed to
retain theories — it deleted them on resolve, the exact moment the real cause became known),
signature-scan (assumed to have hit history — it printed matches to stdout and kept nothing), and
the feedback inbox (assumed to support a rate — it has no denominator and never will, since only
lessons that produced a noticeable event ever enter it). Two of the three were this repo's own
assertions about itself, made by an agent reading its own source.
The distinction that was missing: a mechanism surface — what a system can detect, interrupt, retrieve, gate, test or enforce — and an audit surface — what survives afterward and can be counted, reproduced, compared or challenged. The second is systematically smaller, nothing marks which is which, and a capability then reads as evidence. The same error produced a false sentence in three files at once ("the token budget and retrieval benchmark fail the build" — there is no CI in this repo and neither gate is wired to either build). Two of the three gaps were closed the same day; the inbox's is structural and stays open, which is why it supports a case series and never a percentage.
7. The one loop step still declared, measured at 14%. This system's consultation loop has three
steps in the standing kernel: recall, say what came back, report the outcome with scar_learn.
Two of them acquired mechanisms after failing (recall-gate, and the capture half in scar 5's
family). The write-back step kept only its kernel line — and on 2026-08-19 the ledger priced it:
777 document deliveries, 112 outcome reports, 14.4%; 105 of 139 documents ever delivered had
never once been marked applied. The economics report could not state a return because there was
almost no reported outcome to divide the bill by.
What makes it a scar for this rule rather than a metrics gap is the second-order damage. rank.mjs
scores an unreported delivery on the same decay curve as a rejected one, so silence does not leave
the ranking uninformed — it actively demotes. 35 documents sat in the hard-decay regime having never
drawn a single noise report, among them core/rules/read-before-act and core/rules/write-after- solve: the corpus was quietly sinking its own rules for want of a tool call that prose had been
asking for all along. The step performing worst was, predictably, the only one with nothing behind
it. outcome-gate.mjs now asks once at Stop and names the specific unreported ids — and whether
that moves the rate is an open measurement, not a result: this scar records the 14.4% and the
mechanism, and the next reading of npm run economics decides whether the mechanism deserved it.
8. Two documents describing the same handoff, and no code between them. Two loop skills in this system were designed as a pair: the read-only one ends with a ledger of findings and states that fixes belong to the other; the other opens by surveying the repo, and was documented as the thing the first one feeds. Both sentences were true, both files were internally consistent, and each was written by someone reading the other. The arrow between them existed only in those two sentences — the downstream implementation did not contain the upstream's name anywhere. Every run of the documented sequence re-derived the diagnosis by hand, and the expensive half — the findings the first pass had refuted, so the second would stop re-raising them — was discarded.
What is new here is why review misses it. Scars 1–5 are one document making a claim nothing backs, which a careful reader catches by asking "what enforces this?". This is two documents corroborating each other, and corroboration is what a reader normally treats as evidence. It survived four days and a skill-roster audit. The pair's shared design helped it hide: both carry a state file, a Stop hook, a ledger and a completion marker, so they look wired together in a way they are not. Sameness of shape reads as connection.
The check is mechanical and takes a second: grep the downstream implementation for the upstream's name. No hit means no handoff, whatever the prose on either end says.
How to apply
- For every convention, ask: what breaks if someone ignores this? If the answer is "nothing, until much later," it needs a mechanism.
- Prefer, in order: a pre-commit hook > a lint script run in CI > a script someone runs > a documented rule. Only drop a tier when the tier above is genuinely impossible.
- Pair declaration with enforcement: state the rule in prose (so the why survives) and wire it in config (so the what is guaranteed). Neither alone is sufficient — config without rationale gets deleted by whoever finds it inconvenient.
- Put lifecycle fields in machine-readable frontmatter, never in body prose, and lint them. Define transitions explicitly: who moves a status, and on what evidence.
- When you write a rule you cannot yet enforce, mark it as aspirational rather than pretending — an honestly-labeled unenforced rule is fine; one that looks operative and isn't will be trusted exactly when it matters.
- A documented handoff between two components is a claim about code, so check it in code: grep the downstream implementation for the upstream's name — its output path, its state file, its command. Two docs agreeing is not corroboration; each was written by someone reading the other. Be more suspicious when the two components share a design, because shape is what makes an absent wire look present.
- Having wired the mechanism, ask the second question: what does it retain? A gate that fires and forgets is still enforcement, and it is not evidence. Before claiming the system measures something, open the file the measurement would be read from — if there is no such file, you are describing a capability, not a result. Instrument when the historical question is worth answering, not merely because the mechanism happens to run; and when you name a gate, name the surface it actually runs on, never the one it sounds like it should.
This rule can be wrong
A hypothesis with 8 confirmations, not a law. If an agent applies it and still fails, that is recorded against the rule. Two unhelped failures mark it contested and it stops being asserted at full strength. A knowledge base that cannot demote its own claims only grows.