Do not re-derive what is already confirmed
- Evidence
- 3 recorded incidents
- Derived from
- 3 narrated failures
- Status
- active
- In force since
- Retrieved
- 10 times by an agent
- Outcome
- untested· no recorded outcome yet
3 independent failures forced this. Each is narrated below rather than summarised away: the incident is what makes the rule credible to the next person, and to an agent deciding whether to apply it.
Why this rule exists
- Rule3 scars
Do not re-derive what is already confirmed
- Promotion
A mechanism that keeps recurring has proved that writing it down didn’t prevent it. That is the promotion test: prose to rule, and where possible rule to runnable check.
- Incidents3 narrated
Each kept in full below, with what actually happened.
- Original events
The records behind these incidents are private and never rendered here. An agent running locally traces one with
scar_evidence; nothing on this site can.
The incidents behind it
- 01
The same confirmed root cause was investigated across three separate sessions. Each one re-read the same source files, reached the same conclusion, and produced no new information. The cause had been documented, with file and line references, after the first s…
- 02
The inverse, and the reason the status distinction exists: a session read an entry's *full* contents including a long investigation narrative, when the entry's fix section contained ready-to-apply steps. The relevant content was four lines; the session read fo…
- 03(2026-08-31, the expensive one)
A session was asked, in one plain sentence, to research harness and agent-memory improvements. It answered by spawning a background research agent, which spawned two more of its own — three agents, ~106k tokens in the parent alone, none of it announced before …
Summarised. The full narrative for each is in the rule below.
The rule
When an entry documents a confirmed root cause, that documentation is authoritative. Read its summary block first, check its status, and act on it. Do not re-investigate the codebase to re-derive a cause someone already proved.
Read strategy, by status:
| Status | What to do |
|---|---|
| Fix designed, not applied | Read the apply/fix section only. Implement directly. |
| Root cause confirmed | Read the cause section only. Do not re-investigate. |
| Cause unknown / unconfirmed | Full investigation warranted. Read the entry, then the codebase. |
The same rule governs scale, not just re-reading. Match the cost of the approach to what the task actually needs. Maximum thoroughness is a choice that has to be justified by the request in front of you — and when it spends someone else's metered budget, it has to be agreed before it runs, not reported after.
The scar
Incident 1. The same confirmed root cause was investigated across three separate sessions. Each one re-read the same source files, reached the same conclusion, and produced no new information. The cause had been documented, with file and line references, after the first session. Nothing in the workflow told later sessions to trust it — so each one "verified" it from scratch, at full cost.
Incident 2. The inverse, and the reason the status distinction exists: a session read an entry's full contents including a long investigation narrative, when the entry's fix section contained ready-to-apply steps. The relevant content was four lines; the session read four hundred. Reading everything is not the safe default — it's just a slower way to be wrong about what matters.
Incident 3 (2026-08-31, the expensive one). A session was asked, in one plain sentence, to
research harness and agent-memory improvements. It answered by spawning a background research
agent, which spawned two more of its own — three agents, ~106k tokens in the parent alone, none of
it announced before it ran. The person on the other end was mid-shift against a real metered budget
shared across the whole day's work, had produced no shippable work yet in that shift, and had
already joked about cost earlier in the same session. That joke was the check-in signal, and it was
read as banter rather than as a budget statement. Separately the same day, a /workflow run had
fanned an ordinary pre-commit review into ~30 agent invocations; the day's usage came out 82%
subagent-heavy, 48% from workflow-subagent alone, and the shift's budget was gone before any
product work landed. The tool being enabled answered "am I allowed to," and got read as "should I."
Orchestration cost is front-loaded and invisible: it commits in full at launch, so there is no
midpoint where the person paying can see the trajectory and say no. The research itself was
useful — the scale was never asked for.
How to apply
- Every entry gets a summary block immediately after its frontmatter: cause, fix, files, status. Four lines maximum. Write it for the session that arrives six weeks later.
- Check status before reading the body, and let it dictate how much you read.
- A session that investigates without resolving must increment the investigation counter and add a dated addendum — never leave an entry looking untouched after spending a session on it.
- If you genuinely believe a documented cause is wrong, that's a legitimate override — but say so explicitly, prove it, and update the entry. Silent re-derivation is the failure mode; a deliberate, evidenced correction is not.
- Before spawning more than one subagent, say what you are about to launch and roughly what it costs, and let the person answer. An enabled capability is permission, never instruction — re-derive the sizing decision from the actual task every time, no matter what session-wide state says. If a task sits between "one careful pass" and "worth the expensive tool," that ambiguity is itself the signal to ask. Any cost remark, including a joke, is a standing instruction to bias cheaper and to stop in-flight agents that are not earning their keep.
This rule can be wrong
A hypothesis with 3 confirmations, not a law. If an agent applies it and still fails, that is recorded against the rule. Two unhelped failures mark it contested and it stops being asserted at full strength. A knowledge base that cannot demote its own claims only grows.