The full argument, in the order it was written. The home page shows a gate firing; this is what sits behind it.
Repository search, git and docs all answer the same kind of question. This answers the other kind.
What every tool already gives you
This function does X.
It was last changed in commit a1f59f2.
It is called from four places.
Its tests pass.
Here is how it is written.
Repository search, git history and documentation all answer this well.
What nothing gives you
We tried X.
It failed because Y.
The failure was verified — 5 times.
The safe approach is Z.
Do not repeat X.
Nothing in the toolchain produces this, because nothing in the toolchain was there when it hurt.
Agents write nearly all of this, so the design assumes their output is wrong until something outside the writer says otherwise. Each control is a mechanism on disk — a gate, a script, a committed file — not a policy.
An agent’s feedback never edits a lesson. It lands in an inbox that consolidates repeats, gets reproduced before it is believed, and is excluded from the searchable corpus entirely.
One recorded outcome nudges a lesson’s rank a bounded step; it takes two confirmations to change what the claim asserts. Nothing the writer can simply say about its own work is an input.
A lesson that misleads sinks in ranking and stays on disk with its provenance. Wrong for this task is evidence, not garbage — and a sunk lesson can climb back by being useful.
The always-loaded kernel and the per-call payload each have a budget script in the verify suite. Raising a ceiling is a diff to that script — a reviewed decision, never drift.
Every token Scar actually put into a context is priced. The measured half stops at break-even, because the cost a lesson avoided is unobservable — anything called a saving is labelled a projection.
Reverted attempts are written down with the incident that forced the reversal, so the next session argues with the record instead of re-running the experiment. Several of this page’s own layouts died that way.
Agents don’t browse this site — they call a local process over MCP. There’s deliberately no tool that searches raw records: an agent that can load 2.2M tokens of narrative will try to.
brain/ — markdown + YAML frontmatter, private, in git
scripts/ — frontmatter is the database. One write path, so the sanitizer is unavoidable
tools/distill — IDF clustering (8.9M chars → 181k of briefs), then an LLM per cluster, behind a security quarantine and a denylist gate
core/rules — promoted when prose has demonstrably stopped preventing it
tools/recall — BM25 decides whether to answer: triggers weighted 3× plus an intent bonus, corpus-derived stopwords, a coverage floor that returns nothing rather than noise. A 23 MB fine-tuned encoder then reorders and extends an answer keyword recall already gave, and never answers alone. Benchmarked in the verify gate
tools/mcp-server — stdio, local. scar_context · scar_recall · scar_evidence · scar_learn · scar_hypothesis · scar_feedback
scar_learn → .recall/feedback.json — the one derived file that can’t be regenerated
Markdown and YAML frontmatter, in git. Two runtime dependencies: the MCP SDK, because the protocol is the point, and a pinned WASM runtime for the encoder.
BM25 from scratch, plus a 23 MB MiniLM encoder fine-tuned on task → lesson pairs. Keyword recall decides whether to answer; the encoder only reorders and extends that answer. No vector database, no network call, and it fails open to keywords.
Next.js 16 and Tailwind 4, static at build. It reads the corpus directly, and the private records sit structurally outside the deploy artifact.
82 skills and tools, and harness hooks that deny a step before the agent takes it, in Claude Code and Codex. Around them: a denylist sanitizer and a frontmatter linter on the pre-commit hook, a published-output leak check on every site build, and a token budget plus a retrieval benchmark in the verify suite.
Scar trains small models on its own outcome data. Each has to beat the rule it would replace, on data it never saw, before it changes anything. Until then the rule keeps deciding.
Reorders and extends an answer keyword recall already gave, reaching lessons the task never names in their words. It never answers alone, and when it fails, recall falls back to keywords.
Would drop a recall result the agent is about to report as noise, before it costs context. It retrains on every upkeep run and is read only once a run promotes it.
Would decide whether a gate was right to fire on a case its rules cannot settle. Its last training could not beat shuffled labels, so the rules keep deciding.
Would flag a signature-scan firing as a likely false alarm, so a noisy signature gets tightened instead of ignored.
Would write the recall query a gate forces, whichever model is driving the session, instead of leaving the phrasing to the agent.
Would flag an agent’s report about Scar’s own behaviour as false before it reaches the ranking, and guards the other models’ labels the same way.
Would write a generic, client-free account of one private incident, so a lesson can be learned from it without the raw record ever leaving.
Same task, run twice. The agent is simulated; the knowledge isn't — the failure on the left is recorded here 5 times, and the retrieval on the right is a real query, run when this page was built.
Hide the extend-booking control once the reservation's last day has passed.
Search the repository for the booking model
Find the existing date comparison and reuse its shape
Compare today against the end date
new Date().toISOString().slice(0, 10)
Tests pass. Review passes. Ship.
Bug report: the control vanishes for some users
Cannot reproduce — it works every morning
The failure is time-dependent, which is why it survived review.
Rediscover the timezone boundary from scratch
Cost
scar_recall("…once the last day has passed")
1 ranked documents, ~954 tokens
Deriving today's date from a UTC timestamp rolls the day over early for anyone west of UTC
score 33.9 · 5 recorded failures · high confidence · matched on control, day, last
Agent states the recall in one line before planning
A silent consultation and a broken install look identical, so saying it is part of the loop.
Build the day key from local date components
The shared helper already existed and was simply not reached for.
scar_learn(id, "help" | "noise" | "dismiss")
The claim gains support; ranking follows the outcome.
Result
Simulated agent behaviour; real corpus, real ranking. The right-hand retrieval is scar_recall, run at build time against 325 distilled lessons and 9 rules, and its token cost is measured from the documents actually returned — not estimated from their titles. The task sentence deliberately shares no vocabulary with the lesson it finds.
Not an error log. Failure + evidence + resolution + recurrence + lesson — drop any one and it’s a note again. Below is a real one, read from the file it lives in.
On a long-lived codebase, the agent kept finding the code, reading the docs, checking the history — and still walking into a wall someone had already hit. The expensive knowledge was never in the repository. It was in sentences nobody writes down: we tried that before. That fix caused a regression last month.
So the question isn’t how much an agent can remember. It’s what did we learn the hard way, and how do we stop the agent learning it again — a ranking problem, not a storage one.
Nothing here is simply stored. An incident is observed → distilled → validated → enforced → evaluated, feeding back into itself. Every step below is a real artifact with a real count.
The agent writes the record, in the session where the surprise happened. Predictable work leaves nothing: a log of everything is a log of nothing.
The failure, the confirmed cause, the fix. An audit trail — no agent ever loads this layer.
An agent distils records sharing one mechanism into a single claim — through a gate it does not control, which refuses drafts rather than filling slots.
A lesson is a hypothesis. Using it is the experiment, and rank follows the result — in both directions. Self-assessment is not an input.
A mechanism that keeps recurring proves writing it down didn’t prevent it. It hardens into a rule, or better, a runnable check.
The agent describes the task in a sentence and gets the few claims that bear on it. Never the history behind them.
The knowledge changes the plan before the code is written — the only point at which it is cheap.
The agent reports back: helped, harmed, or noise — and dismiss for a scanner firing it ruled out, which is counted against the signature rather than the search ranking. This is what makes the system improve instead of just grow, and it feeds straight back to 01.
Agents write nearly all of this, so nothing here takes an agent’s word for its own work. A draft is refused unless it names a mechanism and at least two trigger phrases; refusals are recorded rather than retried. A signature that is prose instead of something you can grep for is rejected, because it claims a check and delivers advice.
Rank comes only from things the writer cannot assert: how many independent records shared the mechanism, whether the claim is checkable, and what happened when it was used. A lesson calling itself confidence: high off one record with no signature is capped — at write time, and again by a linter over everything already on disk. It can still climb, by being useful.
The corpus shrinks twice. Distillation drops the narrative and keeps the mechanism; ranking then keeps each claim and fetches the rest only when asked.
private engineering records — never published, never loaded
325 lessons and 9 rules, in full
loaded once per agent session
An empty context is maximally compressed and worth nothing. The target is the most signal that fits a real budget — naming 325 claims beats explaining twelve and dropping the rest.
Raw history is for investigation; distilled knowledge is for execution. The path back from a claim to its records stays open — compression that destroys provenance is just confident advice nobody can check.
Not authored best practice. Every rule traces back to the incidents that forced it, and says how many there were — one scar stays marked a guess until reality confirms it.
Ranked ×1.15 above lessons in recall — a rule has already earned its place.
"ABSENT" did not survive contact with the device. A code read concluded a feature was missing.
Two reported bugs were refuted, and in both the *report's* read was wrong, not the app. Both had the same failure shape: the feature was looked for in the place the report assumed it would be, rather than the place the reference implementation actually put it.
A hit counter caught a test aimed at the wrong target. A fault-injection harness was armed against one endpoint while the screen under test used a differently-named one.
A code-read verdict is a hypothesis, not a result — read the full chain, incident by incident.
A lesson that surfaces for every query is noise. These are real results from the pass that serves scar_recall — the scores and token counts are what an agent gets. Try the console yourself.
Task, described in a sentence
“a list renders duplicate keys after toggling a view mode”
Task, described in a sentence
“two clients disagree on a record count from the same endpoint”
A claim that gets used and fails has to be able to lose rank. Otherwise a knowledge base can only ever grow.
A new lesson starts neutral at 1.0, so learning something can’t bury everything learned next. Bounded both ways: evidence nudges the ranking, it never captures it.
Shown at real size — this is a prototype, not a usage graph. What matters is that the edge exists and is wired into ranking. Everything else in .recall/ regenerates on every build; this file can’t, because outcomes only happen once.
Honest scope. Every number on this site is computed from the corpus at build, not asserted. The system is used on real working sessions by the person who built it, and by nobody else yet, so this evidence is early, self-reported, and has no comparative benchmark behind it — 476 recorded outcomes so far, read from .recall/feedback.json. The source records are real work across six projects including client engagements, so they stay private and only the distilled layer is published. Read the claims as a small and growing sample, not as a result.
See per-document status in the console →Because the repository records what the code does, and none of its tools record what already went wrong. This adds that one category — validated experience, including approaches tried and abandoned, which leave no trace in the code.
The repository
tells you what the system does today
Git history
tells you what changed, and when
Documentation
tells you what someone intended at the time
Scar
tells you what the developer learned the hard way
It replaces none of them. Each source answers a question the others cannot, and the last column is the one this system adds.
The repository is private — it holds the engineering records Scar was distilled from, and some of those are client work. Get in touch and I’ll grant access; once you have it, one prompt to your agent sets everything up. Your records stay on your machine. Nothing is uploaded.
Built by Joshua Diniega · jdiniega202@gmail.com