The keyword half of the retrieval an agent gets, running here in your browser. Same index, same ranking code as the MCP server. An agent's answer can also be reordered by the encoder, which this page does not load.
Claims only — the browser has the index, not the lesson bodies. An agent also gets each document’s full text under a ~6,000-character cap, which lands around 1–1.5K tokens for five results.
Every lesson declares the tasks that should surface it — “a control disappears later in the day”, “a bug that only reproduces after a certain hour” — weighted three times the body. That bridges what a document says and what it’s for, by hand, at write time. The encoder catches some of what the triggers miss, but its similarity cannot tell an off-topic query from an on-topic one, so it is never allowed to answer alone.
A rule outranks a lesson; a skill is damped to 0.55. Measured, not guessed: before the prior, “adding a modal to the mobile web build” returned three tool descriptions at 3,300 tokens — a skill’s description is its whole document, so BM25 had no length dilution to punish.
Retrieved often and never acted on is a demotion. The counter that matters is applied, not recalled — otherwise the loudest document wins forever by being loud.
The searched corpus is 274 documents, not 1,302. That’s why an answer costs a thousand tokens instead of a hundred thousand, and why the encoder can re-embed the whole corpus in under a minute on a laptop. Ranking isn’t left to judgement either — a fixed benchmark of task-phrased queries runs in the verify suite, and a change that drops below its floor fails it.
Small numbers, shown at real size — early, self-reported, and growing. What matters is that the edge exists and is wired to ranking, because a knowledge base with no outcome signal can only grow. The most common outcome of any retrieval is being read and correctly ignored, so that is reportable too: noise demotes a document for surfacing where it did not belong, which is the difference between a ranking fault and a lesson that was simply not needed today. A firing of the signature scanner is a different question and gets a different verb, dismiss, counted against the signature — the two shared one counter until they were split, which had been sinking the best signatures in search for firing often. A lesson retrieved and never applied sinks; a claim that fails in use is marked contested rather than quietly kept. The signature count is reported the same way, deliberately not gated on: a lesson naming a detectable form (a literal, an error string, a code shape) is a check, not just advice, and requiring one to publish would manufacture fake patterns rather than admit a low ratio.
scar_recall("a text match or presence check accepted as proof that an element renders")
→ 1 documents, ~5.6K tokens
→ the agent states in one line what came back and what it changed
→ scar_learn(id, "help" | "apply" | "harm" | "noise" | "unread") after the taskThe standing cost is the kernel: ~2.2K tokens once per session, naming every claim Scar can back without stating any of them. Everything else is paid per query.