Context a hook injects above the harness size limit arrives as a preview, and a check keyed on the hook being wired reports it delivered
An agent harness caps how much text a hook may inject into the model's context. Output above the cap is not rejected and raises no error: the harness writes it to a file and injects a short preview plus the file path in its place. A hook that emits a document just over the cap therefore exits cleanly, its budget gate measures the emitted text and passes, and a fallback that asks 'is the hook wired?' concludes the document is already in context. Every layer that checks the sender's side agrees; only the receiver's side shows the loss.
Resurfaces when
- writing a session-start hook that injects a document into the agent's context
- setting a token or size budget for text that a hook, plugin or tool result delivers into a model context
- adding a 'skip re-sending, it was already injected' guard to a tool
- pricing context spend from what a hook emitted rather than what the model received
- a standing digest grows by a few hundred bytes after new rules or lessons land
- an agent behaves as if it never saw rules that the startup hook is known to inject
- a transcript shows 'output too large' and a preview where injected context was expected
The lesson’s retrieval contract, weighted three times heavier than its body. A lesson that states when it applies doesn’t need a semantic search to find it.
The lesson
When a harness sits between your code and the model, "I emitted it" and "the model received it" are separate facts, and a size limit is where they part. Harnesses commonly handle oversized injected context gracefully — persist the full text to a file, inject a short preview and a path — which means an over-limit hook produces no error anywhere. The hook exits 0. A budget gate that measures the emitted text passes. A guard that avoids re-sending the document because "the hook is wired" tells the agent the document is already in context. Each check is correct about the side it looks at, and none of them looks at the receiving side.
The failure is quiet in a second way: a preview is the beginning of the document, so whatever was written first still arrives and the session looks roughly right. What disappears is everything after the cut — typically the dense middle and end, which in a digest is where the actual rules and index live.
So a budget for injected context has two ceilings, and only one is yours. Your own ceiling (how many tokens you are willing to spend) is a choice. The harness's delivery limit is a fact, and the working budget is the smaller of the two. A budget set above the delivery limit is not a generous budget; it is a gate that approves text the harness will not deliver.
The failure that taught it
A tooling repo injected a standing digest — operating rules followed by an index of lessons — at the start of every agent session, and priced each injection into a spend ledger. A token budget guarded the digest's size, with a ceiling comfortably above its current length. A tool that could deliver the digest on request declined to when the hook was installed, to avoid paying for it twice.
A diagnosis noticed its own session had received only a preview. Counting across stored transcripts showed the same in roughly four sessions in five, across more than a dozen projects the hook was installed in, for as far back as the transcripts went: the digest had grown past the harness's inline limit and never shrunk back below it. The first ~2KB — the usage instructions — always arrived, so every session looked configured. The rules section, the part the whole budget existed to protect, arrived in none of them. The ledger had meanwhile priced every load at the full size, so the system's measured spend was overstated by its largest standing line, and an experiment comparing sessions with and without the digest was really comparing a preview against nothing.
No test had failed, because every test checked the hook's output, the budget's arithmetic, or the wiring — never what a real session's context contained.
How to apply it
- Before shipping a hook that injects context, find the harness's delivery limit and put it in
the budget. The ceiling is
min(your budget, harness limit); encode the harness limit as its own named constant so raising your budget cannot silently cross it. - Put what must arrive first. If truncation or preview is ever possible, order the document so the load-bearing part survives a cut — rules before indexes, instructions before examples.
- A "don't re-send, it's already there" guard must key on delivery, not on configuration. "The hook is installed" is evidence the hook ran, not that its text arrived. If the emitted size is over the limit, the guard should send.
- Price what was received, not what was emitted. A ledger fed by the sender's byte count inherits every silent loss downstream of the sender.
- Verify from the receiving end at least once. Open a real session transcript and look for the injected text in full. It is the only check that sees what the model saw, and it takes a minute.
- Generalises to any injected context with a cap: tool results the harness truncates or compresses, proxy layers that rewrite large payloads, system prompts assembled from files that grow.
Where this claim comes from
- Lesson
Context a hook injects above the harness size limit arrives as a preview, and a check keyed on the hook being wired reports it delivered
- Distillationmechanism stated
Written from the mechanism, not the incident — which is what lets it transfer to code sharing nothing with the original.
- Scar1 occurrence
One recorded failure — weaker evidence, and ranked accordingly rather than presented as settled.
- Evidencenone recorded
No source records recorded — hand-written and migrated lessons predate the pipeline that captures them.
What happened when it was used
- 5
- retrieved
- 4
- acted on
- 1
- held up
- 0
- did not hold up
Applied 4 times. Outcomes move the ranking both ways, which is what makes this improve rather than just grow.