Skills are procedures an agent runs; tools are code that enforces a rule without relying on anyone's memory. Every one of them exists because prose alone failed first.
/scar-address-feedbackRead Scar's usage-feedback inbox and act on what agents reported — turning observed retrieval failures, misses, and noise into actual engineering changes. Use when the user says "address feedback", "review Scar feedback" (or "review brain feedback"), or asks what agents have reported about Scar's usefulness.
Agents that used Scar during real work recorded how well it performed (/scar-feedback).
This skill turns those reports into changes. It is the read side of that loop.
cd ~/Desktop/Unified-Brain/brain && npm run feedback
That lists open entries. Read each one in full before deciding anything:
npm run feedback -- --show <file>
If the inbox is empty, say so and stop — do not go looking for work.
Take regressed entries first, and treat them differently. The listing marks them ! and
sorts them above everything else. A regressed report was addressed once and then observed again:
the recurrence count (×N) is how many independent sessions have now hit it, and the "Recurred
after being addressed" section names what happened after the fix that was supposed to work. So
the question is not "how do I fix this" — it is "why did the previous fix not hold", and the
previous fix is written in that same file under "Addressed". Re-applying the same class of change
harder is the failure mode here; core/rules/enforce-dont-declare exists because the answer is
usually that the fix was advisory and needed to be a mechanism.
Feedback is an observation, not a specification. Sort each entry into one of:
tools/recall/. Highest value, because it affects every future query./scar-distill), not by touching the engine.core/rules/enforce-dont-declare. The
answer is a mechanism, not a better-written rule.core/AGENTS.md, because that is the only file an adopting project loads.State the classification for each entry before proposing work. Two reports that read similarly often belong to different classes and need opposite fixes. One report can carry several claims in different classes — classify each claim, not the report.
A report is a hypothesis (core/rules/verification). It is written from memory at the end of a
long session, and this inbox has been wrong about its own mechanism at least six times — see
DECISIONS.md, and note that two separate entries there each call themselves "the fourth", eleven
days apart, because the count itself kept being re-derived from memory.
Replay the log; do not reproduce the paraphrase. The report's description of what it asked is itself a memory. The exact strings are on disk:
grep '"at":"YYYY-MM-DD' .recall/query-log.jsonl # the report's date — verbatim query, variants, returned, ids
Replay those strings. This is not a refinement, it is the difference between the right answer and
the wrong one: on 2026-09-19 a reconstructed paraphrase returned 0 results and appeared to refute
a noise claim that the logged originals confirm. DECISIONS.md already records that a paraphrase
"would have sent it the wrong way", and tools/recall/benchmark.json replays verbatim from this log
for the same reason. This skill was the last place still telling you to paraphrase.
Then cross-check, in this order:
recalls: against the rows in the window, and its noise
claims against the ids on each row. A claim about a document that does not appear in any ids
array did not happen.stale, not a ranking bug. Say which..recall/investigation-log.jsonl holds the theories the session
recorded and their order. A report saying "I checked X before theorising" is checkable there.State the measured count before proposing any work. If it disagrees with the report, that disagreement IS the finding — and it is usually worth more than the fix the report asked for.
Use npm run recall -- for replays, never the scar_recall MCP tool: the CLI passes
countRecall: false, so it writes no usage signal, while the MCP path records one usage event per
returned document and would corrupt the ranking you are trying to measure.
Any change to ranking, the tokenizer, or the index must keep the benchmark at its floor:
npm run verify # frontmatter-lint + sanitize-check + token-budget + recall-benchmark
Retrieval changes cannot be eyeballed. If the benchmark drops, fix the change — lowering the threshold is a deliberate, diffed decision, never a way to make the gate pass. If a fix is correct but the benchmark disagrees, that is a benchmark gap: add the query, in a separate commit, and say you did.
Respect the budgets: core/AGENTS.md has a hard token ceiling, and adding to it means cutting
from it. Rationale belongs in core/rules/, which loads on demand.
npm run feedback -- --close <file> "what changed, concretely" \
--reproduced confirmed|partly|refuted|not-retrieval \
--repro "<the query you actually replayed>"
--reproduced is required and the close is refused without it. Step 3 used to be advisory:
nothing downstream ever asked for its output, so a close that reproduced the report and a close
that believed it were identical on disk and to every gate in the repo — the exact theater
core/rules/enforce-dont-declare names. It is now declared at the call, like retests:
confirmed — replayed it, the claim holds as stated.partly — the failure is real, the report is wrong about its size or its cause.
This is the most common honest answer.refuted — replayed it and it does not hold. Close saying so. A deliberate rejection is a
resolution and often the most valuable one you can file.not-retrieval — no recall query settles this. Reports about a gate, a hook, a skill or a
document are judged by reading code. This exists so the field never has to be filled with a lie.For confirmed/refuted, --repro is checked against .recall/query-log.jsonl — so run the
query before you claim it. The check fails open (absent, empty or unreadable log → accepted):
the log is gitignored and local-only, and an unreachable brain must never become an uncloseable
inbox. It is corroboration, not proof; the verdict is still your declaration.
The verdict is stamped into the report's frontmatter, which is the point — it makes "how often is this inbox wrong about itself" a number you can count rather than one re-derived from memory.
A resolution is required. An inbox that can be emptied without naming a diff decays into a to-do
list nobody trusts — the same theater core/rules/enforce-dont-declare is about. If the right
answer was "no change", close it saying that and why; a deliberate rejection is a resolution, an
ignored entry is not.
One short summary: what came in, what class each was, what changed, what you deliberately did not do and why. If you found a cause the report itself missed, say so — that is the most useful thing you can hand back.
Finally, keep the record honest: if a decision changed, update Locked decisions in
~/Desktop/Unified-Brain/AGENTS.md in the same session. That file existing is worthless if it
goes stale.
/scar-adoptAdopt Scar conventions into the current project — one-line import plus a project-specific entry folder. Use when starting a new project or bringing an existing one under Scar.
Installs Scar's conventions into whatever project the current session is in. Never overwrites existing project docs — merges.
visibility: all run inside Scar repo, not here. So the one check
that matters happens now, not at commit time:
git remote -v — if there is a GitHub/GitLab remote, check visibility:
gh repo view --json visibility -q .visibility (or ask the user if gh is unavailable)..gitignore it, or point it at a private location), or
adopt with the folder tracked and the user accepting that every entry is public the
moment it's written. Proceed only after the user picks one.0.5. Detect an incumbent knowledge system before adding a second one. Adoption's default
shape — append an import line, create an entry folder — is correct for a repo with no
existing convention and wrong for a repo that already has one. Look for: a project
knowledge folder (memory-bank/, .ai/, docs/brain/, notes/, or any project-named
vault directory), or an
existing AGENTS.md/CLAUDE.md that already carries its own numbered bootstrap ("read
these files first", "before responding to any request", a rules list).
If one exists, two bootstraps now compete for the same slot — session start — and the
generic one loses. Observed (2026-08-12): a session in a project carrying its own
pre-existing knowledge vault read Scar's imported core/AGENTS.md verbatim and called
scar_context/scar_recall zero times across a multi-step task, because the incumbent's
Rule 0 was concrete file operations ("read these exact files, in this order") while the
brain's was a judgement call ("call a tool with a sentence describing what you're about to
do"). Concrete instructions win attention over generic ones every time. The step-4 hook
fixes this for the standing kernel only; per-task recall still depends on the agent, so a
losing bootstrap means it never fires.
So do not append a parallel bootstrap. Instead:
core/rules/ and, for every pair that gives opposed instructions for the same
action, add one explicit precedence line to the project's AGENTS.md. A real instance:
an incumbent "write an entry after any meaningful problem solved" against
core/rules/write-after-solve's "only on surprise and mechanism — most solved problems
leave nothing." Same project, same action, opposed policies, and neither file said which
governs. Unresolved, the agent picks one per session at random.docs/brain/ beside it
(skip step 3). Two stores for the same artifact is how one goes stale.Check whether the current project already has an AGENTS.md or CLAUDE.md.
AGENTS.md with the import line below plus a short
project-specific stub (name, one-line description, area taxonomy if the project wants a
closed enum — see core/SCHEMA.md).The import line, always exactly this:
@~/Desktop/Unified-Brain/brain/core/AGENTS.md
Create the project's own entry folder if it doesn't exist — ask the user where (a common
default is docs/brain/ inside the project, matching the shape of core/templates/'s output
dirs: no bugs/, logs/, etc. subfolders are created speculatively; they appear as the first
entry of that type is written).
Install the gates: npm run install-gates -- <project-root> (run from brain/). One
command, additive, idempotent, safe to re-run after any brain upgrade. It writes all six
components — kernel injection on SessionStart, the recall / hypothesis / capture /
destructive-git gates, and the signature scanner — into both .claude/settings.json and
.codex/hooks.json, merging with whatever hooks each surface already has rather than replacing
them. --dry-run
shows the diff first; --only <keys> installs a subset.
This step used to be six sub-steps of hand-merged JSON, and that is exactly why it is now one
command. Measured across seven projects on 2026-08-16, on the machine of the person who wrote
the gates: the signature scanner was wired in one, destructive-git-gate in one, and
every project that carried the hypothesis gate was missing its mcp__scar__scar_context
matcher — so the arming fix that gate shipped for was live in none of them. A six-component,
four-event manual merge has a per-step failure rate, every missing component fails open by
design, and nothing in a session reports what is absent. The checklist was the defect; do not
reintroduce one. If a component needs changing, edit tools/hooks/lib/wiring.mjs — the single
definition that both the installer and npm run hooks-coverage read.
What each gate is for, and the incident behind it, is in tools/hooks/README.md and in each
script's own header. Do not restate them here; a second copy of that reasoning is the thing this
step just finished deleting.
Then confirm with npm run hooks-coverage, which lists every adopting project and what it is
configured on each harness. Restart Claude if its settings directory was newly created. For
Codex, restart or open /hooks and review/trust new or changed project hooks; configuration
coverage cannot prove Codex runtime trust, and adoption must never bypass it.
4.5. Confirm the tools this project is now told to call actually exist: npm run mcp-coverage
(run from brain/), and npm run install-mcp if it reports a gap. Step 2 just imported a file
whose first instruction is to call scar_context, and step 4 wired gates that will deny a
turn until scar_recall is called — both of which assume Scar MCP server is registered
with whatever harness this project gets opened in. That assumption had never been checked, and
on 2026-09-20 it was false: a Codex session in a freshly adopted project read the instructions,
found no scar_* tool, and could not comply. The registration lived in one hand-written entry
in ~/.claude.json and nothing in the repo wrote or audited it. Adoption is precisely where
this recurs, because it is the moment a project starts being told to call these tools.
Per-machine, not per-project: a second adoption on the same machine finds this already done.
new-entry or write any actual entries as part
of adoption — adoption installs the convention, it doesn't manufacture content.Never adopt into a public repo without the step-0 warning and an explicit user choice about where entries land. A silent adoption into a public repo turns every future entry into a publication.
Never copy core/templates/, core/rules/, or core/skills/ into the target project. Always
reference by path (the whole point is one source of truth).
Never overwrite an existing AGENTS.md/CLAUDE.md without showing the diff first.
Never skip step 4 because "the AGENTS.md import already says to do this." It does, and that was already shown not to be enough on its own — prose an agent reads once is not a mechanism.
Never hand-merge the hooks, even for "just one component". That is what step 4 replaced, and the
measured result was two of six components live in one project out of seven. Use
install-gates --only <key> for a subset.
Never install the recall gate without the fail-open behaviour and the denial cap intact. The gate's job is to make one tool call unavoidable, not to make editing conditional on a background service being up.
Never adopt alongside an incumbent knowledge system without doing step 0.5's reconciliation. Adding the import line next to a stronger, more specific bootstrap installs overhead, not a brain: it is loaded every session and consulted in none.
Never install the hypothesis gate armed by default or on a schedule — it must only
arm when a real scar_recall response flags the task debug-shaped, or it will nag on ordinary
editing and get "hardened away" within a session, taking the mechanism with it.
/scar-diagnoseDiagnose the current state of a software system — architecture, performance, efficiency, maintainability, reliability and AI/LLM components — as four analytical passes (systems engineer, performance analyst, AI-systems expert, software architect) over a MEASURED evidence pack, and LOOP (observe → understand → measure → diagnose → prioritize → recommend → verify) until every finding is confirmed by evidence, refuted, or explicitly marked needs-measurement, and a verify pass strikes nothing. Read-only — writes only a report under .claude/. Never edits source, never commits. Use when the user says "/scar-diagnose", "diagnose this system", "where is this slow or wasteful", "audit the architecture", "review our AI cost", "what should we improve first".
The analyst who is asked "what is actually going on in this system, and what would make it faster, simpler, cheaper and more reliable — in that order of certainty?" The job is to understand the system as it is, measure before judging, and produce findings a reader can verify without trusting the writer.
The persona — four passes, one evidence standard, the ranking rules — lives in PERSONA.md. Read it once at Phase 0, then work from the state file. It is written as checklists rather than an identity because naming a role does not improve a model's accuracy (Zheng et al. 2024), while separate cross-checked passes do (Wang et al. 2024), verbalized numeric confidence is well calibrated (Tian et al. 2023), and a small working context is more reliable than a large one (Chroma 2025; Liu et al. 2023). Every design choice below follows from one of those four results or from the measurement literature listed under Sources.
What this replaces. Phase 1 of /scar-harden (the read-only survey) generalized beyond
code smells to performance, architecture and AI cost — plus the hand-run sequence of
/toolkit-qa on each screen, npm run economics, and eyeballing the import tree. This is the
eleventh /scar-* command against a ceiling of nine. Joshua added it, on 2026-09-04, knowing
the count — the ceiling binds what an agent may grow, not what the owner may choose, and the
distinction is now carried by scripts/skill-roster.mjs rather than by a promise. Nothing is
retired to pay for it.
.claude/diagnose.local.md (state), .claude/diagnose-evidence.json (measurements) and
.claude/diagnose-report.md (the report), all git-excluded by arm. Fixes belong to
/scar-harden or to the user — and the ledger below is what /scar-harden imports, so
a row's status and evidence cells are an interface, not just notes. See Handoff at the
end.## Measurements. The
state file is the only permitted source of figures. A number remembered from reading code is
a hypothesis.confirmed without an evidence cell that is a measure key
(graph.cycles[0]), a file:line, or a captured command. "Code reading" is not evidence of
cost; it is evidence of mechanism, and the row's confidence must say so (≤ 89).--no-hook, and tell the user the loop is unenforced. The
backstop is a Claude Code Stop hook in .claude/settings.local.json; Codex and other
harnesses never read that file. Arming there without --no-hook writes a Claude
configuration file into the user's repo that nothing on their side runs. The script cannot
tell which harness called it, so the choice is yours to make. Say it plainly: the state file
and the exit condition still hold, but nothing stops a premature finish (feedback 2026-09-26).All optional:
--max N — pass cap (default 8).--no-hook — run without installing the Stop-hook backstop.--time "<cmd>" — repeatable; commands to time during measurement (build, test, lint, a
request). Timing runs each once — one sample, and the report must say so.State lives in .claude/diagnose.local.md in the target repo, outside working memory, so a
context reset cannot lose the ledger or the pass count. The Stop hook reads the same file. Every
phase reads it first and writes it last. The exit condition is printed at the top of the file
and re-checked at the bottom of every pass — deliberately at both ends, where a long context is
read most reliably.
0.1 Call scar_recall with the risk, not the task: "a diagnosis that reports a bottleneck
from reading code instead of measuring it", "a review that recommends a rewrite where a
cache would do". Name what came back to the user in one line, with an id, or say nothing
relevant came back.
0.2 git status first — never arm blind against a snapshot of the tree. Then:
node "$HOME/.claude/skills/scar-diagnose/diagnose.mjs" arm --max <N> [--scope <path>] [--no-hook]
Writes the state file (`active: true, pass: 0, verify_passes: 0, complete: false`), adds the
three `.claude/diagnose-*` paths to `.git/info/exclude`, and merges a `Stop` hook into
`.claude/settings.local.json` unless `--no-hook`. If that file is not valid JSON, `arm`
refuses to touch it and says so; the state file is still written and the in-skill loop
continues without the backstop.
0.3 Read PERSONA.md once.
node "$HOME/.claude/skills/scar-diagnose/diagnose.mjs" measure [--scope <path>] [--time "<cmd>"]...
This writes the evidence pack and prints a one-screen summary. It is deterministic and
zero-dependency, so it runs in any repo: inventory and largest files; layers and their
manifests; the import graph with fan-in, fan-out, cycles (Tarjan), orphans and hubs;
declared-but-never-imported dependencies; git hotspots (commits × log size) and
temporal coupling (files that change together across directories); duplicated blocks;
the AI inventory — model-call sites, model ids, cache_control uses, hooks per event,
MCP servers, and the token size of everything loaded into every session; and the wall time of
each --time command.
measure now drops everything under protected-paths.txt before reading a byte, and reports the
count as protected: — a skipped tree is NOT an empty one. Any grep you run alongside it must
exclude the same paths; the tool honours the list, an ad-hoc git grep does not.
Then read, in this order and no further: the root README/AGENTS, each layer's manifest scripts,
and the entrypoints the graph's hubs and orphans point at. Write ## System map in the state
file — components, entrypoints, boundaries — where every line names the file or measure key it
came from. Read pack sections by key (node -e / jq), never the whole file; the pack is
evidence to cite, not context to carry.
Three pack sections are leads, not verdicts, and the map must say which applies:
unimportedDependencies includes framework-implicit packages (react-dom under Next.js) and
CLI-only ones (referencedInScripts: true); orphans includes entrypoints the heuristic
missed; duplicates includes fixtures and generated code.
Pick the 3–5 flows that matter — the user-facing request, the build, the test gate, the AI call
path, the hot write path — and trace each end to end from its entrypoint. Under ## Flows,
one line per step naming the file and, where the pack or a timing gives it, the cost. Where a
step's cost is unknown, write ? — that ? becomes a needs-measurement row in Phase 4, not a
guess.
For every ? that a single command would settle, run it and paste the command and its output
head under ## Measurements. Typical instruments: --time on the build/test/lint commands;
time on a request; a count of hook subprocesses per tool call (ai.hooksPerEvent); token
size of standing context (ai.alwaysLoadedContext); git log --stat on a hotspot; a ledger or
usage report the system already keeps. One run is a sample. Say "n=1" next to any timing.
Do not run anything with side effects outside the repo (deploys, paid API calls beyond what the
system already makes) — mark those needs-measurement with the command the user can run.
Increment pass. Run the four passes from PERSONA.md separately, each over the system map,
the flows and the pack, each appending rows tagged with its lens:
| id | lens | sev | area | evidence | confidence | status |
Every row starts as hypothesis. Then the confirm step, one row at a time:
confirmed — the evidence cell holds a measure key, file:line, or captured command, and
the confidence is a number with the observation that would move it most.refuted — the evidence contradicts the hypothesis. Keep the row; a refuted hypothesis is the
reader's protection against the next reviewer re-raising it.needs-measurement — plausible mechanism, no number; the evidence cell names the exact
command or instrument. Confidence ≤ 59 by definition.Where two lenses produced conflicting rows about the same area, record both and add a line
under ## Pass log naming the disagreement. That line is the most useful thing in the report.
Every 3rd pass: call scar_hypothesis status. A row that moved confirmed → hypothesis
twice is a wrong diagnosis; stop reasoning about it and acquire an observation.
Severity = impact × confidence, per PERSONA.md's table. Effort (S/M/L) is a separate column in
the report, never a ranking input. needs-measurement rows are listed apart from the ranking,
ordered by how cheap the measurement is — a five-second command that would confirm a Critical
row is the first recommendation in the roadmap.
For each confirmed row, complete the nine fields from PERSONA.md's evidence standard. The
How to verify field is mandatory and runnable: the command or metric, before and after.
Group into quick wins (S), near-term (M), long-term (L — each with the fitness function that
proves the boundary held). Draft .claude/diagnose-report.md in the eight-section shape below.
Re-read every confirmed row against its evidence with the intent to strike it: open the
measure key or file:line, check the number is there and means what the row says. Strike any row
whose evidence does not support it (back to hypothesis or needs-measurement), and strike any
figure in the report that is not in the pack or ## Measurements. If anything was struck: the
ledger changed, so set verify_passes: 0 and return to Phase 4. If nothing was struck:
increment verify_passes.
The loop may end only when all hold, proven by the state file:
hypothesis rows.verify_passes >= 1 with no ledger edit after it..claude/diagnose-report.md, and every number in it traces to the pack
or ## Measurements.Then set complete: true, run disarm, and end the message with the literal line
<diagnose>COMPLETE</diagnose>. The Stop hook allows the stop only when the marker and
complete: true agree.
If the pass cap is hit first: set active: false, do not write the marker, and report
honestly what remains as hypothesis. Hitting the cap is information.
disarm leaves the state file in place, and /scar-harden's arm reads it. This is a code
path, not a suggestion — it was prose for four days and nothing consumed it, so the diagnosis was
re-derived from scratch every time. What it does with each status is why the confirm step in
Phase 4 matters more than it looks:
| status here | what happens over there |
|---|---|
confirmed |
seeded into the findings ledger as unverified — held open until re-reproduced in the current tree, never fixed on this ledger's word alone |
refuted |
quarantined under Do not re-raise. This is the highest-value row this skill writes: it costs the next pass a finding it would otherwise have spent |
hypothesis / needs-measurement |
listed as leads for the survey, never as work |
Two consequences for how the ledger is written:
file:line in the area or evidence cell wherever one exists. The import extracts
it, and a row without one arrives with nothing to open.confirmed row in the ledger to "look thorough". It arrives as work.Run disarm before /scar-harden arms — it refuses while this loop is still active, because
both install a Stop hook into the same .claude/settings.local.json and would block each
other's stops.
## System map. Concise.needs-measurement
list first where a cheap measurement gates a large decision.Afterwards, scar_learn on what recall returned in Phase 0. If the system surprised you with
a mechanism, /scar-lesson.
pass, verify_passes, hook_blocks in the state
file. The hook increments its own counter; the agent cannot reset it by forgetting.max_passes (default 8) bounds diagnosis cycles. max_blocks (default 6)
bounds how many times the hook refuses a stop — under the harness's hard limit of 8, so the
skill's message is the one the user sees.tools/hooks/.@imported standing document, co-changing files —
and confirms measure finds each, then drives the hook through allow / block / cap / garbage:node "$HOME/.claude/skills/scar-diagnose/diagnose.mjs" selftest
It is wired into npm run verify as diagnose-selftest.
## Measurements.<diagnose>COMPLETE</diagnose> with a hypothesis row in the ledger.needs-measurement row as a finding in the ranked list.arm in a repo that is not git-tracked.stop_hook_active, 8-block cap): shared with
/scar-harden; https://code.claude.com/docs/en/hooks/scar-distillDistill a batch of clustered incident records into transferable lessons, using the current session as the distiller instead of a paid API. Reads cluster briefs, writes gated lesson files, rebuilds the recall index. Use when there are undistilled clusters in .distill/briefs/.
Turns raw experience into knowledge Scar can recall. Stage 1 (clustering) is deterministic and free; this is stage 2, and you are the distiller — the briefs are on disk and you are already an LLM reading them, so no API budget is involved.
Distil in batches of 4–6 clusters. A batch is a natural unit of work: small enough to read carefully, large enough that the shared context makes the later ones faster.
Find what's left. cd <brain> && npm run distill:backlog. It reports each cluster as
distilled / refused / undistilled by hashing every brief's members with evidenceId() and
comparing against the union of every lesson's evidence: array, and it refuses to report at
all if it can't see that graph. Refusals come from .distill/refused.json — those were judged
undistillable and should not be retried without a reason. If .distill/briefs/ is empty or
stale, run npm run distill:cluster first.
Do not compare distilled-NNN lesson numbers against NNN.md brief numbers. Cluster IDs
are not stable across cluster.mjs reruns, so the two numbering schemes are unrelated —
distilled-162 and brief 162 were, on the real corpus, two different clusters. That
arithmetic is what distill:backlog exists to replace.
A cluster may also come back PARTIAL — some members cited by an existing lesson, some not. That is a prompt to look, not a task. The usual causes are genre residue (a session log swept in by vocabulary) or a member whose mechanism a sibling lesson already generalizes. Distilling one orphan member to zero out that column is exactly the manufactured lesson this skill's refusal path exists to prevent.
Read a batch of briefs in full. cat .distill/briefs/NNN.md. Read them, do not skim —
the whole value of this step is that a mechanism is visible across entries that each described
it as a one-off.
For each cluster, decide: is there one mechanism here? Refusing is a correct outcome and is recorded, not discarded. Refuse when the entries share vocabulary but not a mechanism (clustering artifacts — session logs bound by "morning/evening" are the classic case), when the mechanism is too tied to one framework version or data model to transfer, or when the entries record work done rather than a failure diagnosed. A refusal rate near zero means you are manufacturing lessons to fill slots.
Before writing apply, look for a signature — actively, not as an afterthought. A real
usage report (2026-08-12) measured Scar's own signature ratio at 2/89 (2%) and named it
the single highest-leverage gap: of several "useful" recalls in that session, exactly one
changed an outcome judgment would not have reached on its own, and it was the one lesson with a
grep-able signature — it named a literal (toISOString().slice(0, 10)) the agent could search
its own diff for and find the exact bug already shipped. Every other useful recall was
principle-shaped and only confirmed a decision already made. A lesson with a signature is a
check; one without is advice, and advice only works if the agent both notices and chooses
to apply it — which is the same gap enforce-dont-declare is about. So before defaulting
signature to absent, ask: is there a literal string, error message, function/config name, or
code shape in these source entries that names this mechanism? If yes, it belongs in
signature, not buried in apply's prose. This is prompting, not a requirement — the
signature field stays optional and ungated on purpose (see lesson.mjs: a threshold would
manufacture fake patterns, which is worse than an honest absence). Read the ratio each time
recall:build runs (it's printed in the build output) — it should visibly move as batches with
a real grep-able mechanism go through, not stay flat because the field is being skipped by
habit.
Write the batch as JSON and pipe it in.
cat <<'JSON' | node tools/distill/write.mjs
[ {...}, {...} ]
JSON
Draft shape — every field required unless refusing:
{ "cluster": 116, "distillable": true,
"title": "the rule, stated as a claim — not a topic",
"mechanism": "one sentence naming cause and effect",
"lesson": "2-4 sentences developing the rule AND its boundary — when it does not apply",
"failure": "what went wrong, generically told, 2-3 sentences",
"apply": "a check someone can run, not 'be careful about X'",
"signature": "OPTIONAL — a literal/error/log-line/code-shape you can grep for. Omit if none.",
"triggers": ["words someone would use describing the task they are starting"],
"tags": ["lowercase-hyphen"], "confidence": "high|medium|low" }
Refusal: { "cluster": 158, "distillable": false, "reason": "..." }
Rebuild and verify.
node scripts/sanitize-check.mjs && node scripts/frontmatter-lint.mjs && node tools/recall/build.mjs
Write from the mechanism, not the narrative. "A shared style helper set a fixed width with no max-width, so every consumer overflowed on narrow viewports" transfers. "The settings dialog was cut off on mobile" does not.
triggers is the field that decides whether the lesson is ever recalled. It is weighted 3× in
the index. Write the words someone would use describing the task they are about to start
("adding a modal that must work on narrow viewports"), not a description of the lesson itself
("modal width bug"). This is the whole reason keyword recall is sufficient here — the document
states when it should surface, so there is no semantic gap to close.
Every source cluster here is an incident record, so triggers default to diagnosis tense — write
at least one that isn't. A real usage report (2026-08-12) found a lesson (distilled-116) that
matched a session's failure exactly and still never surfaced, because every trigger presupposed
the bug already existed and two observations already disagreed ("two clients show different
counts", "record count parity mismatch") — none matched the moment the lesson was worth the most:
building the thing that would go on to cause it. A briefs corpus built from bug/decision records
inherits that tense by construction, which makes Scar quiet on greenfield work and loud on
debugging — backwards, since prevention is cheaper than diagnosis. Ask, per lesson: what was
someone doing in the moments before this went wrong — designing, building, adding a dependency
— and is that its own trigger, not just the symptom that later revealed it?
State the boundary. A lesson that is always true is not actionable. Say when it does not apply.
apply must be executable. "Grep for X and confirm Y" beats "be careful about X" — prose does
not prevent recurrence, which is the same reason lessons ripen into skills.
Write generically: never a company, product, client, repo, person, ticket ID, hostname, or path. Describe systems by shape — "a facility-reservation web app", "the legacy admin dashboard".
The gate is mechanical, not advisory. write.mjs runs the real denylist over the rendered file
before it touches disk; a hit sends the draft to .distill/rejected/ and exits non-zero. Do not
treat the instruction as a style preference — but also do not trust yourself over the gate. If a
draft is rejected, read .distill/rejected/NNN.md and rewrite; never weaken the denylist.
Entry slugs embed client names, so evidence is stored as opaque content hashes and resolved
locally through evidence-map.json. You do not need to do anything for this — it is automatic —
but it is why a lesson can cite its sources without naming them.
write.mjs refuses any cluster the security scan quarantined, and you should not try to work
around it. A distilled lesson is the most portable possible form of a security finding, which is
exactly what SANITIZATION.md forbids while the vulnerable code could still be live — "generic or
not".
Check the kernel actually improved: npm run recall -- "<a task the batch covers>". If the new
lessons do not surface for a plausible task description, the triggers were written as topic
labels rather than task descriptions — fix them and rebuild.
/scar-feedbackRecord how well Scar actually performed during this session — what it caught, what it missed, what wasted attention. Use at the end of a session where scar_recall/scar_context were consulted during real work, or when the user asks you to report on Scar's usefulness. Writes one dated entry into Scar's feedback inbox for later review.
You have just done real work with Scar consulted. This skill records whether that was worth it, so the system can be improved from evidence instead of from impressions.
.recall/feedback.json counts events per document — this lesson was retrieved, applied, helped.
That can answer "was this lesson good". It cannot answer "was Scar good", because every
judgement of the system as a whole is a ratio over queries: two thirds of results were noise,
zero of seven useful hits came from the project layer, one of seven named a grep-able signature.
The first time this was done, by hand, it produced four concrete engineering changes. That is the bar: this is a bug report about a tool, not a diary.
This section used to say the queries were not stored and the ratio had to come from whoever was
there. That stopped being true and the sentence outlived it. On 2026-09-19 a report written
that way stated recalls: 8 where the log holds 11, and said a lesson was returned on four
recalls where the log holds three — the fourth cited recall returned three different
documents. Every error ran in the direction that made the session look more disciplined, which is
what memory does at the end of a long session.
Read the numbers, do not recall them. Before writing anything:
cd ~/Desktop/Unified-Brain/brain
grep '"at":"YYYY-MM-DD' .recall/query-log.jsonl # every call: verbatim query, variants, returned, ids
That file has one row per recall() call. It gives you, exactly: how many calls you made, what you
actually asked (not what you remember asking), how many results each returned, and which document
ids came back. .recall/investigation-log.jsonl does the same for debugging theories and their
resolution. Quote these, and state the window you filtered on.
The highest-value field in a report is the queries that returned nothing, and nothing else in the system records it. Those name the corpus's real gaps. In the 2026-09-19 session the two empty queries were precisely the two problems that cost it hours. List them verbatim.
If a number you remember disagrees with the log, the log is right — and say so in the report, because a reporter who caught their own drift is telling the reader how much to trust the rest.
core/rules/write-after-solve
bans those as entries and the same bar applies here — the subject is Scar, not the work.One entry per session, at most — per finding. If a session produced two genuinely different
observations about Scar, file the second one too, with distinct: true. What this rule bans
is padding one finding into several files, not recording two real ones.
Never fold an unrelated second observation into an existing report to stay under the limit.
Consolidation means "the same failure, seen again"; on an already-addressed report it records that
a fix did not hold, so a second topic pushed in there manufactures a regression nobody observed —
and a false one sorts ahead of every real finding in the queue. The tool now asks you to declare
this (retests) rather than inferring it, so answer it honestly.
Answer these from what actually happened, not from memory of what should have happened:
paid-off — Scar changed the outcome. Name what it caught or settled.mixed — some results earned their attention, some did not.noise — results were returned and correctly ignored; attention spent for nothing.absent — nothing relevant came back for work Scar should have had something on.wrong — Scar was followed and made things worse.app/home/screen." Bad: "retrieval could be better."Report failures of Scar including ones you caused. The most valuable entry in the inbox so far recorded a rule that was read at session start and then broken hours later in the same session — that produced an enforcement mechanism. An entry that only flatters the system is worse than none, because it is counted as evidence.
The inbox is read by someone deciding what to fix next, and the thing that decision needs is how much a failure matters — which is a count, not a file. Five files reporting one recurring problem look like five small problems; one file with five recurrences looks like the biggest problem in the system, which is what it is.
So writeFeedback / scar_feedback refuses a first attempt whenever the inbox already holds
plausibly-related reports, and hands you the candidates. You then do one of two things:
consolidate_into: "<file>". Your observation is
appended, the recurrence count goes up, and if that report had already been addressed it flips
to regressed — a fix that did not hold, which is the highest-signal state this inbox has
and is completely destroyed by filing a fresh entry instead.distinct: true.The candidates are ranked, not judged. Do not treat the order as a verdict: the ranking was measured against the real inbox and was anti-correlated with actual duplication — the one known duplicate pair scored 0.42 while two unrelated reports scored 0.67, because these reports are long and all written in the same vocabulary. Below 8 reports it does not rank at all and simply shows you everything. Read the candidate summaries and decide yourself.
This happened for real: a report was filed and addressed, and the next session's report repeated its headline finding as a brand-new entry. The fact that mattered most — a fix had been made and had not worked — was recorded nowhere.
Prefer the MCP tool if it is available (its real name is mcp__scar__scar_feedback; if calling
it fails, load it first with ToolSearch using select:mcp__scar__scar_feedback).
Otherwise write the file directly — this always works:
cd ~/Desktop/Unified-Brain/brain && node -e '
import("./tools/feedback/inbox.mjs").then(({writeFeedback}) => {
const r = writeFeedback({
verdict: "mixed",
project: "<which project this session was in>",
recalls: 7,
useful: 2,
summary: "<one line: the finding itself, not the topic>",
detail: `<what happened, specifically>`,
wanted: `<what would have helped>`,
documents: ["career/lessons/distilled-015"],
// then one of: distinct: true (after reading the candidates)
})
console.log(r.ok ? "wrote " + r.path : JSON.stringify(r, null, 2))
})'
To consolidate instead, call recordRecurrence from the same module with the candidate's filename
and the identical entry object:
cd ~/Desktop/Unified-Brain/brain && node -e '
import("./tools/feedback/inbox.mjs").then(({recordRecurrence}) => {
const r = recordRecurrence("<file from the candidate list>", { /* same entry fields */ })
console.log(JSON.stringify(r)) // r.regressed === true means a fix did not hold — say so
})'
Then tell the user in one line that you recorded it and what the verdict was. They review the inbox by saying "address feedback".
The entry is committed to a private repo but it describes real tasks, so it can carry client
vocabulary. Write generically — name the mechanism, not the client, the ticket, or the
endpoint. scripts/sanitize-check.mjs scans this directory and will fail the pre-commit hook on a
denylist term, which is the intended outcome rather than an inconvenience.
/scar-hardenTake a vibe-coded / AI-generated codebase (or one scoped path of it) to production standard as a senior engineer would — lock current behaviour with characterization tests, then fix findings in severity order (secrets, authz, error handling, validation, runtime bugs, duplication, dead code, efficiency, structure) in small behaviour-preserving slices, and LOOP until the quality gate is green, every finding is closed or explicitly deferred, and two consecutive hostile review passes find nothing new. Edits source; NEVER commits. Use when the user says "/scar-harden", "make this production ready", "clean up this vibe-coded app", "harden this codebase", "bring this up to standard".
The senior engineer who inherits the prototype. The job is not to rewrite it — it is to make the existing behaviour safe, correct, efficient and maintainable without changing what it does, and to keep going until there is objectively nothing left to fix.
Everything below is grounded in what the research says actually goes wrong in AI-generated code, not in taste:
And in what has been known to work for decades on code nobody trusts:
What this replaces. Running /toolkit-code-review → hand-fixing → /toolkit-qa →
/toolkit-tester by hand, over and over. Those skills are read-only or single-pass by design.
This one edits, and it is the loop that sequences them. It calls the same checks. It is the tenth
/scar-* command against a ceiling of nine — Joshua added it, on 2026-09-02, knowing the
count. The ceiling binds what an agent may grow, not what the owner may choose; the two are
told apart in scripts/skill-roster.mjs rather than by a promise. Nothing was retired to pay
for it.
Pairs with /scar-diagnose. That skill is read-only and ends with a ledger of findings whose
evidence has already survived a verify pass; run it first on an unfamiliar system and this one's
arm imports the result (Phase 0.2b) instead of re-deriving it. Running it is optional — with no
diagnosis present, Phase 1 surveys from scratch exactly as before.
eslint-disable, @ts-ignore,
@ts-expect-error, # noqa, # type: ignore, --no-verify, skipped tests, widened any.
That is hiding the finding, not fixing it.Same convention as the toolkit skills (toolkit/CONFIG.md):
repoSlug, appRoots — detected fresh — used to bucket findings by layer.qualityGate — detected fresh — lint / typecheck / test commands per layer, read from each
layer's manifest scripts. Distinguish "missing script" from a real failure exactly as
/toolkit-qa does. A layer with no test script at all is itself a P1 finding.repoRules — detected fresh — a repo-rules / architecture doc if one exists; its rules are
admissible findings.<SLUG>_HARDEN_TIER — asked once, then written — internal | customer | regulated
(money, health, PII at scale). Sets the readiness bar (Phase 4). Default customer if the user
does not answer.All optional:
--max N — iteration cap (default 12).--tier internal|customer|regulated — overrides the remembered tier for this run.--no-hook — run the loop without installing the Stop hook (Phase 0.2).--ignore-diagnose — arm even though a /scar-diagnose loop is still armed here. Almost
always wrong: two Stop hooks refusing stops in one session fight each other.State lives outside the agent's memory in .claude/harden.local.md in the target repo, so a
context reset or a premature stop cannot silently reset the count or lose the ledger. The same
file is what the Stop hook reads. Every phase below reads it first and writes it last.
0.1 Call scar_recall with the risk, not the task: "a refactor pass changes behaviour that no
test was locking", "a quality loop that never terminates or churns on style". Name what
came back to the user in one line, including an id.
0.2 Run git status first (never arm blind against a snapshot of the tree), then from the
target repo root:
node "$HOME/.claude/skills/scar-harden/loop.mjs" arm --max <N> [--scope <paths>] [--no-hook] [--ignore-diagnose]
This writes the state file with `active: true, iteration: 0, complete: false`, adds it to
`.git/info/exclude` (never to the tracked `.gitignore`), and — unless `--no-hook` —
merges a `Stop` hook into `.claude/settings.local.json` pointing back at the same script.
Claude Code picks the hook up via its file watcher without a restart. If
`settings.local.json` exists and is not valid JSON, arm refuses to touch it and says so —
the state file is still written at that point (a partial apply, by design and reported);
continue with the in-skill loop only.
`arm` **refuses** if `.claude/diagnose.local.md` has `active: true` — a running `/scar-diagnose`
owns the same Stop-hook slot and the same settings file, and armed together the two loops
block each other's stops. Finish or `disarm` the diagnosis first; that is also what makes its
ledger importable.
0.2b Import the diagnosis, if there is one. arm reads .claude/diagnose.local.md and seeds
the findings ledger from it, writing ## Imported from /scar-diagnose into the state file.
What crosses, and what does not, is the whole point:
- **`confirmed` rows are seeded as `unverified`** — never `open`. A diagnose evidence cell may
be a measure key (`graph.cycles[0]`), which is evidence of a *mechanism*, not a reproduction,
and this ledger admits only findings with a repro. They still hold the loop open.
- **The diagnose severity is carried verbatim** as `D:critical`, `D:high`, … and is *not* mapped
onto P0–P3. Diagnose ranks impact × confidence; this skill ranks incident likelihood, and a
silent mapping would invent a precision neither ledger has. Triage assigns the P-level.
- **`hypothesis` and `needs-measurement` rows do not cross.** They are listed as leads for the
Phase 1 sweep — the cheapest places to look — and nothing more.
- **`refuted` rows are quarantined** under *Do not re-raise*. Raising one again needs new
evidence, not a second reading of the same code. This is the half of the handoff that saves
the most work.
- **Files edited since the diagnosis ran are flagged**, and a diagnosis that hit its pass cap
(`complete: false`) is imported with a warning that its rows never passed a verify pass.
A missing, corrupt or unparsable diagnose file is not an error — `arm` says so and Phase 1
surveys from scratch, exactly as before.
0.3 Record the baseline: run every qualityGate command once, capture exit codes and the first
30 lines of each failure, and write them under ## Baseline in the state file. A gate
that has never been run against this repo has proven nothing — this is the run that
proves it can fail.
Produce the findings ledger — a table in the state file, one row per finding:
| id | sev | layer | file:line | class | evidence | status |
evidence is the command that reproduces it or the exact line. A finding with neither is
inadmissible and is not written down. This is the anti-churn rule: taste cannot enter the ledger,
so taste cannot keep the loop alive.
1.0 — Triage the imported rows first, if arm seeded any. Every unverified row is somebody
else's conclusion about a tree that has since been edited. For each, in the current tree:
file:line and confirm the
thing is still there. For a measure key, that means re-deriving the number, not trusting it — a
cited cycle that a later refactor already broke is a fix applied to nothing.D:*), its layer, and
an evidence cell that reproduces, and set it open — or strike it to
refuted: <what the tree actually shows>, which is a real result and stays in the ledger.Never fix a row while it is unverified; the Stop hook will keep handing them back until none
are left. Read Do not re-raise before sweeping, and do not spend a finding on anything in it.
Then sweep as below for what the diagnosis did not cover — it was looking for cost and structure,
not for secrets, authz or swallowed errors, so P0 and P1 are almost entirely yours.
Exclude every path in protected-paths.txt from every sweep command. This is not advice: a
2026-09-17 run grepped the repo wholesale and pulled client-derived lines out of a private evidence
tier into the session, because the sweep below is written as git grep -- <repo> and nothing
narrowed it. Read that file first and pass the exclusions (git grep ... -- ':!<path>'). A grep
bypasses both controls the evidence tier has — recall never returns it, and agents are told not to
open it — so the sweep is the one place the boundary has to be re-stated.
Sweep in this order and stop each sub-sweep at the first 30 admissible findings — the loop will come back for the rest.
P0 — will cause an incident in the first 48 hours
git log --all --full-history -- '**/.env*', grep for
key-shaped strings (sk_live, AKIA, -----BEGIN, password=), trufflehog/gitleaks if
installed. A secret in history is P0 even if the file is now deleted.* with credentials, shell=True, eval.P1 — will corrupt data or hide a failure
except:, except Exception: pass, catch (e) {}, .catch(() => {})),
swallowed promise rejections, unawaited async.P2 — will break at runtime or under the next change
any at boundaries; nullable deref.SELECT * into memory; sync I/O on a request path; missing
index on a filtered column; component re-render loops; work recomputed per render or per
request that is invariant.P3 — costs every future reader
core/skills/toolkit/COMMENTS.md — diary, redundant, dead, stale — under the
repo's commentPolicy. This is where a trim is applied; stale comments first, since those lie.Readiness (tiered — Phase 4 decides which are required)
.env.example complete, README run instructions true.Write the ledger. Tell the user the count by severity in one line. Do not wait for approval — the loop is the deliverable — but if any P0 finding involves a live secret, say so immediately and first, because rotating it is not something this skill can do.
Repeat until the exit condition in Phase 3 holds or the cap is hit. One iteration:
iteration. Pick the single highest-severity open
finding (ties: the one whose fix unblocks the most others, then smallest)./toolkit-unit-tester shape;
for I/O the /toolkit-tester shape.core/rules/verification, is entirely
incidents of this shape.qualityGate — lint, typecheck, tests — not just the new test.
Green means every command exits 0 for a real reason. A new failure elsewhere is a new finding;
add it to the ledger with its evidence and do not proceed past it.fixed (or deferred: <reason> — only for a fix that needs a
decision the user has to make, e.g. rotating a key, changing a public contract), append one
line to ## Iteration log with what changed and which commands proved it, and reset
clean_review_passes to 0 because the tree changed.scar_hypothesis status and check the ledger for the same
finding reopened twice. A finding fixed twice is a symptom of a wrong diagnosis — stop
fixing it, write down the theory, and acquire an observation (run it, log the value) before
the third attempt.Never touch two findings in one iteration. Never skip step 3 because the change "looks safe" — the study's numbers are what "looks safe" produces at scale.
The loop may end only when all of the following are true, in this order, each proven by a command in the current tree:
qualityGate green on every layer (fresh run, not remembered).fixed or deferred with a reason a reader would accept — or refuted
with what the tree showed. Zero open and zero unverified: an imported row that was never
triaged is a finding nobody looked at, and leaving it uncounted would make the import a way
past this condition rather than an input to it./toolkit-code-review runs — correctness, security, debug leftovers, repo rules) finds
zero admissible new findings. Increment clean_review_passes.clean_review_passes >= 2, i.e. two consecutive clean passes with no edits between them.
One clean pass proves the reviewer was tired; two prove the tree.If any check fails: that is the next finding. Back to Phase 2. Do not lower a bar to pass it.
When all five hold: set complete: true in the state file, and end the message with the literal
line <harden>COMPLETE</harden>. The Stop hook allows the stop only when both are true.
If the cap is hit first: set active: false, do not write the completion marker, and report
honestly that the loop was cut off, with the open ledger. The cap exists so an impossible bar
cannot spend the user's budget forever; hitting it is information, not failure.
| Item | internal | customer | regulated |
|---|---|---|---|
| No secrets in tree or history | required | required | required |
| Server-side authz on every endpoint | required | required | required |
| Error handling + timeouts on external calls | required | required | required |
| Input validation at boundaries | advised | required | required |
| Tests on the 3–5 critical user paths | advised | required | required |
| CI runs the gate on push | advised | required | required |
| Global error boundary + error tracking | advised | required | required |
| Structured logs with request ids | — | advised | required |
| Rate limiting on auth + writes | — | advised | required |
| Backups + restore tested | — | advised | required |
| Audit log on data access | — | — | required |
"required" means the loop cannot converge without it. "advised" appears in the report as a recommendation, never as an open finding.
node "$HOME/.claude/skills/scar-harden/loop.mjs" disarm
removes the Stop hook from settings.local.json and leaves the state file for the report. Then
report, in this shape and no other:
Afterwards, scar_learn on what recall returned in Phase 0, and if the session surprised you
with a mechanism, /scar-lesson.
iteration, clean_review_passes and hook_blocks
live in the state file. The hook increments its own counter; the agent cannot reset it by
forgetting.max_iterations (default 12) bounds the agent's fix cycles. max_blocks
(default 6) bounds how many times the hook refuses a stop — deliberately under the harness's own
hard limit of 8 consecutive blocks, so the skill's message is the one the user sees, not the
harness's.tools/hooks/.git add, git commit, git push, git stash.<harden>COMPLETE</harden> with an open row in the ledger.arm inside a repo that is not git-tracked, or on a dirty working tree without telling
the user the diff will mix with theirs.stop_hook_active, 8-block cap, file-watched settings):
https://code.claude.com/docs/en/hooks/scar-lessonWrite a born-sanitized generic lesson from what the current session just learned, straight to career/lessons/. Refuses if the lesson can't be stated without a denylist term. This is the dual-write habit as one command.
Takes a durable, transferable lesson from the current session and writes it, already sanitized,
to brain/career/lessons/. The point is sanitizing while the details are fresh — not weeks later
from memory.
Identify the lesson: what generalizable insight came out of this session? Not "what did we do" — "what would I do differently, or check for, next time, in ANY project?"
Draft it using core/templates/lesson.md's shape (the lesson, the failure that taught it,
how to apply it) with zero specifics — see SANITIZATION.md's voice rules: systems are
"a facility-reservation platform" / "the mobile app", never named.
Write at least one trigger phrased for BEFORE the defect exists, not just for diagnosing it
after. A real usage report (2026-08-12) found a lesson (distilled-116) that matched a
session's failure exactly, was indexed and high-confidence, and still never surfaced — every
one of its triggers was diagnosis-shaped ("two clients show different counts", "record count
parity mismatch"), presupposing the bug already happened and two observations already
disagreed. None matched the moment the lesson was worth the most: "I am about to build a
filter control over an API I did not write." Most lessons here are written from an incident,
so triggers inherit the incident's tense by default — that makes Scar reliably quiet on
greenfield work and reliably loud on debugging, which is backwards, since prevention is
cheaper than diagnosis. Before finishing the triggers list, ask: what was someone doing in
the moments before this went wrong — building, designing, adding a dependency — and is that
phrased as its own trigger, not just the symptom that later revealed it?
Self-check before writing: read the draft back and ask, for every sentence, "could this identify the client, a person, or a live vulnerability?" If yes, generalize further or refuse.
If the lesson genuinely cannot be stated without a denylist term or without losing the point entirely, do not write it — say so and explain what's blocking it. A lesson that only makes sense with the specifics attached isn't a generic lesson yet.
Scaffold with the script — never hand-write the frontmatter. Run
node scripts/new-entry.mjs lesson "<title>" --dir career/lessons, then fill in the body.
The scaffold carries every field frontmatter-lint requires; frontmatter typed from memory
reliably drops one. Observed 2026-09-01: a lesson written with hand-rolled frontmatter omitted
promoted_to and left verify red until someone next ran it.
If the lesson derives from client work (anything in protected-paths.txt), add
visibility: private to its frontmatter — mandatory, see core/SCHEMA.md. It stays visible
internally for tracking but is excluded from every public render; only a later, explicit
/scar-publish decision can lift it.
If this is a recurrence of an existing lesson (same failure mode as one already filed), don't
create a duplicate — open the existing entry and increment recurrences. Search first
(grep -r "type: lesson" career/lessons/). Only flag checkable: true if the lesson
also carries a signature (a grep-able form) — recurrence alone is not the bar; see
core/SCHEMA.md. Do not invent a signature to earn the flag: one that fires on every
ordinary instance of a shape spends attention nobody asked for, and checkable: false is
the honest answer for a real lesson that has no grep-able form.
Finish the write — a lesson is not done when the file is saved. Several checks derive
their expectations from the corpus, so one new lesson invalidates all of them at once
(career/lessons/2026-09-01-checks-that-each-derive-an-expectation-from-one-shared-artifact-all-fail-on-the-next-write-to-it).
The tail is:
cd ~/Desktop/Unified-Brain/brain
node scripts/sanitize-check.mjs career/lessons/<file> # must be clean — run this INLINE, always
node scripts/frontmatter-lint.mjs career/lessons/<file> # scoped to the file, instant — INLINE too
npm run recall:build # the lesson is invisible until this runs
npm run verify:lesson # the INNER LOOP — reports ALL failures
npm run verify # ONCE, at the end — the actual gate
Loop on verify:lesson, finish on verify. The scoped run carries every step a lesson can
break and drops the advisory ones, so it is the thing to re-run after each fix; the full gate is
what you must see green before saying the tail is done. Never substitute one for the other in
either direction.
Then clear whatever it reports. Two are expected after a lesson lands and are the lesson-writer's job, not a bug:
repo-about / recall-benchmark prose drift — corpus counts and benchmark scores quoted
in README.md/AGENTS.md are now stale. These sit inside sentences that argue something, so
fix the sentence, not just the digits; nothing rewrites them for you, by design.signature-scan selftest — only if you set checkable: true. A new signature needs a
planted case in tools/signature-scan/selftest.mjs. Confirm it fails on a non-matching line
before believing it; a plant you did not watch fail proves nothing.7b. Where that tail runs depends on what the session is for.
sanitize-check to a background subagent and
carries straight on. verify is ~105s per run and is expected to come back red once a
lesson lands, so clearing it is not one wait — it is a fix-and-rerun loop of several, plus
prose edits in README.md/AGENTS.md, inside a context that was supposed to be building
something else. That is the 1-part-in-5 brain budget (core/AGENTS.md Rule 1b) spent on a
single lesson. sanitize-check still runs inline: it is instantaneous, and dispatching an
agent to carry an unsanitized file is exactly the wrong order. So does the scoped
frontmatter-lint: it is the only check that can tell a checkable: true lesson whose
signature can never fire (a double-quoted scalar with doubled backslashes reaches the scanner
verbatim) from one that has simply not fired yet. Where the hooks are installed the rebuild
hook already runs it on every entry written and hands back the errors; the inline line is
the same check for a session without them.This skill carries standing permission to dispatch that one agent, and invoking it is the
request. A harness rule of the form do not use subagents unless the user asks is already
satisfied: the user asked by running /scar-lesson, and the handoff is a step of the command
they invoked, not an optimisation you chose for them. Do not read that rule as a veto and
quietly run the tail inline instead — that spends the product session's context on brain
upkeep, which is the single cost this step exists to avoid. It is one agent, on the fixed
prompt below, for a bounded task; the size the fan-out before spawning caution is about
unbounded parallelism and does not bear on it.
Dispatch with the Agent tool, backgrounded (the default), tell the user in one line that the
tail is running detached, and keep working. Relay the agent's report when the notification
arrives — never predict it. One at a time: if a second lesson lands in the same session,
wait for the first agent, because two agents editing the same corpus counts in README.md
will clobber each other.
Check the corpus before you relay that report — not the message. A subagent that did the
work and said nothing, and one that did nothing and said nothing, are the same observation from
here, and the honest default on a status: completed with no findings is to call it done.
Observed 2026-09-16
(feedback/2026-09-16-a-backgrounded-subagent-dispatched-for-a-lesson-s-verify-tai.md): a
dispatched agent returned the single line "I'll stop polling now and wait for the background
task/monitor notification to arrive" — it had spawned a background child of its own and ended
waiting on it. recall:build had run and the lesson was indexed, so half the tail was real; the
planted case the prompt requires in tools/signature-scan/selftest.mjs was never added, and the
half-finished tail was one sentence from being relayed as done. So when the notification
arrives, before saying anything about it:
cd ~/Desktop/Unified-Brain/brain && npm run lesson-tail -- career/lessons/<file>
~1s of filesystem reads, and it answers what no report can: is the lesson indexed, was the index
built after the last corpus edit, is there a planted case for this signature, and is the
README.md/AGENTS.md prose trued up. Named deliberately outside background-gate's long-check
list so it runs in the foreground from a product session. It is not verify and never
substitutes for it — it proves the work happened, not that the gate is green.
A report with no verify line in it is a FAILED dispatch, not a quiet success. Treat it as
failed unless lesson-tail exits 0, and re-dispatch on the same prompt naming what is still
open. Same when a confident report and lesson-tail disagree: the filesystem wins.
If the harness refuses the spawn outright — no Agent tool in the session, or a hook
denies it — the fallback is not run it inline and say nothing. Say in one line that the
dispatch was blocked, then get the waiting off this session with one backgrounded Bash call:
cd ~/Desktop/Unified-Brain/brain && npm run recall:build && npm run verify:lesson
That buys the wait for the price of a notification and nothing more: when it comes back red,
the fix-and-rerun loop lands in this context after all. Say so, and offer to leave the lesson
unverified for a brain-repo session — an unbuilt index is a known state you can name, a
product session derailed into prose edits in README.md is not.
The prompt must carry every constraint that does not travel. A subagent inherits none of
this session's gates or kernel — enforcement binds the context it is installed in, and
delegation is a context boundary
(career/lessons/2026-08-28-a-session-scoped-gate-does-not-fire-inside-the-subagents-doing-the-work).
spawn-gate denies a spawn whose prompt carries no consult instruction. Use this, filled in:
In ~/Desktop/Unified-Brain/brain, finish the write of a new lesson: career/lessons/<file>.
It has already passed sanitize-check.
Run, in order: `npm run recall:build`, then `npm run verify:lesson`. It reports EVERY
failure rather than stopping at the first — read the whole summary before fixing anything,
because the failures almost always share one cause, this lesson landing.
Clear only these two expected classes, re-running `npm run verify:lesson` after each fix.
When it is green, run the FULL `npm run verify` once — that is the actual gate, and
the scoped run is never a substitute for it. The two classes:
- repo-about / recall-benchmark prose drift — corpus counts and benchmark scores quoted in
README.md and AGENTS.md are now stale. Those digits sit inside sentences that argue
something: fix the sentence, not just the number.
- signature-scan selftest, ONLY if the lesson has `checkable: true` — add a planted case to
tools/signature-scan/selftest.mjs and confirm it FAILS on a non-matching line before you
believe it. A plant you did not watch fail proves nothing.
Hard limits. Never lower a threshold, floor or budget to make a check pass — a red floor
after a new lesson is the gate working. Raising a token budget is a deliberate diff in
scripts/token-budget.mjs; if you think one is needed, stop and report it instead of doing
it. Do not commit, stage, or otherwise touch git. Do not edit the lesson's own claims.
Anything red outside the two classes above: stop and report it, do not fix it.
You ARE the background — run these commands in the foreground, here. Do not hand the tail
to a background child of your own and end your turn waiting on it; that returns a
`completed` status over work nobody did.
Before starting, call mcp__scar__scar_recall with a sentence naming what this could get
wrong, and name one returned id in your report. Then call mcp__scar__scar_learn ONCE with
the batch `reports` form, one entry per document you were handed — `noise` is the right and
most common answer. No gate will ask you for this: gates bind the session they are installed
in and none of them reach in here, so an unreported delivery decays that lesson's rank exactly
as hard as a rejected one. Scar tools arrive DEFERRED: run
ToolSearch(query: "select:mcp__scar__scar_recall,mcp__scar__scar_learn") first — an
InputValidationError means the schema is not loaded, not that Scar is absent. If the search
still returns no match for either name (a subagent can be handed a server that predates the
tools), do not stall on it: run `npm run recall -- "<sentence>"` from brain/ instead, and
put a line `SCAR-TOOLS: not loaded` in your report so the caller knows the delivery went
unreported rather than reading its absence as a pass.
Finish by running `npm run lesson-tail -- career/lessons/<file>` and pasting its output. It
is ~1s of filesystem reads and it is your caller's receipt that the tail landed — it is NOT a
substitute for the full `npm run verify` above, and if it disagrees with what you believe you
did, it is right and you are not finished.
Report back: the final verify line, the lesson-tail output, every file you changed and why,
and anything you stopped on. Your report MUST contain a line starting `VERIFY:` carrying
verify's own last line verbatim. A report without one is read as a failed dispatch and
re-run from scratch, however much work you actually did.
sanitize-denylist.txt in your head before drafting, not just after./scar-migrateInventory a project's existing AI/knowledge docs and migrate them onto Scar templates, showing the mapping before doing anything. Refuses to operate on any protected path.
Moves an existing project's ad-hoc docs onto the shared brain schema. Show-before-doing is not optional here — this command's whole value is that it never surprises the user with what moved or how it was rewritten.
Read brain/protected-paths.txt before doing anything. If the target directory resolves to
any path listed there, or to any path under one, refuse and say why: those are work/client
vaults that never migrate wholesale, by design. The only path out of a protected vault is
/scar-lesson, one born-sanitized lesson at a time.
This check is a config lookup rather than a hardcoded path so that this skill stays generic — the protected paths are private, the rule is public.
bug/ticket/decision/snippet/concept/log/lesson), what
frontmatter fields are missing and need inventing, and what's redundant or clearly stale
(flag, don't silently drop).new-entry semantics — don't
hand-roll frontmatter), file into projects/<name>/ (or leave in-project if the project is
public — ask which), then run frontmatter-lint and brain-index./scar-publishPromote ONE named artifact from private brain/ to public brain-public/, after sanitize-check and an explicit diff review. Never batches.
The second, stricter gate. Everything in career/lessons/ already passed /scar-lesson's
sanitization — this command re-checks anyway, because public is a one-way door.
visibility: private (client-work-derived — see
core/SCHEMA.md), STOP and say so. Publishing one of these requires the user to first remove
the field in a separate, deliberate edit — this skill never removes it on their behalf.node scripts/sanitize-check.mjs <path>. Any hit stops here — fix and re-run, don't
proceed with known hits "to be cleaned up after."core/rules/ internals a public reader can't see?brain-public/,
commit there with a clear message, and note the promotion in the source artifact's frontmatter
(related: [[brain-public path or note]]``) so it's traceable which private entries have public
twins.sanitize-check passed — the mechanical check catches
names, not context ("a large district in the city's northeast" still identifies).phase0-vault-census.md and any
later equivalent) in any form./scar-skillTurn ONE lesson into a walkthrough command — the rare case where a check needs a person's judgment rather than a grep the scanner can run unattended. Most checkable lessons belong in tools/signature-scan/ instead, so check there first. The lesson's failure story becomes the skill's rationale block; shows the draft before writing and backlinks the skill to its source lesson.
A skill is the least-automatic destination a lesson has — not the top of a ladder. This file
used to open observation → lesson → rule → skill → tool, which reads as a promotion ranking and
routed lessons exactly wrong. tools/signature-scan/lib/signatures.mjs records the incident that
killed that framing: eight lessons at checkable: true, and the ladder said write eight commands,
each firing only when a user remembered to pick it, to run eight greps.
Rank the destinations by what actually makes them fire:
| Destination | Fires when | Who has to remember |
|---|---|---|
| indexed in the corpus | a matching scar_recall |
nobody |
tools/signature-scan/ |
every edit | nobody |
a /scar-* command |
it is typed | the user |
A command fires least often of the three, because a person has to choose it. That is not a defect to engineer away — it is the single property that makes this destination right for a check needing judgment a grep cannot supply, and wrong for everything else.
So a signature is a reason NOT to come here. If the signature is a grep-able pattern, the
scanner already runs it against every edit; the lesson needs promoted_to: signature-scan and
nothing else. Nine lessons were resolved that way rather than becoming nine commands. Reach for
this skill only when the check cannot be a line-level grep:
create policy / drop policy if exists lesson is the standing example);Recurrence is not the bar either (tightened 2026-08-12): cluster size already drives recall's ranking, so promoting on it a second time duplicates a signal instead of adding one, and produces a queue nobody can safely clear.
Promoting is not free. Every skill written here is symlinked into ~/.claude/skills/ and becomes
another command in the user's list — the surface AGENTS.md caps at nine for anything an agent
adds, and scripts/skill-roster.mjs fails verify on a twelfth that nobody claimed. (The
roster's two rows past nine are Joshua's own additions; adding one for him is not yours to do.)
So a skill promoted from here is over the cap, full stop. Promote when the check is real,
the scanner cannot run it, and a person is the right trigger — not to work through a backlog.
State what it replaces before drafting.
/scar-update's final report lists
candidates: checkable: true with no promoted_to).core/skills/<slug>/SKILL.md, following the shape of the other scar-*
skills in this directory): a description that states when to invoke it, and a body whose
rationale section is drawn directly from the lesson's failure story — the scar is why the
skill exists, and it stays visible in the skill text, not just in a linked-away lesson.promoted_to: <skill-slug>, and change status to deprecated if the skill fully supersedes needing to
read the lesson directly — but keep the lesson file, never delete it), and run
node scripts/install-skills.mjs to re-link ~/.claude/skills/ so the new skill is live in
every project immediately.promoted_to and typically deprecated.sanitize-check on the lesson before drafting the skill from it./scar-updateScar's one upkeep command — clear whatever backlog exists (distill, feedback), rebuild every kind of derived state so edits actually take effect (the recall index/kernel here, and the hook wiring out in each adopting project), run the verify gate, and report promotion candidates. Use when the user says "update Scar", "sync Scar", "evolve Scar" (or the pre-rename "update the brain"), or after a session that edited lessons/rules/skills. Say "sync only" to skip the backlog work and just rebuild + verify.
Editing a lesson, rule, or skill file does not change what an agent sees. scar_recall and
scar_context read .recall/index.json and .recall/kernel.md — gitignored, derived, and
stale until rebuilt. This is the routine that makes that not a trap, and the one place the
brain's own upkeep happens.
Derived state means both kinds. The index is one; each adopting project's
.claude/settings.json is the other, derived from lib/wiring.mjs and just as stale — a gate
added here runs nowhere until it is merged out there. Step 4 rebuilds the first, step 5c the
second. A sync that does only the first leaves "Scar is in sync" true of retrieval and false
of every gate, which is how a shipped gate sat live in 1 project of 8 (2026-08-19).
This command absorbed /brain-evolve and /brain-maintain (2026-08-12). Those were a
passive/active pair plus a health pass, and the split cost more than it bought: "check the backlog"
and "clear the backlog" were the same routine with one decision in the middle, and every user had
to pick which one they meant before knowing whether a backlog even existed. Now there is one
command, it looks first, and it does the work if there is work — see AGENTS.md, "Scar grows
inward": the /scar-* list is the one surface that stays small.
Look before doing anything. npm run distill:backlog and npm run feedback. The first
derives coverage by evidence-hash overlap and exits non-zero if it cannot see the evidence
graph, so a number it prints is a number it could actually derive — do not hand-roll it from
ls .distill/briefs/ | wc -l minus lesson files, cluster IDs are not stable across
cluster.mjs reruns and that arithmetic is meaningless.
Say what you found, then decide the shape of the run. If both are empty, this is a sync: skip to step 4 and say so. If either has work, name the size ("6 undistilled clusters, 2 open feedback reports") before starting — a distill pass is long, and a user who typed "sync the brain" after editing one lesson should get the choice rather than a surprise. If they said "sync only", skip step 3 entirely.
Clear the backlog, by routing — never inline.
/scar-distill, batches of 4–6 as that skill directs.
Refusals are a legitimate outcome; don't force a cluster that lacks a shared mechanism into
a lesson to clear the count. A PARTIAL cluster is a prompt to look, not a task./scar-address-feedback, working the inbox as that skill directs:
classify each report, reproduce the claim before acting on it, close with a concrete
resolution. regressed reports come first, and ask why the previous fix did not hold.Rebuild derived state. npm run recall:build — regenerates .recall/index.json and
.recall/kernel.md from what is currently on disk in career/, core/rules/, core/skills/.
This is the step that makes prior edits visible to recall; skipping it is the single most
common way a fix "doesn't work" after being made. Also run npm run brain-index if entry
frontmatter changed, to regenerate the per-folder indexes. The build prints a health line —
the overlap/distilled ratio and its delta since the last build. Read it: a rising ratio is
new lessons restating old ones, and that is the moment to prefer superseded_by: over a
tenth neighbour, not after the next distill batch.
4b. Regenerate the system map. npm run system-map — rewrites brain/SYSTEM-MAP.md and
brain/SYSTEM-MAP.html from the same data pass, a generated, human-readable explanation of the
whole system (the four layers, every gate's own WHY THIS EXISTS incident, every rule ranked
by scars, every skill's description, a pulse of recent DECISIONS.md entries). The .html is
a self-contained local page — open it straight off disk, no server. It exists so the person
maintaining this system doesn't have to hold all of it in their head as it grows. Private
only — brain/ never ships to ui/'s public deploy or brain-public/, and neither file must
ever be added to either.
4c. True up the figures the rebuild just invalidated. npm run repo-about:fix — rewrites the
inventory numbers README.md, AGENTS.md and core/AGENTS.md state about the corpus (record count,
lesson count, rules, skills, kernel size, compression ratio) to match what was just built, and
prints every digit it moved. These went stale on essentially every session that landed a lesson,
and the repair was mechanical every time; a gate that fires constantly and is always cleared the
same way trains its reader to clear it unread.
It repairs digits, never sentences. It also does not touch the recall benchmark's score —
that figure carries a long history in README narrating each movement and whether the run was
isolated, so rewriting the number without the argument would leave a stale claim behind a green
gate. If verify reports benchmark drift in the next step, that prose is yours to write.
Read the repairs it prints; do not just watch them scroll. Running this before the gate
means verify can no longer fail on a figure that moved, so the printed lines are now the only
place an unexpected jump shows up. A claim whose pattern stopped matching is still reported as
drift and still fails — that is the case --fix deliberately refuses to guess at.
4d. Retrain the recall noise filter. npm run noise-filter — refits it from every graded
recall and promotes it only past its bar, both baselines, a label-shuffle test, a used-result
cost ceiling and the recall benchmark (decision 233). A run that fails any of them writes the
filter OFF, so this step can switch filtering off as well as on. Before verify on purpose: the
benchmark in the next step then measures the engine with whatever this run left promoted. Quote
its last line; nothing to decide unless it says PROMOTED or turned a promoted filter off.
Then npm run retrain — refits SCAR Model v1 (the gate-verdict judge) and the worker-escalation
model from the labels that have arrived (decision 245), and checks the report-integrity bar
(decision 247). None is read by a gate, so this promotes nothing. Quote its last line; a DECISION line means a model's usable flag flipped,
and that goes to the user as a decision, never acted on here.
Run the gate. npm run verify — every check the repo has, in one run: the linters
(frontmatter, citations, gates, static, decisions numbering), sanitize-check, token-budget, the
recall benchmark and its negatives, reachability, the signature variant benchmark, the hooks
selftest, payload-selftest, engine-parity, index-fresh, repo-about drift, enforcement and
governance. Do not quote a count of checks or a benchmark score from this file — it prints
them, and every figure ever written here went stale (Rule 0.4). If it fails, fix the regression
before reporting done; never lower a threshold to make it pass — raising a budget is allowed
but it is a diff in scripts/token-budget.mjs, a decision, never a side effect of this routine.
Read the token-budget lines even on a pass. What is gated is the pool total per session
type; a single file sitting over its advisory cap prints as OVER … Not fatal and passes. That
is the early warning, and the next lesson that lands in the kernel is what turns it into a
failure. Say it in the report when a pool has under ~250 tokens of headroom.
engine-parity passing here does not mean the plugin shipped. It fails only if the
vendored copy is enabled in this repo. A change under tools/hooks/ or tools/recall/ in
this run still leads the published engine until the next plugin release — name that in the
report so the user knows a sync is pending (AGENTS.md, "has not shipped until it reaches the
plugin's vendored copy").
5b. Check the gates are actually installed. npm run hooks-coverage. The gates and the
signature scanner live in this repo but run from each adopting project's own
.claude/settings.json, merged by hand, one entry at a time — so a partial install looks
exactly like a complete one from inside a session, and every component that is missing fails
open and says nothing. It had drifted to the scanner being wired in one project of seven, which
is why its hit history read as empty rather than as absent. This is not part of verify:
coverage in another repository is not this repo's build to fail. Report any project short a
component — never a hand-merge, which is the thing that produced the drift in the first place.
The component list grows (it was 7 when this was written and is read from lib/wiring.mjs, not
from here); a new gate is not installed anywhere until 5c runs.
5b-ii. Check Scar's tools are reachable at all. npm run mcp-coverage, then
npm run install-mcp to close anything it reports. 5b asks whether each project runs the
gates; this asks whether each agent harness on this machine can reach scar_* in the first
place, and the two are independent — a project can be fully wired and still open in a harness
where the MCP server was never registered. That happened on 2026-09-20: a Codex session in an
adopting project loaded core/AGENTS.md, was told to call scar_context before non-trivial
work, and had no such tool; codex mcp list showed five servers and none was the brain. It had
been registered in exactly one place — ~/.claude.json — by hand, and nothing had ever checked.
Like 5b, deliberately not in verify: which harnesses a machine has is local state, not a
property of this repository. Flag a registration reported as elsewhere too — it points at a
different checkout, so the tools answer from another corpus and nothing looks wrong.
Read the config diff before accepting the install. The writers are the vendors' own CLIs,
and codex mcp add rewrites and re-normalises the whole of ~/.codex/config.toml rather than
appending. On 2026-09-20 that silently dropped another tool's marker comment — the delimiter
that tool uses to find its own block — which is
career/lessons/2026-09-04-a-guard-that-repairs-a-shared-config-key-to-its-own-known-good-value-silently-reverts-every-other-owner-s-entry
arriving from the vendor's side rather than ours.
5c. Close whatever 5b reported. npm run install-gates -- --all --dry-run, show the user the
diff, then npm run install-gates -- --all. Skip both if 5b was clean.
Why this belongs to sync and not to the reader. Read the top of this file: the whole reason
/scar-update exists is that .recall/index.json and .recall/kernel.md are derived state,
stale until rebuilt, so editing a lesson changes nothing an agent sees until the routine runs.
Each adopting project's .claude/settings.json is derived state by the same definition —
derived from lib/wiring.mjs — and goes stale the same way: adding a component there changes
nothing any project runs until it is merged out. Until 2026-08-19 this routine rebuilt the
first kind and merely reported the second, which made "Scar is in sync" true of the
index and false of every gate. outcome-gate shipped and sat live in 1 project of 8 while
every sync printed the same seven gaps and closed none. Decision 87 mechanized detecting
that drift; nothing mechanized repairing it, so the routine converged on nothing — the same
enforce-dont-declare shape, one notch downstream of where it was last found.
What makes the sweep safe to put in a routine, and what to keep true: it visits only
projects whose settings ALREADY point at Scar (discovery is lib/projects.mjs, shared
with hooks-coverage so the checked set and the installed set cannot diverge). It is additive
and idempotent, never replaces a hooks array, refuses an unparseable settings file by name
rather than overwriting it, and re-parses everything it writes. Those properties are asserted
in npm run hooks-selftest against a sandbox — a non-adopting project and a broken settings
file are planted every run. --all must never be able to adopt a new project; that is
/scar-adopt, which does more than hooks. Show the dry run first regardless: this writes into
repositories the user is not looking at.
Report, including unrun checks. One summary: what backlog existed and what was
done with it, what got rebuilt (including whether SYSTEM-MAP.md changed), whether verify held.
If this run changed a mechanism, refused a requested fix, or found a report wrong about its own
cause, that is a DECISIONS.md entry — numbered in sequence (decisions-lint gates it), written
in the same session, with the reason. The feedback closure text should cite the number.
Then classify every finding as a decision or a repair, and record the two counts. A
decision is the owner's by design — raise a budget, refuse a gate, cut a version. A repair
is anything that fixed drift or re-derived state: a stale rebuild mid-run, a figure that moved,
a verify failure with a mechanical cause, a hand-kept table missing a row. npm run upkeep -- --decisions N --repairs M --note "…" appends the run to .recall/upkeep.jsonl; a pure sync is
0 0 and is worth recording, since syncs are the evidence the system is settling. Say the two
numbers in the report. A repair that recurs across runs is a mechanism nobody has built yet —
name it. The counts are self-reported and feed nothing ranked; npm run upkeep reads the
trend back by release. Then list every type: lesson with
checkable: true and no promoted_to — lessons naming a grep-able signature that nothing
runs yet. The default destination is tools/signature-scan/, not a command: if the
signature is a real pattern, npm run signature-scan already executes it and the lesson just
needs promoted_to: signature-scan. Only a check the scanner cannot run — cross-line, or about
behaviour rather than a code shape — is a /scar-skill question at all. Say "0 issues" plainly when it is
clean; a check that ran and found nothing is worth stating, otherwise the user cannot tell it
from a check that never ran.
Then npm run enforcement and quote the marginal rate, not the level. That list of
promotion candidates is a numerator with no denominator on its own: it cannot distinguish a
corpus converting lessons into checks from one accumulating advice faster than it wires any of
it up. Read the marginal rate against the lifetime one and say which way it moved. Never
propose closing the gap as a goal — the ceiling is not 100%, and a lesson with no grep-able
form is supposed to stay advisory.
Cut a version only when the user asks, and only after step 5. npm run economics:version -- <label> --note "one line on what changed" closes the release that is accruing spend,
freezes its accounting, and opens the next label (0.NN is the convention; check the list the
command prints when run bare). It is the one write economics never makes, and it is not
automatic on purpose: nothing derivable marks a version boundary, so the person who did the work
says so. Cut it after verify, never before — the release freezes on whatever state exists at
that moment, and a cut before the rebuild attributes stale counts to the new label. Do not
import or probe any hook module to check the result; a self-invoking module records phantom
loads into the window you just opened (2026-09-01-a-script-that-invokes-itself-at-module-scope).
npm run economics afterwards shows the new release open.
/brain-maintain also ran a staleness pass over entries with status: open/in-progress and
a duplicate-topic pass. Both were dropped when it merged in here, for reasons, not for brevity:
writeLesson (tools/distill/lib/lesson.mjs) already
scores a new draft's triggers/mechanism against the existing index and surfaces overlaps ≥ 0.5
as a warning. Same economics as core/rules/sanitize-at-write-time: paid once, at the moment
the writer can still act on it, instead of every upkeep run forever.If either turns out to be needed again, bring it back here, as a step. Do not bring back a separate command.
tools/distill/lib/lesson.mjs's denylist gate, writeFeedback's
consolidation check) and duplicating that logic here is exactly the "two paths, one gate" trap
AGENTS.md warns about. This skill routes; it does not re-implement.core/rules/enforce-dont-declare's failure shape — a
backlog that exists and is worked honestly beats one that reads empty and lies./scar-skill's call, one lesson at a time, with a draft shown first: it adds to
the /scar-* surface, which is capped, so it is a decision and not a chore./toolkit-analyzeAnalyzes an existing Jira ticket by key (e.g. /toolkit-analyze PROJ-807), fetches its description and comments, produces a structured execution plan, waits for approval, then posts the approved plan as a Jira comment and transitions the ticket to In Development. Use this to think through a ticket again before coding — whether it's brand new or you're picking it back up after a break. Not for generating new tickets (use /toolkit-plan + /toolkit-jira for that) and not for pasted ticket text (use /toolkit-jira resume for that).
You are the ticket analysis agent. Your job is to read an existing Jira ticket, understand (or re-understand) what needs to be done, produce a clear execution plan, get user approval, and post the plan back to Jira as a comment for documentation.
You never write code directly. You produce a plan and wait for approval before handing off to /toolkit-code (or whichever skill the consuming repo uses for implementation).
This is the tool for stepping back and reconsidering a ticket before touching code — whether that's the first time you're looking at it, or you're returning to it after time away and want a fresh read rather than trusting old assumptions.
No config file to set up — values are detected from the repo or asked once and remembered in ~/.dev-skills.env (see CONFIG.md for the full read/write procedure and category definitions):
jiraProjectKey / ticketPrefix — asked once, then written — used to validate the ticket key looks right and to label the Epic fieldappRoots — detected fresh every run — used to phrase the "files to touch" section in terms of this repo's actual layer names (e.g. backend/frontend, or just app for a single-layer repo)referenceSource — asked once, then written — if this repo audits its work against another reference (a sibling repo, a Figma spec, a design doc), the plan's reference section asks this instead of assuming a specific cross-reference workflow/toolkit-analyze <TICKET-KEY>
Examples:
/toolkit-analyze PROJ-806/toolkit-analyze PROJ-813This skill's entire purpose depends on mcp__claude_ai_Atlassian_Rovo__* tools being callable — unlike create-pr's optional Jira-comment step, there's no fallback path here if they're not. Before Step 1, check your Claude session's available tool list for any mcp__claude_ai_Atlassian_Rovo__* entry. If none appear, the MCP isn't loaded — tell the user and stop.
If the Atlassian MCP is not available, stop immediately and tell the user:
The Atlassian MCP isn't available in this session —
/toolkit-analyzecan't fetch or update Jira tickets without it. Check your MCP server configuration, or paste the ticket's description manually and use/toolkit-jira resumeinstead (it works from pasted text, no MCP needed).
Do not let an unhandled tool-call error be the user's first signal that something's wrong — this check (or the explicit message above once a real call fails) is what distinguishes "the MCP isn't configured" from "this Jira ticket doesn't exist."
Use mcp__claude_ai_Atlassian_Rovo__getJiraIssue with the ticket key.
Extract:
Based on the description, determine:
Type of work — pick whatever categories make sense for this repo's appRoots. Common shapes:
backend — service/API-layer changes onlyfrontend — UI-layer changes onlyfullstack — bothbug — fix a specific broken behaviorchore — non-feature work (config, docs, refactor)Reference material — if referenceSource is configured, identify which specific files/specs/docs in that source need cross-referencing for this ticket. If no referenceSource is configured, skip this section.
Files to touch — list the specific files that will be created or modified, scoped by this repo's appRoots paths.
Endpoints / units of work — list every API endpoint, function, or component that needs to be created, modified, or audited.
Dependencies — any other tickets or features this depends on.
Risks / unknowns — anything that needs investigation before coding starts.
Present the plan in this format:
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Execution Plan — <TICKET-KEY>: <Summary>
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Type: <backend | frontend | fullstack | bug | chore>
Epic: <PREFIX-XXX — Epic name>
Branch: <suggested branch name, e.g. PROJ-806>
── OVERVIEW ──────────────────────────────
<2–3 sentences explaining what this ticket does and why, in plain English.>
── REFERENCE MATERIAL ──────────────────── (omit if no referenceSource configured)
<Reference path/spec>: <specific file(s) or section to cross-reference>
── CHANGES BY LAYER ──────────────────────
<One subsection per appRoots key touched>
File: <path under that layer's root>
Unit of work Action
─────────────────────────────────────────────────────────
<function/endpoint/component> [ ] Create — <what it does>
<function/endpoint/component> [ ] Modify — <what changes>
<function/endpoint/component> [ ] Audit parity — <what to verify>
── ACCEPTANCE CRITERIA ───────────────────
[ ] <criterion from ticket>
[ ] <criterion from ticket>
[ ] <repo's detected quality-gate commands pass>
── QA TESTING GUIDE ────────────────────── (omit if not applicable)
⚠️ <any login requirement or special test-account note>
| Area | Where | What to verify |
|------|-------|----------------|
| <area> | <path/URL> | <what to check> |
── RISKS / UNKNOWNS ──────────────────────
• <anything that needs investigation or could be tricky>
── SKILL TO INVOKE AFTER APPROVAL ───────
/toolkit-code (or whichever this repo's own local build skill is called)
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Approve this plan? Reply 'yes' to start, or tell me what to adjust.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Do NOT invoke any other skill or write any code until the user explicitly approves.
If the user asks for adjustments, revise the plan and present it again.
Once approved, use mcp__claude_ai_Atlassian_Rovo__addCommentToJiraIssue to post the approved execution plan to the ticket in markdown format. Include:
Then hand off to /toolkit-code (or whichever skill the user identified for implementation).
After posting the comment, use mcp__claude_ai_Atlassian_Rovo__transitionJiraIssue to move the ticket to In Development status. Transition IDs vary per Jira project — if unknown, fetch them first via mcp__claude_ai_Atlassian_Rovo__getTransitionsForJiraIssue rather than guessing a hardcoded ID. Match by target status name (e.g. "In Development"), not the transition button label — these can differ.
mcp__claude_ai_Atlassian_Rovo__getAccessibleAtlassianResources and remembered in ~/.dev-skills.env, not asked for./toolkit-analyze vs. /toolkit-jira resume — Read This First/toolkit-analyze <KEY> fetches a ticket live from Jira and is meant for reconsidering scope — re-reading a ticket with fresh eyes, whether for the first time or after time away. Posts back to Jira and transitions status./toolkit-jira resume works from a ticket the user pastes directly into the chat, with no live Jira fetch and no comment/transition side effects — a lighter-weight "I already have the text, just show me the plan" path.If you're not sure which to use: if you want this session's plan recorded back on the ticket itself, use /toolkit-analyze. If you just have the text in front of you and want to move fast, use /toolkit-jira resume.
Run /toolkit-analyze <TICKET-KEY> with the Jira ticket key you want to (re)analyze.
/toolkit-check-bugsInteractively scans a Jira bug board/project and, for every ticket not already logged from a prior run, checks whether the described bug/feature exists in a configured target repo — found, not yet ported, or inconclusive. Persists a scan log so repeat runs skip already-checked tickets unless the user names specific tickets to rescan. Read-only — never writes to Jira, never fixes code. Use when the user says "/toolkit-check-bugs", "scan the bugs board", or "is this bug in <repo>".
You are the bug-check agent. Your job is to triage a Jira bug board against a configured target repo (typically one under active migration/port from a legacy system) and report, per ticket, whether the described bug/feature exists there. You never fix anything and you never write back to Jira. You produce a table with file:line evidence for every claim, and you never scan a ticket twice unless the user explicitly asks you to.
No config file to set up — values are asked once and remembered in ~/.dev-skills.env (see CONFIG.md for the full read/write procedure). This skill is a cross-repo tool rather than a per-consuming-repo one, so its keys use a CHECK_BUGS_ prefix instead of a repo slug:
CHECK_BUGS_JIRA_CLOUD_ID — detected once, then written — via mcp__claude_ai_Atlassian_Rovo__getAccessibleAtlassianResources (a lookup, not a question).CHECK_BUGS_LAST_PROJECT_KEY — detected once, then written, always re-confirmed — remembers the last project/board key used, purely as a convenience default offered back to the user at the start of every run. Never used silently — Step 1 always asks, this just makes the suggested answer faster to accept.CHECK_BUGS_TARGET_PATH — asked once, then written — absolute path to the target repo's checkout, i.e. the repo being checked for whether each Jira-reported bug/feature exists there. There is nothing skill-specific about which repo this is — it's whatever repo the user is triaging against, asked plainly ("Which repo should I check tickets against? Give me its absolute path.").CHECK_BUGS_TARGET_NAME — asked once, then written — a short human-readable label for the target repo (e.g. its directory/project name), used purely for display in prose and report headings so output doesn't read as a generic "the target repo" everywhere. Default to the last path segment of CHECK_BUGS_TARGET_PATH if the user doesn't offer a preferred name.CHECK_BUGS_PRIORITY_FILTER — asked once, then written — no universal default; ask what priority value(s) matter for this board (many Jira setups don't even use severity/priority the same way). Board swimlanes/columns (e.g. a custom "Expedite" grouping) are not reachable via the Atlassian MCP — there is no API to read a board's saved swimlane queries, only issue-level JQL search, so a priority-field filter is the closest reliable substitute if the user is trying to approximate a specific board column.CHECK_BUGS_STATUS_FILTER — asked once, then written — no universal default; ask which status value(s) to include (a status list, OR'd together in JQL via status in (...)). A common useful pairing is "the open/newly-reported status" plus "the resolved/fixed status" — the latter is useful even though it's closed, since it lets this skill confirm whether a fix that shipped upstream actually made it into the target repo's port. Whatever the user picks, treat it as a deliberate narrowing versus "all statuses" (which usually pulls in noise like Duplicate/Not-a-Bug/Blocked).CHECK_BUGS_COMPONENT_FILTER — asked once, then written, only if the project actually has a components field with more than one value — check via the ticket sample fetched in Step 2 before asking; don't ask about a field that doesn't vary. If a Jira project's tickets span multiple product lines/components and the target repo only covers one of them, filtering out-of-scope components avoids wasting a full triage pass on tickets that will always read "Not yet ported" for the wrong reason (out of scope, not un-ported). First-run mistake to not repeat: an early run of this skill scanned without a component filter on a multi-component board and wasted a full triage pass on over a dozen out-of-scope tickets before this was caught — always check whether this filter is relevant before skipping it.Every ticket this skill has ever scanned is recorded in <CHECK_BUGS_TARGET_PATH>/.claude/bug-triage-log.md, one entry per ticket:
## <TICKET-KEY> — scanned <YYYY-MM-DD>
- Jira status at scan time: <status>
- <CHECK_BUGS_TARGET_NAME>: <symbol> <finding> — `<file:line or n/a>`
Create the file (and .claude/ if needed) on first use. Never overwrite the whole file — append new entries; when re-scanning a ticket (explicit rescan only, see Step 3), replace that ticket's existing section in place rather than appending a duplicate.
Check the session's tool list for any mcp__claude_ai_Atlassian_Rovo__* entry before proceeding. This skill's entire premise — scanning many live tickets — has no offline fallback (unlike /toolkit-jira resume, which works from pasted text).
If the Atlassian MCP is not available, stop and tell the user:
The Atlassian MCP isn't available in this session —
/toolkit-check-bugscan't fetch tickets from Jira without it. Check your MCP server configuration and try again.
Always ask, every invocation — never silently reuse a remembered value without asking first.
CHECK_BUGS_LAST_PROJECT_KEY is set, ask via AskUserQuestion with the previous key as the first option, labeled so it reads as the default (e.g. BUGS (last time)), plus an option for "different project" — the user can pick the previous one in one click/Enter, or choose to type a new key.Resolve the user's answer to a project key (e.g. "the BUGS board" → BUGS). Update CHECK_BUGS_LAST_PROJECT_KEY to whatever was used this run.
If the user's invocation already included an explicit board/project name (e.g. "/toolkit-check-bugs BUGS"), skip asking and use it directly — but still confirm back in one line ("Checking project BUGS...") so a typo doesn't silently scan the wrong project.
If CHECK_BUGS_TARGET_PATH isn't set yet, ask which repo to check tickets against and its absolute path, plus what to call it for display (CHECK_BUGS_TARGET_NAME). If already set, confirm it back in one line rather than re-asking every run ("Checking against <CHECK_BUGS_TARGET_NAME>...") — this is a stable fact about the setup, not something that varies run to run the way board/priority/status choices might.
Then confirm the priority scope — ask via AskUserQuestion with CHECK_BUGS_PRIORITY_FILTER (if previously set) as the first option, labeled <value> (last time), plus an option for "different priority / all priorities" so the user can type another value or opt out of filtering entirely. If never set, ask plainly what priority value(s) matter, or whether to skip priority filtering. Update CHECK_BUGS_PRIORITY_FILTER to whatever was used this run. Skip asking only if the user's invocation already named a priority explicitly (e.g. "/toolkit-check-bugs BUGS priority=High").
Then confirm the status scope the same way — ask via AskUserQuestion with CHECK_BUGS_STATUS_FILTER (if previously set) as the first option, labeled <values> (last time), plus an option for "different statuses / all statuses". If never set, ask plainly which status(es) to include. Update CHECK_BUGS_STATUS_FILTER to whatever was used this run.
Rescan mode: if the user says /toolkit-check-bugs rescan <KEY> [<KEY> ...] (or "rescan BUGS-123"), skip the board-wide fetch in Step 2 entirely — jump straight to Step 4 for exactly those ticket keys, using mcp__claude_ai_Atlassian_Rovo__getJiraIssue per key instead of a board-wide search.
Verify this project's actual "bug" issue-type name once via mcp__claude_ai_Atlassian_Rovo__getJiraProjectIssueTypesMetadata — don't assume the literal word "Bug" if this project calls it something else.
While building the query, also check (via the first small batch of results, or mcp__claude_ai_Atlassian_Rovo__getJiraIssueTypeMetaWithFields) whether this project's tickets actually carry a components field with more than one distinct value. If so, and the target repo only covers one product line, ask about CHECK_BUGS_COMPONENT_FILTER per the config note above before finishing the query.
Call mcp__claude_ai_Atlassian_Rovo__searchJiraIssuesUsingJql, scoped to whichever of CHECK_BUGS_COMPONENT_FILTER, CHECK_BUGS_PRIORITY_FILTER, and CHECK_BUGS_STATUS_FILTER are set. Build the status in (...) clause by quoting each comma-separated status value:
cloudId: <CHECK_BUGS_JIRA_CLOUD_ID>
jql: project = <KEY> AND issuetype = <resolved bug type> [AND component = "<CHECK_BUGS_COMPONENT_FILTER>"] [AND priority = "<CHECK_BUGS_PRIORITY_FILTER>"] [AND status in (<CHECK_BUGS_STATUS_FILTER values>)] ORDER BY created DESC
fields: ["summary", "description", "status", "priority", "labels", "components", "created", "updated"]
maxResults: 100
If the fetched set is unbounded/large (e.g. spans years of history with mostly-resolved tickets), consider proposing a created >= date bound to the user rather than silently paging to the full ~300 cap — see the cap-exceeded handling below.
Keep field lists lean on the first fetch — description alone can blow past the response size limit on boards with long ticket bodies. If a call errors on response size, drop description from fields for the list-fetch pass and pull it per-ticket in Step 4 only for tickets that survive the Step 3 log filter.
If nextPageToken is returned, loop until exhausted. Cap total fetched at ~300 — if exceeded, tell the user and ask whether to narrow scope (date range, label, priority) rather than silently truncating.
There is no API to fetch a board's literal Kanban/sprint view (swimlanes, columns) — this is project-scoped issue search filtered by the fields above as the closest reachable equivalent. Say so if the user seems to expect the true swimlane/column boundary back — a field-based filter is a verified proxy, not a guarantee of an exact match to the visual board.
Read <CHECK_BUGS_TARGET_PATH>/.claude/bug-triage-log.md (if it exists) and extract every ticket key already logged.
Drop any fetched ticket already in the log — do not re-scan it. Keep only tickets that have never been scanned before.
If the user invoked with explicit rescan targets (Step 1), this filtering step doesn't apply — those tickets proceed to Step 4 regardless of log state, and their existing log entries (if any) will be replaced, not duplicated.
If zero tickets remain after filtering (nothing new, no rescan requested), skip straight to reporting: "0 new tickets since last check on project <KEY> (<N> already logged). Say /toolkit-check-bugs rescan <KEY> to re-check a specific one."
From the ticket's summary + description, extract candidate screen/feature terms:
<CHECK_BUGS_TARGET_PATH>/.claude/screen-specs/*.md filenames first, if that directory exists in the target repo.If a screen-spec match is found: read it. It typically gives an exact screen/file path (and a backend-procedures table, if the ticket is data/logic-shaped rather than purely visual) — use it directly, skip guessing.
If no screen-spec match: fall back to a plain recursive grep across <CHECK_BUGS_TARGET_PATH> for 2-3 keyword variants, e.g. grep -rli "<keyword>" <CHECK_BUGS_TARGET_PATH> --include="*.ts*" (adjust the include pattern to the target repo's actual language/extension once observed). Don't hardcode this target repo's internal directory layout (e.g. a specific apps/*/... structure) in this file — it's repo-specific and will silently stop matching if that repo's layout changes, and this skill may be pointed at a different target repo entirely on another run. If a directory pattern is worth remembering for future runs against this same target (screens live under one particular subtree, source under another), that's a <CHECK_BUGS_TARGET_PATH>-scoped detail — note it in that repo's scan-state log header, not in this skill file.
If still ambiguous, or multiple equally-plausible candidates exist: do not pick one arbitrarily. List every candidate considered and mark the ticket Inconclusive — a wrong citation is worse than an honest "couldn't narrow it down."
For the resolved candidate file(s), grep/read targeted offsets (never a whole file from line 1) for the specific behavior described. Classify:
file:line + a one-line quote/paraphrase of the matching logic.For every ticket just processed (new scan or explicit rescan), write or replace its section in <CHECK_BUGS_TARGET_PATH>/.claude/bug-triage-log.md per the format under "Scan-state log" above. Create the file (with a # Bug Triage Log header) if it doesn't exist yet.
Bug Check — Project <KEY> vs. <CHECK_BUGS_TARGET_NAME>
══════════════════════════════════════════════════════
Newly scanned this run: <N> | Skipped (already logged): <N> | Fetched: <timestamp>
| Ticket | Summary | Jira Status | <CHECK_BUGS_TARGET_NAME> Status |
|---|---|---|---|
| <KEY> | <summary> | <status> | <✅/🚧/❓ finding — file:line or reason> |
Legend: ✅ exists in target repo · 🚧 not yet ported · ❓ inconclusive
Full report: <artifact URL from Step 7>
To re-check a specific ticket next time: /toolkit-check-bugs rescan <TICKET-KEY>
If nothing new was found (Step 3), skip the table and report the "0 new tickets" message instead.
Every run that produces a non-empty table (i.e. Step 6 wasn't the "0 new tickets" case) also gets published as a readable HTML artifact — this is the primary deliverable a user shares/reopens, not just the chat table.
Load the artifact-design skill before building it (per its own instructions), then build a single self-contained HTML file with:
<script>.https://<jira-site>/browse/<KEY>, monospace) · Summary · Jira status · target-repo status (as a colored pill, not just an emoji) · Finding (prose, with file:line citations in <code>).prefers-color-scheme + data-theme overrides) per the artifact-design skill's standard pattern — this report gets reopened later, possibly in the opposite theme from when it was made.Write the file to the scratchpad directory, then call the Artifact tool on it. Give it a stable file path per project/target-repo pairing (e.g. bug-triage-report-<target-name>.html) so re-running /toolkit-check-bugs on the same project later can redeploy to the same artifact URL rather than minting a new one each time — pass the same file_path again in-session, or the artifact's url if resuming in a new conversation (use action: "list" to find it if not already known).
Report both the chat table (Step 6) and the artifact link together — the chat table is the quick scan, the artifact is the shareable version.
CHECK_BUGS_TARGET_PATH./toolkit-check-bugs rescan <KEY>.Say /toolkit-check-bugs and I'll ask which Jira board/project to check, which repo to check it against, and what priority/status scope to use. Or say /toolkit-check-bugs <PROJECT_KEY> to skip the first question. To re-check specific tickets already logged: /toolkit-check-bugs rescan <TICKET-KEY> [<TICKET-KEY> ...].
/toolkit-codeBuild coordinator — takes an approved spec/ticket and builds it, backend layer first then frontend layer, with a mandatory responsiveness pass and security/locale/theme checklist baked in. Use when the user says "/toolkit-code <name>", "build the X screen", or "implement X" after a ticket/spec is approved.
You are the build coordinator. You orchestrate getting a unit of work from "approved spec" to "built, audited, ready for tests" — backend layer before frontend layer, with a mandatory responsiveness pass once the UI exists. You write code directly when this repo has no finer-grained local skill to delegate to; if this repo does have its own layer-specific skills (a backend-builder, a frontend-builder), invoke those instead of duplicating their work.
This skill's shape is a merge of two patterns that converged independently across consuming repos: backend-coverage-audit-then-frontend-build, finished with a non-optional responsiveness fix pass. Don't treat any of this as repo-specific — every repo building a UI on top of an API benefits from checking the API exists and matches expectations before building the screen that depends on it.
No config file to set up — values are detected from the repo or asked once and remembered in ~/.dev-skills.env (see CONFIG.md for the full read/write procedure and category definitions):
appRoots — detected fresh every run — which folder is "backend" and which is "frontend" (or whatever layer names this repo actually uses).referenceSource — asked once, then written to ~/.dev-skills.env — if this repo cross-references another source for parity, asked the same way /toolkit-qa//toolkit-jira//toolkit-plan ask it.localeConvention / themeConvention / localeDictionaryPath — detected once, then written — drives the locale/theme checklist below, same convention /toolkit-qa detects. See Phase 2 step 4.responsiveness — asked once, then written — breakpoints for the mandatory Phase 3 sweep.repoRules — detected fresh every run — checks conventional paths for a repo-rules file.specCachePath — detected fresh every run — checks what's already in use under .claude/.buildSkillName / buildSkillNames — detected once, then written — if this repo has its own finer-grained local build skill(s) for specific layers (e.g. a dedicated backend-only or frontend-only skill), invoke those instead of writing that layer's code directly here. Detected by scanning .claude/skills/ for genuinely local (non-symlinked) skills; most repos won't have one — /toolkit-code is meant to be sufficient on its own./toolkit-code <unit of work>
Examples: /toolkit-code invoice export, /toolkit-code account summary screen
Before touching any file, output an Implementation Plan and wait for explicit approval.
specCachePath (or whatever's already cached under .claude/) for a spec from /toolkit-plan//toolkit-jira//toolkit-analyze matching this unit of work. Look for .md files matching this unit of work's name (slugified, e.g. work-orders-export.md) in specCachePath. If found, read it instead of re-deriving requirements — it already has scope, parity decision, and acceptance criteria.referenceSource is configured and parity applies (per the spec, or ask once if no spec exists — same question /toolkit-plan//toolkit-jira ask): read the reference material (grep-first, targeted offsets, never a whole file from the top) to determine what each layer needs to do/return/render.getWorkOrders, grep: grep -r 'getWorkOrders\|GET.*work-orders' <appRoots>/ to find function definitions, route registrations, or tests.Implementation Plan — <Unit of Work>
════════════════════════════════════
Reference: <referenceSource summary, if parity applies — omit otherwise>
PHASE 1 — Backend
Unit of work Status Action
─────────────────────────────────────────────────────────────────
<endpoint/function> ✅ exists Verify response parity
<endpoint/function> ❌ missing Implement
<endpoint/function> ⚠️ shape mismatch Fix field names
(one row per required unit of work; omit this phase entirely if no backend layer involved)
PHASE 2 — Frontend
File: <appRoots.frontend>/<path>
[ ] Wire <N> queries / <N> mutations to the Phase 1 procedures
[ ] <N> filters / <N> fields / <N> columns matching the reference (omit if parity not applicable)
[ ] Locale: all visible strings wrapped in <localeConvention> — zero hardcoded copy (omit if localeConvention not configured)
[ ] No orphaned locale keys — every interface key added has a matching dictionary value in every locale, and every dictionary value added is actually wired into the screen (omit if localeConvention not configured)
[ ] Theme: all colors from <themeConvention> — no hardcoded hex (omit if themeConvention not configured)
[ ] Responsive layout — handled by Phase 3 below, mandatory before this is done
SECURITY CHECKLIST
Access control Backend endpoints scoped to the right auth/role level? [ ] yes / [ ] needs fix
Injection/XSS Inputs validated, no raw string-built queries, no
dangerouslySetInnerHTML/innerHTML on user content? [ ] yes / [ ] needs fix
Sensitive data No secrets/PII/tokens in responses, logs, or client state? [ ] yes / [ ] needs fix
Mass assignment Explicit field allowlist on writes, not a raw input spread? [ ] yes / [ ] needs fix
Screen-specific risks: <flag anything unusual — file uploads, payment data, PII fields, admin-only actions>
ACCESSIBILITY CHECKLIST (frontend only — skip if no UI changes)
Every interactive element with NO visible text label must have an
accessibilityLabel (icon buttons, close/X buttons, fab buttons, etc.)
Use the repo's locale system for the label value — never hardcode English.
[ ] Icon-only buttons have accessibilityLabel={locale.<key>}
[ ] Close / dismiss / X buttons have accessibilityLabel={locale.modals.close} (or equivalent)
[ ] Loading-state buttons retain their accessibilityLabel (label doesn't disappear when spinner shows)
[ ] Error banners have accessibilityRole="alert" so screen readers announce them immediately
[ ] Form inputs that show errors have aria-invalid={!!error} + aria-describedby pointing to the error element's nativeID
If any item above is missing: add it before Phase 2 is marked complete — these are QA requirements, not optional.
Scope: Simple / Medium / Complex
Approve this plan? Reply 'yes' to start, or tell me what to adjust.
For each required unit of work identified in the plan:
appRoots's backend path) and is actually wired up/registered, not just present as a dangling file.buildSkillName/buildSkillNames entry mapped to the backend layer, invoke that skill instead of writing the code directly here.Do not proceed to Phase 2 until every required backend unit of work passes. A frontend built against a missing or shape-mismatched endpoint will silently render broken or empty. If a backend unit fails its smoke test: stop Phase 2, return to Phase 1, fix the failure, re-run the smoke test, then resume Phase 2.
Before starting Phase 2, confirm new/changed backend code actually returns data when exercised (restart the dev server if this repo requires it for changes to take effect). E.g. for a REST endpoint, curl it with valid auth and confirm the response body is non-empty and matches the expected schema. For a tRPC procedure, call it from the client and inspect the network response. If something returns unexpectedly empty, investigate now — a silent empty response produces a blank screen that's hard to debug once the frontend is built on top of it.
Once Phase 1 is verified complete and parity-correct:
Build (or modify) the UI for this unit of work under appRoots's frontend path. If this repo has its own frontend-specific build skill mapped via buildSkillName/buildSkillNames, invoke that instead of writing the screen directly here.
Wire data calls to the Phase 1 procedures confirmed above.
If parity applies, mirror the reference's element placement, defaults, and behavior — not just its data shape.
Apply the locale/theme conventions — read from ~/.dev-skills.env if already set; otherwise detect once from an existing screen (same detection /toolkit-qa's Structural Audit B/C sections run) and write the result to ~/.dev-skills.env as <SLUG>_LOCALE_CONVENTION/<SLUG>_THEME_CONVENTION so this is a one-time cost, not paid on every screen. To detect once: grep for the locale hook pattern (e.g. useTranslation(), i18n.t()) and theme hook pattern (e.g. useTheme(), styled()) in existing screens — record the exact function names and import paths. Skip either cleanly only if genuinely nothing of that shape exists anywhere yet.
When adding a new locale-dictionary key: the interface declaration, the dictionary value in every locale, and the screen's wiring to use it are one atomic unit of work — never commit or hand off with only one or two of the three done, even temporarily. This is the single most common way a locale change breaks the whole repo's typecheck (an interface key with no matching dictionary value in one or more locales). Before considering this step done, re-check for hardcoded strings hiding in component props (title=, placeholder=, label=), multi-line <Text> blocks, and ternaries/template literals — a same-line-only grep sweep reporting "0 remaining" is not sufficient proof of completeness on its own.
Run this repo's lint/typecheck commands (read from qualityGate in config if set) and fix everything before proceeding to Phase 3.
A screen is not done without this phase, regardless of how it looked when first built.
/toolkit-responsive <route> — it's report-only by design, sweeping this repo's configured breakpoints and returning PASS/FAIL findings without editing anything./toolkit-responsive <route> to confirm PASS at every breakpoint before calling this unit of work done.After all phases complete:
Unit of work: <name>
Backend status:
✅ <unit of work> — exists, parity verified
✅ <unit of work> — implemented now, parity verified
⚠️ <unit of work> — shape corrected (was: X, now: Y)
Frontend status:
✅ Built and wired to backend
✅ Mirrors reference element placement (omit if parity not applicable)
✅ Lint clean
✅ Typecheck clean
Responsiveness status:
✅ /toolkit-responsive audit run, fixes applied, re-verified PASS
Next: /toolkit-tester and /toolkit-unit-tester to write tests, then /toolkit-qa for the full audit.
buildSkillName/buildSkillNames), invoke those for the layers they cover instead of duplicating their work here — /toolkit-code writes code directly only for layers with no dedicated local skill.Tell me which unit of work to build. I'll check for an approved spec/ticket, audit the backend, build the frontend, and finish with a mandatory responsiveness pass.
/toolkit-code-reviewReview uncommitted working-tree changes for correctness, architecture, and repo-rule violations before committing. Review is read-only — never edits source. Reports PASS/FAIL/WARN with file:line. On a PASS, drafts a commit message. On explicit request after a PASS, can commit the reviewed files (optionally a named subset) and advance a linked Jira ticket. Use when the user says "/toolkit-code-review", "review my changes", "review before commit", "check my code", "is this ready to commit", or "review + commit and move the ticket".
Deep pre-commit review of the current working-tree diff. Covers correctness, security, debug leftovers, repo-specific architecture rules, and (in repos with more than one platform/surface) cross-platform safety. The review itself is read-only — never edits source or stages files.
On PASS, drafts a commit message. On FAIL, stop and report — do NOT draft a commit message until the user fixes the failures.
When (and only when) the user explicitly asks, a final step then commits the reviewed files and advances a linked Jira ticket to a review status. A bare /toolkit-code-review never commits.
No config file to set up — values are detected from the repo or asked once and remembered in ~/.dev-skills.env (see CONFIG.md for the full read/write procedure and category definitions):
appRoots — detected fresh every run — used to bucket changed files by layer for the report.qualityGate — detected fresh every run — lint/typecheck commands per layer, read from each layer's manifest scripts.repoRules — detected fresh every run — checks conventional paths for a repo-rules file. Skips the repo-architecture-rules step and notes it as not found if none exists.localeDictionaryPath — detected once, then written — the file holding this repo's actual locale/translation data, if one exists (see CONFIG.md). Enables the diff-scoped locale dictionary hygiene check (step 5.5). Skips that step cleanly if not configured.ticketPrefix — asked once, then written — used to resolve a ticket key from the branch/commit message.jiraCloudId — detected once via live MCP lookup, then written — required only if the user requests the ticket-advance step.commitMessageConvention — detected fresh every run — checks conventional paths for a commit-style doc. Defaults to Conventional Commits (type(scope): imperative description) if none found, and says so.All optional:
/toolkit-code-review path/a.tsx path/b.tsx) — scope the review and the eventual commit to just those files. Without paths, the review covers the whole git diff (staged + unstaged). Scoping to named files is the common case: the working tree often holds unrelated in-progress changes that must NOT ride along in the commit./toolkit-code-review PROJ-1715) — the ticket to advance in the final step. If omitted, it's derived from the commit message / branch.Run in parallel:
git status
git diff --staged
git diff
git log --oneline -5
If the working tree is clean, say so and stop.
Identify which files changed and bucket them by appRoots layer, plus generic buckets for docs/**/*.md (skip most checks) and tooling/config files (skip most checks).
Run this layer's qualityGate lint/typecheck commands, scoped to the layers actually touched.
Distinguish a missing script from a real failure. If a command errors with "missing script"/"command not found" rather than running and reporting actual lint/type errors, that's a detection issue — report it as ⏭ SKIPPED — <command> not found in this layer's manifest scripts rather than ❌ FAIL. A misdetected command isn't a code-quality finding. Also handle the case where the script exists but immediately errors on startup (e.g. linter binary not found) — mark as ⏭ SKIPPED — <command> misconfigured rather than a test failure.
FAIL if either command runs and exits non-zero for a real reason.
Scoped to the same diff as everything else in this skill — this is not a whole-repo security audit, it's "did this change introduce a new risk." Skip a sub-check entirely if no changed file matches its relevant file type (e.g. skip 3d on a docs-only diff).
grep -nEi 'password\s*=|secret\s*=|api_?key\s*=|token\s*=.*["\x27][A-Za-z0-9+/]{20,}|-----BEGIN' <changed-files>
Also check for .env* or other secret-bearing config files in the diff that contain real values rather than placeholders.
FAIL on any match. Tell the user to unstage the file.
# Raw SQL/query string concatenation instead of parameterized queries
grep -nE '(query|exec)\(.*\+.*\)|`.*\$\{.*\}.*(SELECT|INSERT|UPDATE|DELETE)' <changed-files>
# Shell command construction from variables
grep -nE '(exec|spawn|system)\(.*\$\{|child_process' <changed-files>
# Unescaped regex built from user input
grep -nE 'new RegExp\(.*\binput\b|new RegExp\(.*\bparams?\b' <changed-files>
FAIL on string-concatenated queries or shell commands built from variables that could carry user input — these are the classic injection vector regardless of language/framework. WARN on dynamic RegExp construction from anything that looks like user input (potential ReDoS or injection depending on context) — verify the input is sanitized/bounded first.
For changed frontend files:
grep -nE 'dangerouslySetInnerHTML|\.innerHTML\s*=|v-html|\{\{\{.*\}\}\}' <changed-files>
WARN on any match — confirm the rendered content is sanitized (e.g. via a known sanitization library) before allowing it. Don't assume it's fine just because it predates this diff if the diff touches that line.
For changed backend route/procedure/controller files:
grep -nE '(router\.|app\.)(get|post|put|delete|patch)\(' <changed-files>
For each new or modified route/procedure found, verify it has an auth/permission check (this repo's standard middleware, guard, or protected-procedure wrapper — check an existing route in this repo for the convention). Look for a middleware decorator (e.g. @Auth(), @UseGuards()), a wrapper function (e.g. requireAuth, checkPermissions()), or a guard wrapping the route handler. If none found, FAIL. FAIL if a new route handling user/org-scoped data has no auth wrapper at all. WARN if it has some auth check but the scoping looks broader than the data being touched (e.g. checks "is logged in" but not "owns this org/record").
# Sensitive fields potentially returned in an API response or logged
grep -nEi '\b(ssn|password|card_?number|cvv|api_?secret)\b' <changed-files>
WARN on any match in a response-shaping function or a console.log/logger call — confirm the field is masked, excluded, or this is a false-positive variable name unrelated to actual sensitive data.
# Spreading raw request input directly into a DB write
grep -nE '\.(create|update|insertOne|updateOne)\(\s*\{?\s*\.\.\.\s*(req\.body|input|params)' <changed-files>
FAIL on any match — writes should use an explicit, validated field allowlist (e.g. a schema-validated input type), never a raw spread of unvalidated request input.
# Overly permissive CORS, disabled TLS verification, debug flags left on
grep -nE "Access-Control-Allow-Origin.*\*|rejectUnauthorized:\s*false|NODE_TLS_REJECT_UNAUTHORIZED" <changed-files>
FAIL on any match — these disable a real security boundary and are rarely intentional in committed code.
Severity roll-up for this step: any 3a/3d/3f/3g FAIL blocks the commit. 3b's injection FAIL also blocks; its ReDoS WARN does not. 3c/3e WARNs don't block but must be addressed or explicitly justified before commit.
For each changed source file, grep for:
grep -nE 'console\.(log|error|warn|debug)\(' <file>
grep -nE '(debugger;)' <file>
grep -nE '(\/\/ ?TODO[: ]REMOVE|\/\/ ?FIXME[: ]REMOVE|\/\/ ?HACK)' <file>
if (__DEV__)-style guard) → OK, skip.debugger statement (or this repo's equivalent breakpoint call) → FAIL.TODO REMOVE / FIXME REMOVE → WARN.Apply COMMENTS.md to the comment lines this diff adds — the cheapest moment to stop a diary comment is before it is committed. Same classes, same commentPolicy, same never-flag list, same severity table, reported with the exact trim.
For each changed file in a typed language:
grep -n ': any\b\|as any\b' <file> # or this language's equivalent escape hatch
grep -n '!\.' <file> # non-null assertion, if the language has one
*.test.ts, *.spec.ts, *.d.ts, files in shims/, __mocks__/, or types/ directories.localeDictionaryPath is configured and the diff touches locale files)Scoped to only the keys actually touched by this diff — not a whole-dictionary sweep (that's /toolkit-qa's job). Skip entirely if the diff doesn't touch the interface/type file or the dictionary data file.
/toolkit-qa's Section B2 for the technique) has a value for that key in every locale block. FAIL on any key present in the diff's interface changes with no corresponding dictionary value in one or more locales — this is the orphaned-key defect that breaks the whole repo's typecheck.repoRules is configured)Read the repoRules file and run whatever banned-pattern/required-pattern checks it lists. Typical categories to expect in such a file (adapt to what's actually documented, don't assume all apply):
fetch/axios/equivalent) directly in view/component code; they belong in a service/store layer.Report each violation as <file>:<line>: <rule violated>, citing which rule in repoRules it came from.
Skip this step entirely if the repo is single-surface.
For any changed file that is NOT the platform-specific variant of a shared screen, grep for platform-specific API usage without a guard. WARN if found unguarded.
If a changed file for one platform imports something that only exists for another platform and this repo has a documented shim/equivalent directory, FAIL if no shim exists yet for that import. E.g. if the repo has web/ and native/ layers with a shims/ folder, check if the missing platform variant exists there; if not, FAIL.
git diff --name-only -- <this repo's shared utils/stores/services paths, if configured>
If shared utility/store/service files changed and this otherwise looks like single-surface work, WARN with each path — confirm these changes are intentional and won't break the other surface.
For changed backend files:
For changed service-layer files:
# Code Review — <YYYY-MM-DD HH:MM>
Changed files: <N> (<breakdown by appRoots layer + docs/tooling>)
## Check Results
| # | Check | Result | Details |
|---|---|---|---|
| 1 | Lint | ✅ PASS / ❌ FAIL / ⏭ SKIPPED | … |
| 2 | Typecheck | ✅ PASS / ❌ FAIL / ⏭ SKIPPED | … |
| 3a | Secrets/credentials | ✅ PASS / ❌ FAIL | … |
| 3b | Injection (A03) | ✅ PASS / ❌ FAIL / ⚠ WARN | … |
| 3c | XSS / unsafe rendering (A03) | ✅ PASS / ⚠ WARN / ⏭ n/a | … |
| 3d | Broken access control (A01) | ✅ PASS / ❌ FAIL / ⚠ WARN / ⏭ n/a | … |
| 3e | Sensitive data exposure (A02) | ✅ PASS / ⚠ WARN | … |
| 3f | Mass assignment (A04/A08) | ✅ PASS / ❌ FAIL / ⏭ n/a | … |
| 3g | Security misconfiguration (A05) | ✅ PASS / ❌ FAIL | … |
| 4 | Debug leftovers | ✅ PASS / ⚠ WARN | … |
| 5 | Type-safety quality | ✅ PASS / ⚠ WARN | … |
| 5.5 | Locale dictionary hygiene | ✅ PASS / ❌ FAIL / ⚠ WARN / ⏭ n/a | … |
| 6 | Repo architecture rules | ✅ PASS / ❌ FAIL / ⏭ not configured | … |
| 7 | Cross-platform safety | ✅ PASS / ⚠ WARN / ⏭ n/a | … |
| 8 | Backend/network patterns | ✅ PASS / ⚠ WARN / ⏭ n/a | … |
## Failures (must fix before commit)
<file>:<line>: <description>
## Warnings (review — may commit with justification)
<file>:<line>: <description>
## Logic / architecture notes
<Free-text observations about design, correctness, or patterns that don't fit neatly into the above checks. Keep to ≤5 bullet points. Be specific.>
## Verdict
✅ REVIEW PASSED — N warnings. Drafting commit message now…
OR
❌ REVIEW FAILED — N failures, N warnings. Fix failures and re-run `/toolkit-code-review`.
On FAIL, stop. Do not draft a commit message.
On PASS (zero failures), immediately continue and draft the commit message so the user gets both in one command.
If commitMessageConvention is configured, read it and follow it exactly. Otherwise, default to Conventional Commits:
<type>(<scope>): <imperative description>. If a ticket key is identifiable, this repo may prefer leading with it (e.g. <TICKET> - <type>(<scope>): …) — check recent git log for the actual house style and match it rather than assuming.git log --format=%B -5 and look for trailers like Co-Authored-By:, Signed-off-by:, Closes:. Replicate the exact format. Check recent commits for whether Co-Authored-By or similar trailers are used here before adding or omitting one. Drafting the message never stages or commits on its own — that's the final step, and only on request.Types:
| Intent | Type |
|---|---|
| New user-visible capability | feat |
| Bug fix | fix |
| Visual/layout polish, no behavior change | style |
| Code restructure, no behavior change | refactor |
| Undo a prior commit's change | revert |
| Dependencies, tooling, scripts, build config | chore |
| Documentation only | docs |
| Tests only | test |
| Measurable performance change | perf |
When in doubt between feat and refactor: if a user would notice, it's feat; if only a developer would notice, it's refactor. style = visual/layout polish with no behavior change; chore = tooling or dependencies. If both apply, prefer style if user-visible, else chore.
Scopes: derive from appRoots layer names touched (e.g. a layer named frontend → scope frontend, or follow whatever scope convention recent commits in this repo already use). If changes span multiple scopes for one logical task, pick the primary scope and mention the cross-cutting nature in the body if needed.
Group changes by intent. A "logical task" is a single concern a reviewer can verify in isolation (bug fix, feature, refactor, style pass, doc update, revert). If the diff contains more than one logical task, list each with its own suggested message and note which files belong to which task — the user can stage selectively.
Optimised for copy-paste. Emit ONLY:
## Task N — <one-liner> heading (only when there's more than one task).Changes: bulleted body if the change has multiple sub-parts. Single-concern commits: drop the body entirely.Bullet body rules:
- <area / file label> - <one short clause of context>. Keep each bullet to ~8–15 words.After all tasks, if the security scan (step 3, any sub-check) surfaced anything worth flagging, append one short **Flags**: paragraph — single short sentence per flag.
Runs ONLY when the user explicitly asks to commit (e.g. "commit those files", "commit and move the ticket", /toolkit-code-review --commit). A bare /toolkit-code-review stops after step 10. Steps 1–10 are always read-only; this step is the one place the skill writes (a commit) and mutates external state (a ticket transition), and only after a PASS.
git add ONLY those and commit just them. Leave every other modified/untracked path untouched — the working tree often holds unrelated in-progress work that must not ride along.git push unless the user explicitly asks. "Commit" does not imply "push".git log --oneline -1 afterward so the user sees the SHA.jiraCloudId is configured)Resolve the ticket key in priority order:
ticketPrefix) in the commit message/subject.Transition via the Atlassian MCP:
getTransitionsForJiraIssue for the ticket — never hardcode a transition ID, workflows vary per project and can change.If the MCP is not available, or jiraCloudId is not configured, skip silently — do not error.
git push or git reset --hard unless the user explicitly asks. Committing when asked is fine; pushing is not implied.repoRules/multi-surface config), mark it ⏭ SKIPPED.git show HEAD:<path> + the diff over re-reading the current disk state (avoids noise from unrelated changes)./toolkit-code-review, review my changes, review before commit, check my code, ready to commit?, pre-commit review, review + commit and move the ticket, /toolkit-code-review --commit, /toolkit-code-review <files…>, /toolkit-code-review <TICKET-KEY>
/toolkit-create-prCreates a pull request from the current branch into this repo's configured base branch. Reads git log and diff, writes a structured PR title and body, then opens the PR on GitHub. Run this when you're ready to ship a feature branch.
You are the PR creator. When invoked, you inspect the current branch, build a PR description from the diff and commits, and create the PR targeting this repo's configured base branch.
No config file to set up — values are detected from the repo or asked once and remembered in ~/.dev-skills.env (see CONFIG.md for the full read/write procedure and category definitions):
baseBranch — detected once with confirmation, then written — the branch new work branches from and PRs target. Detect git's actual default branch, show it, and let the user confirm/override before writing — this repo's actual PR-target convention may differ from its git default branch.githubRepo — detected fresh every run — org/repo for gh pr create, parsed from git remote get-url origin.ticketPrefix — asked once, then written — used to derive the JIRA ticket from the branch name.appRoots — detected fresh every run — used to scope the quality-gate checks to whichever layers actually changed.qualityGate — detected fresh every run — command lists read from each layer's manifest scripts, run before opening the PR.jiraCloudId — detected once via live MCP lookup, then written — if available and the Atlassian MCP is reachable, Step 6 posts a JIRA comment. If the lookup fails, Step 6 is skipped silently.qaTestUrl — asked once, then written, optional — a staging/dev URL QA can use to manually verify the change. If the user says there isn't one, the "How QA Can Test" comment omits the Prerequisites URL line and just lists the steps.Run these commands in parallel (substituting <baseBranch> from config):
git branch --show-current # current branch name
git log <baseBranch>..HEAD --oneline # commits ahead of base
git diff <baseBranch>...HEAD --stat # files changed summary
git diff <baseBranch>...HEAD # full diff for description
If the branch has no commits ahead of <baseBranch>, stop and tell the user:
"Nothing to PR — this branch has no commits ahead of
<baseBranch>."
Based on the --stat output from Step 1, determine which appRoots layers were touched, then run that layer's detected qualityGate commands. For example, if appRoots.frontend files changed, run every command detected for the frontend layer; if appRoots.backend files changed, run the backend layer's.
If CI runs a real build (not just a typecheck) on certain layers, prefer running the equivalent local command (e.g. tsc --noEmit) if it catches the same errors without writing build output — but only if the repo's qualityGate config specifies that; don't assume that equivalence yourself.
Run every test command listed, even if this PR's own changes didn't add new tests — regression tests already in the suite must stay green.
Distinguish a missing script from a real failure. If a command errors with "missing script"/"command not found" rather than actually running, that's a config/setup issue, not a quality-gate failure — stop and tell the user "qualityGate.<layer> lists <command>, but it doesn't exist in this repo" rather than reporting it as a failed check in the PR body.
If any command runs and fails for a real reason: stop. Do not create the PR. Show the errors to the user and either fix them (if the fix is small and obviously correct) or ask how to proceed. A red CI build is more disruptive to fix after merge than before the PR opens. A test failure is the same severity as a lint/typecheck failure here — it means this PR broke previously-verified behavior.
If all commands pass: record the results — they go into the "Test plan" checklist in Step 4 as already-completed items ([x], not [ ]).
Extract the ticket number from the branch name (matching the remembered ticketPrefix, e.g. PROJ-825 from branch PROJ-825).
If the branch name doesn't contain a ticket number, check the most recent commit message.
If still unclear, ask the user: "Which ticket does this PR close?"
Format: [<TICKET>] <concise description of what changed>
Rules:
Examples:
[PROJ-825] Add MS_Sign_On and SAML2 custom-login parity to login screen[PROJ-826] Audit admin dashboard — all widgets verifiedUse this template exactly, substituting the actual appRoots layer names this PR touched in place of the generic placeholders:
## Summary
• <bullet: what this PR does — 1 sentence>
• <bullet: any backend-layer changes, if appRoots.backend touched>
• <bullet: any frontend-layer changes, if appRoots.frontend touched>
• <bullet: any config/env changes, if applicable>
## Ticket
<TICKET-NUMBER> — <ticket title>
## Test plan
- [ ] <specific thing QA should verify>
- [ ] <specific thing QA should verify>
- [x] <quality-gate command> passes (<layer>, if touched — ran in Step 1.5)
- [x] <quality-gate command> passes — <N> passed (<layer>, if touched and a test command exists — ran in Step 1.5)
## Files changed
<paste the --stat output here>
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Rules for the body:
git diff and commit messages — be specific, not generic[x] for a suite that didn't run.Confirm gh is authenticated before pushing anything:
gh auth status
If this fails, stop here — don't push first and discover the auth problem at the gh pr create step. Report: "Run gh auth login first, then re-run /ai-skills:create-pr."
Check if the branch has a remote tracking branch:
git status -sb
If no upstream, push with: git push -u origin <branch>
If already tracking, push with: git push
Create the PR:
gh pr create \
--base <baseBranch> \
--title "<title>" \
--body "$(cat <<'EOF'
<body>
EOF
)"
Return the PR URL to the user.
jiraCloudId is configured)If the ticket number was identified and the Atlassian MCP is available, post a comment using mcp__claude_ai_Atlassian_Rovo__addCommentToJiraIssue with this exact structure:
Dev Update — <one-line description of what this PR does>
PR: <PR URL>
---
**What Was Fixed / Implemented**
| File | Fix | Detail |
|------|-----|--------|
| <file path> | <short label> | <what changed and why — be specific> |
| <file path> | <short label> | <what changed and why> |
(one row per meaningful change — group by file, skip trivial changes)
---
**Commits**
• `<short hash>` — <commit message>
• `<short hash>` — <commit message>
---
**How QA Can Test**
Prerequisites: (omit this block if qaTestUrl is unset)
- Open <qaTestUrl>
- Log in with a staging account
Test 1 — <test name>
1. <step>
2. <step>
3. Confirm: <what to verify>
Test 2 — <test name>
1. <step>
2. <step>
3. Confirm: <what to verify>
(continue for each distinct area touched by the PR)
- [x] <quality-gate command> passes with zero warnings/errors
- [x] <quality-gate command> passes — <N> passed (<layer>, if touched)
Rules for the comment:
git diff — one row per meaningful change, grouped by file. Skip trivial changes (imports, whitespace).git log <baseBranch>..HEAD --oneline — use short hashes and the actual commit messages.cloudId: the jiraCloudId value detected/remembered for this repo.contentFormat: markdownIf the MCP is not available, or jiraCloudId is not configured, skip silently — do not error.
baseBranch — never push to the repo's default/production branch directly unless baseBranch IS that branch.<baseBranch> into <baseBranch>. If the current branch IS the base branch, stop and tell the user.gh is not authenticated, report: "Run gh auth login first, then re-run /ai-skills:create-pr."/toolkit-export-auditAudits an export button (CSV/Excel download) on a built screen against this repo's configured reference implementation. Checks column names, order, casing, scope, and file name — then fixes any discrepancy. Run this whenever you add or review an export button.
You are the export-audit agent. Your job is to find an export button on a named screen, look up the matching export implementation in this repo's configured reference, and make this repo's export an exact copy — same columns, same casing, same scope, same file name (modulo an acceptable file-format divergence, e.g. CSV vs. XLSX).
This skill is invoked with /ai-skills:export-audit <screen name>.
No config file to set up — values are detected from the repo or asked once and remembered in ~/.dev-skills.env (see CONFIG.md for the full read/write procedure and category definitions):
referenceSource — asked once, then written, required for this skill — type: "repo" is the only mode this skill supports meaningfully (comparing against another codebase's export implementation). If nothing's been asked/answered yet, or the answer isn't a "repo"-type reference, ask the user where the reference export implementation lives — this skill needs one to function.appRoots — detected fresh every run — used to locate this repo's screen file.qualityGate — detected fresh every run — lint/typecheck commands to run after fixing.Before running any audit, check both sides for an export button/handler. This determines whether the rest of this skill should run at all:
grep -n "export\|Export\|csv\|CSV\|xlsx\|download\|Download\|ExportTo\|onExport" <this repo's screen file>
grep -n "export\|Export\|ExportToExcel\|download\|buttons:" <referenceSource.path>/<matching script/component path>
Four outcomes:
| This repo | Reference | Outcome |
|---|---|---|
| ❌ no export | ❌ no export | Stop here. Report: "No export button on either side — nothing to audit." Do not proceed to Step 1. |
| ❌ no export | ✅ has export | Missing export — proceed to Step 1/2 to derive the reference's shape, then build it from scratch. |
| ✅ has export | ❌ no export | Flag, don't auto-remove. This repo has an export the reference doesn't. Report it as a ⚠️ divergence and ask the user whether it's intentional (a deliberate addition) before treating it as something to fix — removing a feature is a product decision, not a parity-audit action this skill takes unprompted. |
| ✅ has export | ✅ has export | Normal case — proceed to the full audit (Step 1 onward). |
This gate is what makes /ai-skills:export-audit <screen name> safe to run speculatively — e.g. from /ai-skills:qa or a screen-build pipeline — without it doing unnecessary work or producing a confusing report on a screen that was never meant to have an export.
(Only reached if Step 0 found an export on at least one side.)
Find the screen file under the relevant appRoots layer — reuse the grep result from Step 0 rather than re-running it.
If no export button or handler is found on this repo's side here, check the reference screen (Step 2). If the reference has an export button and this repo does not, that is a missing export — implement it.
Grep first, never read from line 1.
To find the matching file: grep the referenceSource path for the screen name or export function name first (grep -rl 'export\|Export' <referenceSource.path>/), then narrow to the matching file before reading.
grep -n "export\|Export\|ExportToExcel\|download\|buttons:" <referenceSource.path>/<matching script/component path> | head -30
Then read only the export function by targeting the offset returned by grep — read from the function declaration to its closing brace, not the whole file.
Extract the following from the reference export:
| Detail | What to look for |
|---|---|
| Headers | The header row/array pushed first, e.g. [{text:"FIRST NAME"},{text:"LAST NAME"},…] |
| Column order | The exact left-to-right order of fields in each data row |
| Column values | How each field is read/derived: direct field access, nested lookup, computed/ternary status label, etc. |
| Export scope | Does it use the on-screen table's current (filtered/sorted) data, or does it fetch fresh, unfiltered data? |
| File name | The output filename — e.g. "users.xlsx" |
| Status/label strings | Exact label text and the codes/values they map from |
Compare this repo's export function against the reference.
Do not assume a casing convention — derive the exact header strings from the reference and copy them verbatim.
| This repo's header | Reference header | Casing match? | Format note |
|---|---|---|---|
| (each header) | (read from reference) | ✅ / ❌ |
| Position | This repo's column | Reference column | Match? |
|---|---|---|---|
| 1 | ✅ / ❌ | ||
| … |
| Column | In this repo? | In reference? | Action |
|---|---|---|---|
| (one row per column found in either side) | ✅ / ❌ | ✅ / ❌ | add / remove |
| Column | This repo's value expression | Reference value expression | Match? |
|---|---|---|---|
| (each column) | ✅ / ❌ |
| Detail | This repo | Reference | Match? |
|---|---|---|---|
| Source data | filtered/on-screen subset vs. full dataset | (same dimension) | ✅ / ❌ |
| Respects active search/filter? | yes / no | yes / no | ✅ / ❌ |
If the reference's export always pulls a fresh, full dataset (bypassing whatever search/sort is currently applied on screen), this repo's export must match that scope — flag it if it's currently exporting the filtered subset instead.
| This repo | Reference | Match? |
|---|---|---|
| ✅ / acceptable divergence if only the extension differs |
A different file extension (e.g. .csv vs. .xlsx) due to a library/format constraint is an acceptable, explicitly-noted divergence. The base name should still match. If base names differ (e.g. reference is users.xlsx and this repo is user-list.csv), flag as ⚠️ WARN with both names shown — do not auto-fix a filename divergence without confirming with the user.
Check that every column value field referenced by the export is actually declared on the relevant TypeScript/type definition this repo uses for that data:
grep -n "type <DataType>\|<field1>\|<field2>" <screen-file>
If a field used in the export is missing from the type, add it.
After the audit, fix every ❌ item. Do not stop at the report — implement the corrections in the screen file. For complex cases (e.g. reference fetches fresh data server-side but this repo uses on-screen filtered data), describe the required refactor and ask the user to confirm the approach before implementing.
Common fix shapes:
Headers don't match the reference — wrong casing or missing columns: Always derive the correct header strings by reading the reference's header array for this specific screen. Do not assume the headers are the same as another screen's export.
Exporting the filtered/on-screen subset instead of the full dataset (when the reference's scope is "always fresh, full data"): Switch the data source to the full/unfiltered set, matching the reference's scope.
Missing columns: Add them even if the value is frequently empty for many rows — presence parity matters, not just non-empty-value parity.
Missing field on type: Add the field to the type definition backing the export's row-mapping function.
Run this layer's qualityGate commands from config (typically lint + typecheck). Fix all warnings/errors before reporting done.
Export Audit — <Screen Name>
══════════════════════════════════════════════════════
Screen: <path>
Reference: <referenceSource path>/<matching file>
Presence check: ✅ both have an export / ➖ neither has an export — nothing to audit
/ ⚠️ this repo has one, reference doesn't / ❌ reference has one, this repo doesn't
(If presence check is ➖, stop here — omit every section below.)
Export function: <function/handler name>
──────────────────────────────────────────────────────
AUDIT RESULTS
──────────────────────────────────────────────────────
Header casing:
✅ / ❌ <detail>
Column order:
✅ / ❌ <detail>
Missing columns:
✅ none missing / ❌ <list missing columns>
Extra columns (not in reference):
✅ none / ⚠️ <list>
Column values:
✅ / ❌ <detail per column>
Export scope:
✅ matches reference scope / ❌ <detail>
File name:
✅ / ⚠️ <detail>
Type coverage:
✅ all fields declared / ❌ missing: <list>
──────────────────────────────────────────────────────
FIXES APPLIED
──────────────────────────────────────────────────────
<bullet list of what was changed>
──────────────────────────────────────────────────────
QUALITY GATE
──────────────────────────────────────────────────────
<command>: ✅ pass / ❌ N issues
──────────────────────────────────────────────────────
RESULT: ✅ EXPORT MATCHES REFERENCE / ❌ ISSUES REMAIN
Thanks to the Step 0 presence gate, this skill is safe to invoke unconditionally — it self-determines whether there's anything to audit and exits quickly if not.
/ai-skills:qa — call it as part of a structural/parity audit without needing to pre-check for an export button yourself./toolkit-figmaTurn a design (pasted screenshots, or a fetched Figma frame if this repo has Figma MCP configured) into a cached spec file so build/QA skills can read it instead of re-parsing the design every time. Use when the user says "figma <screen name>", "/toolkit-figma <name>", "spec the figma frame", or pastes screenshots and asks for a spec.
Read a design (screenshots, or a Figma MCP fetch if configured) and produce a single spec file the rest of the workflow consumes. Write only one file (the spec). Do not touch app code.
No config file to set up — values are detected from the repo or asked once and remembered in ~/.dev-skills.env (see CONFIG.md for the full read/write procedure and category definitions):
appRoots — detected fresh every run — where to look for the screen's existing implementation, if any, so the spec can link to it.specCachePath — detected fresh every run — directory to write specs into (e.g. .claude/figma-specs/). Checks what's already in use under .claude/; defaults to .claude/specs if nothing found yet.localeConvention — detected once, then written — if this repo has an i18n hook/dictionary convention, the spec's locale-strings section is phrased against it. If nothing's detectable yet, that section just lists literal strings without proposed keys.figmaMcpAvailable — detected live, never cached — checks whether Figma MCP tools are present in the current session. Detect by checking the current Claude session's tool list for any tool matching *figma* or *Figma*. If none found, set to false. If not available, assume no MCP and require pasted screenshots.FIGMA_API_TOKEN — read from ~/.dev-skills.env, only when figmaMcpAvailable is true. Used to authenticate the Figma frame fetch.
If figmaMcpAvailable is true but FIGMA_API_TOKEN is missing or empty, do not error or block silently. Tell the user explicitly:
figmaMcpAvailableis set butFIGMA_API_TOKENisn't in~/.dev-skills.env— add it, or fall back to pasted screenshots for this run.
Then proceed with the screenshot-only path for the current invocation rather than refusing to continue.
If the token is present but returns a 401/403 from the Figma API, tell the user: "Your FIGMA_API_TOKEN in ~/.dev-skills.env appears invalid. Update it manually and re-run." Do not clear the key automatically.
Never print, log, or write the token's value anywhere — not into the spec file, not into chat output. Read it once, use it for the fetch call, discard it.
Required:
figmaMcpAvailable is true and FIGMA_API_TOKEN is present in ~/.dev-skills.env, a Figma URL/node reference to fetch directly. If either the config flag or the token is missing, fall back to requiring pasted screenshots.Optional:
If neither screenshots nor a fetchable Figma reference is available, stop and ask once. Don't guess at layout from a description alone. "Once" means once per invocation — if the user still can't provide input on the next invocation, ask again. Do not remember a 'no input' state.
Lowercase, hyphenated, no spaces. Match this repo's existing screen-folder naming convention if one exists (check how an existing screen under appRoots is named).
Look in specCachePath for <slug>.md (or this repo's spec-cache convention if specCachePath isn't set but a .claude/ spec folder already exists — check before assuming neither exists). If it exists:
Find where this screen lives (or would live) under the relevant appRoots layer so the spec can link to it. If the screen doesn't exist yet, note that — this repo's own build skill will create it.
Walk the design top-to-bottom and capture:
localeConvention is set)For each section that needs server data, look up the data-fetching hook/call that would feed it — grep this repo's data-layer convention for something already matching. For a missing one, note "Server: needs new endpoint" — a hand-off item, not something this skill builds.
localeConvention is configured)Determine which locale dictionary section the new strings belong to. List every string from the design in a "Locale strings" section keyed by the proposed dictionary key. If localeConvention is unset, list the literal strings without proposed keys and note that locale wiring isn't configured for this repo.
For each card/table/panel, point at an existing shipped component in this repo to reuse, rather than inventing new structure. Use the same language verbatim in the spec so the build skill doesn't invent new patterns.
<specCachePath>/<slug>.md — use this template:
# <Screen Name> Spec
Generated by: /ai-skills:figma on <YYYY-MM-DD>
Source: <Figma URL(s) for human reference, or "pasted screenshots">
## Repo location
- File: <path under appRoots, or "does not exist yet">
## Page header
- Title: "<text>" — locale key: `<block>.<key>` (omit key if no localeConvention)
- Sub-content: <if any>
- Actions (top-right): <list>
## Layout sections
### <Section Name>
- Elements:
- <element>: <text or data field> (icon: <name>, locale: `<block>.<key>`)
- Data: <hook/call> — endpoint <verb> <url> — shape `{ <fields> }`
- Empty state: "<text>"
- Loading state: skeleton | spinner | nothing
<!-- repeat per section -->
## Tables
### <Table Name>
| Column | Source field | Type | Sortable | Notes |
|---|---|---|---|---|
| ... | ... | ... | ... | ... |
- Pagination: page-based | cursor
- Default sort: <field> <asc|desc>
- Row click: opens <panel> | navigates to <route>
## Filters
- <name>: <type> — default: <value> — source: <query/hook>
## Locale strings (omit if no localeConvention)
Block: `<dictionary-block>`
: ""
## Visual baseline reuse
- <component category>: match <existing shipped screen/component> — <specific style notes>
- Theme colors only — never invent a new palette; pull from this repo's configured theme convention.
## Cross-references
- Related shipped pages (style mirror): <list>
## Open questions
- <anything you couldn't pin down from the design — needs product/design clarification>
After writing the file:
/ai-skills:code <slug> to build it (or this repo's own local build skill, if it has one).figmaMcpAvailable isn't true. Ask for screenshots instead.Tell me which screen to spec, and share the design (screenshots, or a Figma reference if MCP is configured).
/toolkit-fleetPuts the Unified-Brain fleet — Joshua's seven-role agent team, run under Scar's enforcement gates — to work on a real task in the repo you're working in, on its own branch in a kept worktree, never committing. Also runs the fleet's gates-on vs gates-off experiment on request. Detects the Unified-Brain checkout, dry-runs first in any new repo, states the exact cost shape and waits for an explicit yes before spawning anything, and never decides on its own to deploy the fleet — an "on" flag is permission, not instruction. Use when the user says "/toolkit-fleet", "put the fleet on this", "have the team do X", "fleet work", "run the fleet", "run a fleet experiment", "gates-on vs gates-off", or "is this repo fleet-ready".
Two jobs, one roster (<FLEET_ROOT>/fleet/AGENTS.md, "Two jobs, one roster"):
| you want | driver | what happens |
|---|---|---|
| the team to do a task in this repo | work.mjs |
one lead (the orchestrator by default) starts in a worktree branched off this repo's HEAD, dispatches the other roles only as the task needs, and leaves the result on a branch for you to review. Nothing is committed. |
| the gates measured | run.mjs + compare.mjs |
canned fixtures in fleet/tasks.json, two arms, worktrees off the Unified-Brain checkout, scored mechanically. Does not touch this repo. |
Work mode is the default reading of "run the fleet on this." The experiment is only what the user asks for by name ("experiment", "gates-on vs gates-off", "compare arms").
You never decide, on your own, to deploy the fleet. An agent that can choose to fan out into a paid multi-agent run will, because eligibility reads as instruction the moment it exists. The cost is real token spend, it is metered, and it is the user's, not yours to commit on a hunch.
You MAY evaluate one mechanical threshold and PROPOSE: more than 5 files in scope and the work is write-heavy. Both, not either. If they hold, name them, then stop and ask — never spawn. A standing habit, a prior clean run, or the user once saying "use the fleet when it's useful" is a permission, not an instruction. Re-answer "does this task, right now, warrant the cost" from the task in front of you, every time.
FLEET_ROOT — the Unified-Brain checkout. Shared, no repo prefix (one checkout serves every
repo, like FIGMA_API_TOKEN in CONFIG.md). Detect it from this file's own
realpath (brain/core/skills/toolkit/fleet/SKILL.md lives inside the checkout; walk up to the
git root and confirm fleet/work.mjs exists there). Ask once only if that fails; write the
answer to ~/.dev-skills.env. Before trusting a stored value, confirm <value>/fleet/work.mjs
still exists.fleet/lib/codemap.mjs with dekko on the run's
worktree, not by you. The dry run shows the last summary, labelled with its age; a real run
rebuilds it. If the banner says dekko not found, tell the user the fix is one line,
uv tool install dekko, and that the run still works without it. The owner may keep
fleet/maps/<slug>/notes.md — hand-written, gitignored — and it goes into the prompt whenever
present. --no-map skips the summary but never the notes.<SLUG>_* lines — read by fleet/lib/conventions.mjs, not by you. The driver
appends them to the lead's prompt as REPO CONVENTIONS (locale hook, theme object, base branch,
reference repo — whatever the other toolkit skills have recorded). You do not paste them into the
task; you check the dry-run's --show-prompt output carries them.FLEET_ROOT resolves and <FLEET_ROOT>/fleet/work.mjs, run.mjs, package.json, agents/ exist.git rev-parse --show-toplevel. The isolation is
git worktree add … HEAD; no git, no worktree, no fleet.node --version resolves.cd <FLEET_ROOT>/fleet && npm run selftest exits 0. It drives both drivers in --dry-run and
spawns nothing.node -e "import('<FLEET_ROOT>/fleet/lib/cli.mjs').then(m=>{const b=m.resolveCli();console.log(m.cliVersion(b), b)})"
It tries FLEET_CLAUDE_BIN, then claude on PATH, then the newest VSCode extension's bundled
binary. "Not on PATH" is not "not resolvable"; on 2026-09-14 the extension binary resolved fine.git -C <target> status --porcelain | wc -l and the current branch.Report it in one block, every line naming a check that ran:
fleet ready
roster 7 roles — 🧭 orchestrator · 🔨 implementer · 🎨 ui-designer · 🔍 auditor · 🧪 qa-verifier · 📚 researcher · 💥 adversary
harness selftest <n>/<n>
cli <version from the resolver>
target <repo> (<branch> @ <sha>), <n> uncommitted change(s)
conventions <n> line(s) for this repo (from the dry run's banner)
map <the dry run's map line: last summary + age, "none yet", or "dekko not found">
Any failed check is "fleet not ready" with the failing check and its fix. Readiness is a statement about the harness, not permission to run.
cd <FLEET_ROOT>/fleet
npm run work -- --repo <target> --task "<the task in the user's words>" --dry-run --show-prompt
Free; spawns nothing; creates no worktree. Read the banner and the shown prompt back to the user:
the lead and model with its tier, which roles are available, how many convention lines went in,
the wall-clock cap, and the carries line. The worktree starts from a snapshot of the user's working tree, not from
HEAD — modified, untracked and deleted files are carried in as uncommitted work, because in this
workflow a commit means "verified" and the work to continue is uncommitted by definition. Never
tell the user to commit first. What is not carried: gitignored files (.env, build output,
node_modules). If the task depends on one, the choices are --in-place (no worktree, no
isolation, agents run with permissions bypassed against the real tree — say that out loud and get
a yes) or none. --from-head exists for the rare task that should start clean.
For a long brief, write it to a file and pass --task-file <path> instead of --task. A brief
that names files, the acceptance criterion, and what is out of scope gets a better team than a
sentence does; the orchestrator's own role file says why the delegation message is the whole
contract.
Before any spawn, say:
--look for read-only investigation, researcher on
sonnet with the write tools denied. --fix for one scoped edit whose brief already names the
files and the acceptance criterion, implementer on sonnet, no crew, and the user reviews the
patch. --feature when the work must be split across files or checked by someone other than the
writer, orchestrator on opus with the full roster and a 2400s cap. The banner prints the lead's
tier and, on a capable-tier lead, the cheaper alternative — relay that line if it appears.model: from each <FLEET_ROOT>/fleet/agents/<role>.md;
never assume. --roles a,b narrows the set. Availability costs nothing; a role bills only when
the lead dispatches it, and the lead is instructed to use the smallest set and say which.--look and --fix, 2400s for --feature. Real
runs have taken 378s and 1907s. A run that hits the cap keeps whatever landed in the worktree and
its report is labelled PARTIAL. --timeout overrides.<role:model>, can dispatch <role:model …>, up to <timeout>, real API spend. Proceed?"Then wait. Nothing runs past this point without a yes.
npm run work -- --repo <target> --task "…" # or --task-file
Live view, read-only, loopback only: npm run dashboard → http://127.0.0.1:7777. While a run is
going the working now panel shows the active agent, the tool it is running, the lead's latest
words and a tail of recent activity — point the user at it rather than narrating the run yourself.
Finished runs show their report.md inline; raw transcripts are never browsable, and tool inputs
are masked for credential shapes. The work panel lists every work run: running (the driver's pid answers), stalled (it does not, and no
record was written), done, timed out, or dry run — with gate activity and files changed read live
from the run's own directory. It never shows the task text or the report.
The driver prints the worktree path and branch (fleet/<slug>/<timestamp>-work), gate activity
(ran / fired / denials), the file count changed, and two paths: report.md (the lead's final
message, written before anything is printed) and transcript.jsonl (raw evidence — never paste
or commit it). Relay:
report.md. The driver enforces it: a lead that ends without the block is asked once more in
the same session, and the file is labelled NO STATUS BLOCK if it still does not comply, or
PARTIAL if the run timed out (then it holds the last thing the lead said). Never summarise
"verified" up from "the implementer says it works"; the format keeps those apart on purpose.npm run land -- <runid> writes the team's own diff (snapshot → worktree)
to land.patch, checks it against the user's current tree, and prints the one-line git apply to
run. It changes nothing by default. That is deliberate: a tool that rewrites files across a
repo is correctly read as irreversible by a permission classifier, and a patch the user reads
first is better review anyway. --apply applies it (uncommitted, atomic, all or nothing, refuses
on conflict). The user's carried work is excluded by construction. Never run --apply unasked,
and never commit after either path.teamFiles is the snapshot→worktree diff, filesChanged is everything dirty
including the user's carried work. A run that edits two already-dirty files never moves the second
number, which is how one 2-file contribution got recorded as 13.Reviewing, landing, or discarding the branch is the user's. Never merge, never commit.
npm run recover -- <runid> rebuilds the report and the transcript from the CLI's own session
store, which lives outside runs/ and survives its deletion. It rebuilds the team's patch too,
because every run writes real git objects into the target repo and those outlive the worktree. With
work.json intact that is exact. With the record gone, add --repo <path> --from <snapshot-sha>
to search the object store; it prints candidate trees for the user to pick and --tree <sha> writes
that one's patch. Do not offer to guess the snapshot sha — the tool refuses to, because landing
against a wrong base corrupts a tree. It writes recovered.json, never work.json.
cd <FLEET_ROOT>/fleet
npm run experiment -- --dry-run # first, in any context this hasn't run from before
npm run experiment -- [--only <task>] [--roles a,b | --qa] [--arms gates-on] [--repeats n] [--lead <role>]
npm run experiment:compare [<runid>]
Cost shape before a real run: cells = tasks × arms × repeats (defaults: 3 × 2 × 1), lead
implementer:sonnet plus any named roles at their own declared models. compare.mjs will not
call a difference until n ≥ 30 per arm and provenance is uniform; a single run proves the harness,
not the hypothesis. It prints NO CLAIM — preconditions not met otherwise. CONTROL IS DIRTY
means the gates-off arm recorded a firing; TREATMENT IS EMPTY means gates-on recorded none —
in both cases the arm's settings never reached the CLI and nothing in the table means anything.
runs/ is raw evidence — never commit or publish itfleet/runs/ is gitignored and holds full, unsanitised agent transcripts and the prompts that
produced them. Client repos are in scope for work mode (fleet/AGENTS.md, 2026-09-14), on the
same terms as Scar's private project layer: transcripts, prompts, and branches stay on this
machine; nothing from a run is pasted, published, or made into a lesson without
brain/SANITIZATION.md first. The experiment never runs on a client repo; it has its own fixtures.
Nothing here stages, commits, or pushes, in the target or in <FLEET_ROOT>. No fleet role has
that authority either (fleet/agents/AGENTS.md). Work lands on a branch for a human.
FLEET_ROOT is the only copy. A vendored fleet/ would drift the moment either side changed —
the failure engine-parity polices for Scar's engine. Always cd <FLEET_ROOT>/fleet; never
copy its files anywhere.
--dry-run --show-prompt first in a repo this skill hasn't dry-run from.--in-place without saying what it gives up and getting a yes.<FLEET_ROOT>.fleet/. Run it from <FLEET_ROOT> in place.runs/./toolkit-fleet, "put the fleet on this", "have the team do X", "fleet work", "run the fleet",
"run a fleet experiment", "is this repo fleet-ready", "gates-on vs gates-off", "compare fleet arms",
"start the fleet dashboard"
/toolkit-jiraJIRA ticket generator and resume handler. Two modes — generate a ticket from an approved /toolkit-plan spec (or from scratch for something trivial), or resume implementation from a ticket pasted directly into the chat. Always shows a plan and waits for approval before executing any code. To re-analyze a ticket by key (fetch live, re-plan, post back to Jira), use /toolkit-analyze instead.
You bridge ticket tracking with implementation. You never write code directly — you generate tickets and hand off to /toolkit-code (or whichever local build skill this repo uses), after getting explicit approval.
No config file to set up — values are detected from the repo or asked once and remembered in ~/.dev-skills.env (user's home directory — shared across all repos, never repo-local) (see CONFIG.md for the full read/write procedure and category definitions):
jiraCloudId — detected once, then written — looked up live via mcp__claude_ai_Atlassian_Rovo__getAccessibleAtlassianResources, not asked.jiraProjectKey — asked once, then written — required for creating tickets.baseBranch — detected once with confirmation, then written — branch new ticket branches are cut from. Detect git's actual default branch, but always show it and let the user confirm/override before writing — this repo's actual PR-target convention may differ from its git default branch.referenceSource — asked once, then written — if this repo cross-references another source (a sibling repo, a Figma spec, a design doc) when scoping work, ask instead of assuming a specific reference workflow.appRoots — detected fresh every run — used to phrase the ticket's file paths in terms of this repo's actual layer names.specCachePath — detected fresh every run — where /toolkit-plan writes approved specs. Checked first in Mode 1 before deriving anything from scratch.buildSkillName — detected once, then written — the name of this repo's own local "go build this" skill, used in the handoff message (e.g. /build-screen). Detected by scanning .claude/skills/ for a genuinely local (non-symlinked) build-coordinator-shaped skill; if none found, hand off to /toolkit-code instead.jiraIssueType — fixed at Story — always create tickets as Story. Verify Story exists for this project via mcp__claude_ai_Atlassian_Rovo__getJiraProjectIssueTypesMetadata once jiraProjectKey is known; if this project has no Story type, stop and ask the user which type to use instead (don't silently fall back to Task).jiraAssigneeFieldId — asked once, then written, only if relevant — a custom field ID this Jira instance uses for a "Developer" field, if any. Most Jira projects don't need this; only ask if the user's normal assignment flow fails or they mention a custom field./toolkit-jira <unit of work>)Invoked when the user says:
/toolkit-jira <name>Steps:
specCachePath (or whatever spec-cache convention is already in use under .claude/) for a spec matching this unit of work, written by /toolkit-plan. If one exists and is marked approved, read it — it already has the scope, the parity decision, and the acceptance criteria nailed down. Skip straight to step 5 (the spec replaces steps 2-4 below entirely; don't re-derive anything /toolkit-plan already settled).<name> — run /toolkit-plan first to scope this properly, or describe it now and I'll do a lightweight version of that here." If they choose the lightweight path, continue with steps 3-4.5 below as a fallback. If the lightweight path is chosen, no spec file is written — proceed directly to ticket creation. /toolkit-plan is the recommended default, not a hard blocking dependency — small/trivial work can skip it.appRoots.find/ls/grep on the expected path(s).
4.5. If referenceSource is configured for this repo, ask whether parity applies to this specific unit of work — don't assume it does just because the repo has a reference source configured globally. Ask plainly: "Does this need to match <referenceSource description>, or is this a net-new screen/feature with no existing counterpart?" A repo-wide referenceSource (e.g. always comparing against a legacy app) doesn't mean every new screen has a counterpart there — plenty of new work is genuinely new.
/toolkit-qa doesn't later try to find a reference that was never meant to exist.If the unit of work already exists — switch to Audit mode:
qualityGate in config, if set) scoped to the file and capture any warnings/errors.If it does not exist — use the standard greenfield ticket template below.
Template selection criteria: If 70%+ of the required functions already exist in the codebase, use the Audit Template; if fewer than 30% exist, use the greenfield template; if in between, ask the user.
/toolkit-plan's spec (step 1) or step 4.5's fallback question — even when the spec format wasn't designed with this field in mind: add a top-line note: Parity check: required (vs. <referenceSource description>) or Parity check: not applicable — net-new, no counterpart. This is what lets /toolkit-qa skip its reference-parity audit cleanly instead of reporting a missing reference as a finding.
5.5. Confirm the Atlassian MCP is available before going further — steps 6-7 below depend on it and have no fallback. If mcp__claude_ai_Atlassian_Rovo__* tools aren't available in this session, stop here and ask: "Would you like me to generate the ticket description for you to paste into Jira manually, or wait until the MCP is configured?"mcp__claude_ai_Atlassian_Rovo__atlassianUserInfo — use the returned accountId for assignment. This makes assignment dynamic to whoever is running the skill, not a hardcoded person.mcp__claude_ai_Atlassian_Rovo__createJiraIssue:
cloudId: from configprojectKey: from configissueTypeName: Story (from config — see jiraIssueType above)assignee_account_id: <accountId from step 6>jiraAssigneeFieldId is configured, also set that field to {"accountId": "<accountId from step 6>"}baseBranch:
git checkout <baseBranch>
git pull origin <baseBranch>
git checkout -b <TICKET-KEY>
Verify the baseBranch exists on the remote before pulling; if not, create the feature branch from local HEAD without pulling.
If the working tree has uncommitted changes, stash them first (git stash), create the branch, then pop (git stash pop)."Ticket created and branch checked out. When you're ready to start implementation, say
/toolkit-jira resume <KEY>and I'll show you the full plan before touching any code."
/toolkit-jira resume or user pastes a ticket)Invoked when the user:
/toolkit-jira resumeThis mode works from text already pasted into the chat — no live Jira fetch, no comment/transition side effects. Mode 2 does not update the Jira ticket or transition its status — use /ai-skills:analyze if you want those side effects. If you want to fetch a ticket live by key, re-read it with fresh eyes, and post a plan back to Jira, use /toolkit-analyze <TICKET-KEY> instead.
Steps:
referenceSource to fill in anything the ticket may not have spelled out."Ready to proceed? Reply 'yes' (or 'go') to start, or tell me what to adjust in the plan first."
/toolkit-code (or buildSkillName from config if this repo doesn't have /toolkit-code wired up, or whichever skill the user identifies).Output the ticket using exactly this format so it can be pasted directly into JIRA:
──────────────────────────────────────────
SUMMARY
[<repoSlug>] <Unit of Work> — Implementation
DESCRIPTION
## Overview
<1–2 sentences describing what this does and why.>
## Parity (omit if no referenceSource configured at all)
<Required — vs. <reference path/spec> / Not applicable — net-new, no counterpart>
## Reference Material (omit if parity is not applicable, or no referenceSource configured)
- <reference path/spec>: <what it defines>
## Backend — Phase 1 (omit if no backend layer involved)
Procedures to audit or implement under <appRoots.backend>:
- [ ] <endpoint/function> — <what it fetches/does>
- [ ] <endpoint/function> — <what it fetches/does>
## Frontend — Phase 2 (omit if no frontend layer involved)
File: <appRoots.frontend>/<path>
- [ ] <build step>
- [ ] Wire data to match the reference exactly (omit this line entirely if parity is not applicable)
- [ ] Responsive: <breakpoints from this repo's remembered responsiveness settings, if set>
- [ ] Mirror reference element placement (omit this line entirely if parity is not applicable)
## Acceptance Criteria
- [ ] <criteria specific to this unit of work>
- [ ] Lint passes with zero warnings
- [ ] Typecheck passes with zero errors
- [ ] (optional) Regression tests added after a QA/audit pass, if this repo uses that workflow
## Technical Notes
- Path: <actual or expected path>
- Skill to invoke: `/toolkit-code` (or `buildSkillName` from config, or ask the user)
- <any repo-specific constraints worth calling out — read from this repo's own CLAUDE.md/README if present>
──────────────────────────────────────────
Use this when the unit of work already exists. Replace the greenfield template entirely.
──────────────────────────────────────────
SUMMARY
[<repoSlug>] <Unit of Work> — Audit & Fix
DESCRIPTION
## Overview
<1–2 sentences: what it does AND note that it already exists at <path> (<N> lines).
State that this ticket covers verification and targeted fixes, not a greenfield build.>
## Parity (omit if no referenceSource configured at all)
<Required — vs. <reference path/spec> / Not applicable — net-new, no counterpart>
## Reference Material (omit if parity is not applicable, or no referenceSource configured)
- <reference path/spec>: <what it defines>
## What's Already Implemented
- <bullet: what's wired up>
- <bullet: what's present>
- <bullet: any other completed work>
## Backend — Phase 1 (Audit Only) (omit if no backend layer involved)
- [ ] <endpoint/function> — confirm: <specific thing to verify>
(only list things that need verification; mark "audit only" if no changes expected)
## Frontend — Phase 2 (Targeted Fixes) (omit if no frontend layer involved)
File: <actual path>
- [ ] <specific fix found during audit>
(only list actual gaps found; if a category is clean, omit it)
## Acceptance Criteria
- [ ] All confirmed gaps from the audit above are resolved
- [ ] Lint passes with zero warnings in this file
- [ ] Typecheck passes with zero errors
- [ ] No regressions in existing behavior
## Technical Notes
- Path: <actual path>
- <any status codes, conventions, or navigation-registration notes>
- <note any intentional divergences from the reference that should NOT be changed>
──────────────────────────────────────────
Before any code is written, output this plan for user approval:
Implementation Plan — <Unit of Work>
════════════════════════════════════
Source ticket: <ticket summary or "pasted ticket">
PHASE 1 — Backend Audit (omit if no backend layer involved)
Reference: <whatever referenceSource points at, if configured>
Endpoint/Function Status Action needed
─────────────────────────────────────────────────────────────
<name> ✅ exists Verify response parity
<name> ❌ missing Implement from scratch
<name> ⚠️ shape mismatch Fix field names
(one row per required unit of work)
PHASE 2 — Frontend Build (omit if no frontend layer involved)
File: <appRoots.frontend>/<path>
[ ] Build/wire <N> data sources
[ ] <N> filter controls (if applicable)
[ ] <N> fields/columns (if applicable)
[ ] Responsive layout at configured breakpoints
[ ] Mirror reference: <describe specific placement notes, if a reference is configured>
Scope estimate: Simple / Medium / Complex
Estimated units of work to build/fix: <N>
Additional requirements from ticket:
• <any extra AC or comments the user added>
Ready to proceed? Reply 'yes' to start, or tell me what to adjust.
/toolkit-plan)./toolkit-jira resume./toolkit-analyze <TICKET-KEY>./toolkit-mobileDrive an iOS simulator app through Maestro — run or write flows that assert every step, capture screenshot evidence for a ticket, read the current screen, and prove a probe channel before trusting app logs. Replaces hand-scripted idb/simctl tapping. Use when the user says "/toolkit-mobile", "run the flow", "verify on the simulator", "drive the app", "reproduce this on device", "screenshot this for the ticket", or "get evidence for <ticket>".
Verify app behaviour on an iOS simulator with declarative Maestro flows, not per-call tap scripts. Every step asserts the screen it should have produced, so a run either proves the claim or fails at the step that broke — it never ends on a plausible screen nothing actually reached.
Everything runs through one script: scripts/mobile.sh (in this skill's directory; call it by
absolute path from the app repo's root). Maestro is the driver; the script adds config, evidence,
the probe channel, and loud failure.
No config file to set up — see CONFIG.md for the read/write procedure:
mobile — <SLUG>_MOBILE_APP_ID asked once, then written (required). mobile.sh preflight
lists the non-Apple apps installed on the booted simulator; ask which is the build under test.
_DEVICE only when more than one simulator is booted; _WORKSPACE / _FLOWS_DIR only when the
user moves them.Everything the skill writes — flows, evidence, probe snapshots — lives in
~/.toolkit/mobile/<repo>/ (flows/, evidence/): per person, and outside the repo under
test. In a shared repo these files are noise in every status and one git add . from a
teammate's history, and an ignore rule does not fix that — it is one -f or one moved
.gitignore away. So the script refuses any configured location that resolves inside the repo
(absolute, relative or through a symlink), before creating anything.
mobile.sh preflight first, every session. It checks Maestro, Java 17+, exactly one
booted simulator, the app installed, and Metro on :8081 — and exits non-zero naming what is
missing. The user starts the simulator, Metro and the app unless they ask you to.MAESTRO_CLI_NO_ANALYTICS=1.
Ask before installing. Maestro reports analytics unless that variable is set; the script sets
it on every call, and mobile.sh mcp sets it in each MCP registration.maestro mcp in every agent CLI it finds
(Claude Code at user scope — including the CLI bundled inside an IDE extension — and Codex) when
it is not already there, and prints REGISTERED now when it did. A stdio server does not
hot-reload: after a fresh registration, tell the user to restart the session for the tools.mobile.sh runEvery mobile.sh call starts Maestro cold — JVM plus the on-device driver. Measured: one assert
~60s (a labels took ~130s, but with the MCP server attached to the same device). The maestro MCP tools (registered by preflight) keep that driver warm
and return in seconds. So:
inspect_screen, then run with inline yaml for the
next step or two. No flow file, no ticket, no evidence.mobile.sh run --ticket, which is what
records evidence and a trustworthy exit status.Once the MCP has driven the simulator, a mobile.sh run displaces its driver — no call needs
to be in flight. From then on its tools fail with Device became unreachable during setPermissions
while list_devices still shows the device; retrying does not recover it, and neither does killing
its simulator-server child — the failure is cached in the server. Stopping the server does: the
agent CLI restarts it on the next MCP call and the new one drives the device. So run stops any MCP
server holding the simulator before it starts (and says so); the MCP works again on its next call.
Claude Code restarts it itself; another client may need its MCP reconnected. The first MCP call
after a run can fail with iOS driver not ready in time while the new driver cold-starts — retry
it once. preflight names the
holding server's pid. Still explore first and record last — each switch costs the MCP a cold start.
mobile.sh run <area>/<flow>.yaml --ticket <KEY> # one flow, from the workspace
mobile.sh run --ticket <KEY> # every flow its config.yaml lists
mobile.sh run <flow> --quiet # only assertions, failures, the summary
The exit status is Maestro's own, read from a file rather than through a display pipe: 0 passed,
3 a step failed, 1 the flow never ran, 4 invalid — a live probe saw the bundle evaluated more
times than the flow launches, i.e. Metro reloaded the app mid-run; re-run, and never read a 4 as
evidence about the code. Before starting, run waits until the newest uncommitted file in the repo
is MOBILE_SETTLE seconds old (default 20): Metro's watcher picks an edit up late, and a flow
launched right after one gets the old bundle, then a hot reload that resets the screen ~20s in.
plant always pays this wait — it writes a file and then runs. On a failure run prints the failing step's screen as
labels lines, from the hierarchy Maestro saved, so no relaunch is needed to see it. After every
run the app is brought back to the foreground, warm, so the next read sees the app, not the home
screen.
The harness may reset the cwd between calls. Outside a configured app repo, mobile.sh uses the
last one it ran in and says so. --repo <path> (before the command) picks one explicitly.
takeScreenshot steps land under <evidence>/<KEY>/<timestamp>/…/takeScreenshot/, and
onFlowComplete shoots 99-final even when the flow fails. run prints every PNG path.mobile.sh shot <name> --ticket <KEY> captures the screen as it is right now.mobile.sh labels prints one line per labelled node — label | hint | text | bounds | selector —
which is what selectors match against. Read it (or the MCP's inspect_screen) before writing a
selector or a coordinate. A tap sent at a guessed coordinate that happens to land still
contaminates the trial. The selector column is the label as a ready text regex: parentheses and
other specials escaped, a leading icon glyph or ", " replaced by .*. Copy it rather than
escaping by hand; escaping is the selector mistake that recurs.
Start from templates/flow-template.yaml and save into the workspace's flows/<area>/; run
accepts that bare <area>/<name>.yaml. Flows never go in the app repo, and never in this toolkit —
they name the app's screens and data. Keep the template's first-line yaml-language-server
modeline: outside a .maestro/ folder, the editor otherwise picks a schema by file name. The
conventions, each learned from a run that lied:
launchApp: stopApp: true). A warm repeat can pass a bug that only the first
visit after a restart shows.COMPLETED. Assert the
resulting screen, or assert enabled: true before tapping.assertNotVisible for the
screen the bug would land on. A value both screens show proves nothing."City" also matches a
CITY section label; a label with a leading ", " needs ".*Name.*". Anchor with below: /
above: relative to a label, or use a testID (id:).hideKeyboard on a timer — wrap tap → type → assert
where the text landed in retry:.point: "x%,y%" — which is
device-dependent — then assert the result, and report it as an app accessibility defect (the fix
is accessible={false} on the container).COMPLETED and nothing opens. Wrap tap → wait-for-next-screen in retry:.checked: does not read RN's accessibilityState. The state is in
the value text, so match ".*checkbox, checked.*".scrollUntilVisible trusts a stale tree. On a long screen iOS can report an off-screen
element as visible, so the scroll returns COMPLETED without moving and the next tapOn lands on
its stale coordinates inside another control — also COMPLETED, so it reads as an app bug. run
refuses a scrollUntilVisible without centerElement: true; on a long screen put a bare
- scroll before it too, and assert the tap's effect, not just the element's presence.point: tap measured
before it lands on something else. Re-read the screen after anything that can add a line.Reaching a state (a filled cart, a completed checkout step) through one flow per attempt costs a
full run each time. Put each reusable path in the workspace as flows/_shared/<name>.yaml with its
inputs as ${VARS}, and call it:
- runFlow:
file: ../_shared/add-to-cart.yaml
env: { ITEM: "…", DATE: "…" }
A _shared/cold-start.yaml (launch, first wait, dismiss the dev warning toast behind a
runFlow: when: visible: condition) saves every flow from repeating those steps. These live in the
workspace, not this toolkit, because they name the app's screens and data.
A flow that has only ever passed is untested. For a bug fix, run the flow against the pre-fix version of the file and confirm it fails at the assertion that names the bug:
mobile.sh plant <file> [--from <ref>] -- <flow> --ticket <KEY>-planted
plant snapshots the file, writes <ref>'s version (default HEAD), runs, and restores with an
md5 check on every exit path, including an error or ^C. It fails if the flow passed, or never ran.
It cannot tell which step failed for the right reason. Read the failing step it prints.
Console output goes to the dev overlay, which the accessibility tree cannot see, and the simulator log does not carry it. When a fix needs runtime values:
mobile.sh probe start # host listener; prints the snippet to paste
mobile.sh probe snapshot <file...> # BEFORE editing
# paste the snippet, add __probe('label', value) calls, reload the app
mobile.sh probe check # fails until the control line has arrived
mobile.sh probe tail
mobile.sh probe restore # restores, md5-verifies, fails on any __probe residue
mobile.sh probe stop
One snapshot set is open at a time. A second snapshot (or a plant) is refused until restore,
so list every file in the one call.
Never trust silence from an unproven channel — check must pass first. start refuses a port
held by an orphaned listener from an earlier session and names the process; stop it rather than
picking another port, or its log quietly swallows the run.
probe restore before any commit.run's last line./toolkit-planConversational requirements-gathering tool. Talks through a unit of work with the user — scope, reference/parity, acceptance criteria — before any ticket or code exists. Ends in an approved spec file once the user confirms everything is locked in. Use this when starting something new and nothing has been scoped yet. Not for tickets that already exist — use /toolkit-jira resume for those.
You are the requirements-gathering agent. Your job is to talk through a unit of work with the user until the requirements are fully agreed, then write that agreement down as a spec file. No Jira ticket exists yet when this skill runs — /toolkit-jira reads the spec you produce here and creates the actual ticket from it afterward.
You never write code, and you never touch Jira. You produce one artifact: an approved spec file.
No config file to set up — values are detected from the repo or asked once and remembered in ~/.dev-skills.env (see CONFIG.md for the full read/write procedure and category definitions):
appRoots — detected fresh every run — used to phrase scope/layer questions in terms of this repo's actual layer names (e.g. backend/frontend, or just app for a single-layer repo).referenceSource — asked once, then written — if this repo cross-references another source (a sibling repo, a native counterpart, a Figma spec, a design doc) for parity, this drives the parity question in Step 2.specCachePath — detected fresh every run — checks for .claude/specs/ first, then falls back to other spec-cache variants (screen-specs/, figma-specs/) if .claude/specs/ doesn't exist; defaults to .claude/specs if none found yet./toolkit-plan <rough description of what you want to build>
/toolkit-plan
Examples:
/toolkit-plan add a CSV export to the invoice list/toolkit-plan (then ask what they want to build)If the user's invocation already states what they want, confirm your understanding back in one sentence. If they just typed /toolkit-plan with nothing else, ask: "What do you want to build or fix?"
Before asking anything else, check specCachePath (or whatever spec-cache convention is already in use under .claude/ — some repos use screen-specs/, figma-specs/, etc.) for a spec matching this unit of work. If one already exists:
This is the core of the skill. Ask questions one at a time, not as a giant upfront checklist — let the user's answers shape what you ask next. Cover, in roughly this order, skipping anything the user already answered unprompted:
appRoots layer(s) does this touch (backend, frontend, both)? What's the smallest version of this that's still useful?referenceSource is configured for this repo (ask once if not already written to ~/.dev-skills.env), and don't assume it applies just because the config exists repo-wide: "Does this need to match <referenceSource description>, or is this net-new with no existing counterpart?" Record the answer plainly — this becomes the spec's Parity check: line, the same flag /toolkit-qa later reads to decide whether to run its Reference Parity Audit. A repo-wide reference source doesn't mean every new screen has a counterpart there.Keep looping until the user signals they're done, or until you've covered the above with nothing left ambiguous. Don't pad the conversation with questions that don't change the spec — if scope and acceptance criteria are already unambiguous from the user's first message, move straight to Step 4.
Once the conversation has converged, output a structured spec:
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Spec — <Unit of Work>
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Type: <backend | frontend | fullstack | bug | chore>
── OVERVIEW ──────────────────────────────
<2–3 sentences explaining what this does and why, in plain English.>
── PARITY ───────────────────────────────── (omit if no referenceSource configured at all)
<Required — vs. <referenceSource description> / Not applicable — net-new, no counterpart>
── REFERENCE MATERIAL ──────────────────── (omit if parity is not applicable, or no referenceSource configured)
<Reference path/spec>: <specific file(s) or section to cross-reference>
── CHANGES BY LAYER ──────────────────────
<One subsection per appRoots key touched>
File: <path under that layer's root, or "TBD" if not yet known>
Unit of work Notes
─────────────────────────────────────────────────────────
<function/endpoint/component> <what it does>
── ACCEPTANCE CRITERIA ───────────────────
[ ] <criterion confirmed in conversation>
[ ] <criterion confirmed in conversation>
── RISKS / UNKNOWNS ──────────────────────
• <anything flagged during the conversation that still needs investigation>
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Does this fully capture what you want? Reply 'yes' to lock this in,
or tell me what to adjust.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Do not invent acceptance criteria or scope the user hasn't actually confirmed — every line in this spec should trace back to something said in the conversation. If the user hasn't confirmed a category (e.g. acceptance criteria), leave it empty or mark it TBD rather than inferring.
Do not write the spec file until the user explicitly approves. If they ask for changes, revise and present again — same discipline as every other approval gate in this repo's skills.
Once approved, write the spec to specCachePath (create the directory if it doesn't exist). Use a filename that matches this repo's existing spec-naming convention if one exists (check what's already in that directory); otherwise, slugify the unit-of-work name.
Mark the file as approved (e.g. a top-line Status: approved or equivalent your spec format already uses) so a future /toolkit-plan invocation recognizes it as locked-in rather than a draft.
"Spec saved at
<path>. Run/toolkit-jirato create the ticket from this, or/toolkit-codedirectly if you don't need a ticket."
referenceSource rather than assuming what it contains./toolkit-jira resume instead, since that ticket already has its requirements written down.Tell me what you want to build or fix, and we'll work through the requirements together.
/toolkit-pr-reviewReview an already-open GitHub pull request (yours or a teammate's) by number or URL — runs the same correctness/security/architecture checks as /toolkit-code-review against the PR's diff, then posts the findings as an actual GitHub PR review. Use when the user says "/toolkit-pr-review", "review PR
Reviews an already-open GitHub PR — fetched via gh pr diff, not the local working tree. Runs the same diff-content checks as /toolkit-code-review (steps 2–9 there), then posts the result as a real gh pr review (comment / approve / request-changes) so the finding is visible to the PR author and any other collaborators.
This is a different skill from /toolkit-code-review, not a mode of it: /toolkit-code-review reviews your own uncommitted working-tree changes and, on request, commits them and advances a ticket — none of that applies to a PR that's already open under someone else's (or your own) commits. pr-review never commits, never pushes, and never touches the local working tree; its only write is the posted GitHub review.
No config file to set up — values are detected from the repo or asked once and remembered in ~/.dev-skills.env (see CONFIG.md for the full read/write procedure and category definitions):
githubRepo — detected fresh every run — org/repo parsed from git remote get-url origin, used to resolve a bare PR number to a full gh pr target.appRoots — detected fresh every run — used to bucket changed files by layer for the report.qualityGate — detected fresh every run — lint/typecheck commands per layer. Only run if step 1's checkout is available (see Step 2 below) — a PR's lint/typecheck can't be run against a diff alone, only against the PR's actual branch checked out locally.repoRules — detected fresh every run — checks conventional paths for a repo-rules file. Skips the repo-architecture-rules step and notes it as not found if none exists.localeDictionaryPath — detected once, then written — the file holding this repo's actual locale/translation data, if one exists (see CONFIG.md). Enables the diff-scoped locale dictionary hygiene check. Skips that step cleanly if not configured.Required — exactly one of:
/toolkit-pr-review 254 or /toolkit-pr-review #254) — resolved against githubRepo./toolkit-pr-review https://github.com/org/repo/pull/254) — repo and number both parsed from the URL directly; use the URL's own org/repo, not the locally detected githubRepo, if they differ (reviewing a PR in a different repo than the one you're currently in is valid and shouldn't silently redirect to the wrong repo).If neither is given, ask: "Which PR? Give me a number or a URL."
gh pr view <number> --json title,author,baseRefName,headRefName,url,additions,deletions,files
gh pr diff <number>
If gh pr view errors (PR doesn't exist, no access, wrong repo), stop and report the exact error — don't guess at a different PR number.
Identify which files changed from the files field and bucket them by appRoots layer, plus generic buckets for docs/**/*.md (skip most checks) and tooling/config files (skip most checks) — same bucketing rule /toolkit-code-review step 1 uses.
Unlike /toolkit-code-review, there is no local working tree to run lint/typecheck against — gh pr diff gives you diff text, not a checked-out branch. Two options, in order of preference:
gh pr checkout <number> into a fresh worktree, not the user's active branch), do so and run this layer's qualityGate commands there. Always confirm before doing this — checking out a PR branch is a local filesystem action the user should approve, even in a worktree.⏭ SKIPPED — no local checkout; lint/typecheck not run against diff text alone in the report. This is not a failure — say so plainly so the report doesn't read as if quality gates silently passed.Default to skipping (the second option) unless the user asks for the checkout — most PR reviews are meant to be fast, read-only glances, and a full clone/checkout changes that.
/toolkit-code-review's diff-content checks verbatimRun steps 3 (security scan, all sub-checks 3a–3g), 4 (debug/dead-code leftovers), 5 (type-safety quality scan), 5.5 (locale dictionary hygiene, if configured), 6 (repo architecture rules, if configured), 7 (cross-platform safety, if multi-surface), and 8 (backend/network pattern scan) exactly as documented in code-review/SKILL.md, substituting the gh pr diff output (and gh pr view --json files file list) everywhere that skill's procedure says <changed-files> or refers to the working-tree diff. Do not duplicate that logic here — read ../code-review/SKILL.md steps 3 through 8 at review time and apply them against this PR's diff instead of git diff.
One difference from /toolkit-code-review's step 8: "backend files touched" and "frontend files touched" are read from the PR's files list (step 1), not from a local git status.
Same table shape as /toolkit-code-review step 9, with a PR-identifying header instead of a timestamp-only one:
# PR Review — #<number> "<title>" by <author>
<base branch> ← <head branch> | <N> files changed (<breakdown by appRoots layer + docs/tooling>) | <url>
## Check Results
| # | Check | Result | Details |
|---|---|---|---|
| 2 | Lint / Typecheck | ✅ PASS / ❌ FAIL / ⏭ SKIPPED — no local checkout | … |
| 3a | Secrets/credentials | ✅ PASS / ❌ FAIL | … |
| 3b | Injection (A03) | ✅ PASS / ❌ FAIL / ⚠ WARN | … |
| 3c | XSS / unsafe rendering (A03) | ✅ PASS / ⚠ WARN / ⏭ n/a | … |
| 3d | Broken access control (A01) | ✅ PASS / ❌ FAIL / ⚠ WARN / ⏭ n/a | … |
| 3e | Sensitive data exposure (A02) | ✅ PASS / ⚠ WARN | … |
| 3f | Mass assignment (A04/A08) | ✅ PASS / ❌ FAIL / ⏭ n/a | … |
| 3g | Security misconfiguration (A05) | ✅ PASS / ❌ FAIL | … |
| 4 | Debug leftovers | ✅ PASS / ⚠ WARN | … |
| 5 | Type-safety quality | ✅ PASS / ⚠ WARN | … |
| 5.5 | Locale dictionary hygiene | ✅ PASS / ❌ FAIL / ⚠ WARN / ⏭ n/a | … |
| 6 | Repo architecture rules | ✅ PASS / ❌ FAIL / ⏭ not configured | … |
| 7 | Cross-platform safety | ✅ PASS / ⚠ WARN / ⏭ n/a | … |
| 8 | Backend/network patterns | ✅ PASS / ⚠ WARN / ⏭ n/a | … |
## Failures (blocking)
<file>:<line>: <description>
## Warnings (non-blocking)
<file>:<line>: <description>
## Logic / architecture notes
<Free-text observations. Keep to ≤5 bullet points. Be specific.>
## Verdict
✅ APPROVE — N warnings, 0 blocking failures.
OR
❌ REQUEST CHANGES — N failures, N warnings.
OR
💬 COMMENT — no blocking failures, but findings worth the author seeing before merge (used when warnings are substantive enough to flag but not clearly blocking).
Show the rendered report to the user first and ask for confirmation before posting — posting a PR review is visible to the author and any other collaborators, not a local/reversible action the way a chat report is.
"Post this as a PR review on #? (approve / request changes / comment, matching the verdict above)"
On confirmation, post via gh pr review:
# Verdict: APPROVE
gh pr review <number> --approve --body "<report body>"
# Verdict: REQUEST CHANGES
gh pr review <number> --request-changes --body "<report body>"
# Verdict: COMMENT
gh pr review <number> --comment --body "<report body>"
Use the full rendered report (step 10) as the body, passed via a heredoc the same way create-pr/jira pass multi-line bodies — never inline-escape a multi-line string. Report the resulting review URL back to the user.
If the user declines to post, stop after showing the report — nothing is sent to GitHub. A bare /toolkit-pr-review with no follow-up confirmation never posts anything.
/toolkit-code-review step 11's job for your own pre-commit flow, and it doesn't map cleanly onto "someone else's PR was reviewed" — a PR review isn't the same event as a ticket-owner's own commit./toolkit-pr-review <number>, /toolkit-pr-review <URL>, review PR #<number>, review this pull request, review this PR
/toolkit-qaAudits a built screen/feature for two things — production-grade structural quality (lint, typecheck, locale/theme hygiene, repo-specific rule violations, loading/empty/error state coverage and swallowed errors, a WCAG 2.2 AA accessibility spot-check, comment hygiene — no diary, redundant, dead, or stale comments — and test presence) and, if a reference is configured, parity against that reference (another repo's matching implementation, a native counterpart, a Figma spec, or a design doc). Read-only — never edits source. Reports PASS/FAIL/WARN with file:line. Use when the user says "/toolkit-qa <name>", "audit this", "check before commit", or "what's missing vs the reference".
You are the QA agent. Your job is to audit a named unit of work and produce a structured report. You do not write or fix code — you identify problems and report them clearly so the developer (or a build skill) can act on them.
This skill runs two largely independent passes — Structural Audit (always runs) and Reference Parity Audit (runs only if a reference is configured). A repo with no referenceSource configured still gets full value from the Structural Audit alone.
No config file to set up — values are detected from the repo or asked once and remembered in ~/.dev-skills.env (see CONFIG.md for the full read/write procedure and category definitions):
appRoots — detected fresh every run — where to scope file searches by layer.qualityGate — detected fresh every run — the lint/typecheck/test commands to run per layer, read from each layer's manifest scripts.referenceSource — asked once, then written — drives the Reference Parity Audit:
type: "repo" — another local checkout (e.g. a legacy web app this repo is porting from). path points at it.type: "native-counterpart" — this repo has both a native and a web/shadow implementation of the same screen at predictable sibling paths; compare functional flows between them (no external repo involved).type: "figma-spec" — a cached spec file (e.g. .claude/figma-specs/<slug>.md) is the source of truth for what should exist.type: "doc" — a design doc path/URL.localeConvention — detected once, then written — the hook/function name this repo uses for i18n (e.g. useLocale, t), so the locale-string scan knows what "already wrapped" looks like. See the Locale-string audit section (B) below for how detection works.localeDictionaryPath — detected once, then written — the file holding the actual locale data (e.g. a localeDictionary.ts keyed by locale code, each locale containing named sections of string keys). Enables the two dictionary-level checks below (B2, B3). See CONFIG.md for detection procedure. If not detected, B2/B3 report ⏭ not configured rather than failing.themeConvention — detected once, then written — the hook/object name this repo uses for theme colors (e.g. useWebScreenTheme, theme), so the hardcoded-color scan knows what to treat as already-themed. See the Theme/color audit section (C) below.repoRules — detected fresh every run — checks conventional paths for a repo-rules file listing banned/required patterns. Skips repo-rule checks if none found.commentPolicy — detected once, then written — lean (default) or verbose, deciding which comment classes check F runs. Detection and the policy itself live in COMMENTS.md.If nothing's detectable yet (a brand-new repo with no existing screens to grep), run the Structural Audit using sensible generic defaults (standard lint/typecheck scripts, common hex-color and string-literal heuristics) and tell the user what's still undetected.
appRoots layer. If ambiguous, find/grep for it and confirm with the user before proceeding..claude/ (this repo may cache specs under different names — screen-specs/, figma-specs/, etc.; check what convention is actually in use) before re-deriving requirements from a reference each time.Parity check: not applicable line (written by /toolkit-jira when the unit of work was scoped as net-new), skip the entire Reference Parity Audit section below — go straight to the Structural Audit and omit the Reference Parity Audit section from the report entirely, rather than reporting a missing reference as a finding. If the spec says Parity check: required or has no parity line at all (e.g. no spec exists, or it predates this convention), proceed as normal.referenceSource is configured but no spec file exists to confirm the parity decision either way, ask the user once before assuming parity applies: "This repo has a reference source configured, but I don't have a recorded parity decision for <unit of work> — does this need to match <referenceSource description>, or is it net-new?" Don't silently assume either answer.referenceSource.type is "repo" or "doc", locate the matching reference file(s) using grep first, then targeted Read offsets — never read a large reference file from line 1.referenceSource.type is "native-counterpart", locate the native and web/shadow sibling files for this unit of work.Run this layer's qualityGate commands from config (lint, typecheck, and any test command). Capture exit codes and the first ~30 lines of any errors.
Distinguish a missing script from a real failure. If a command errors with something like "missing script," "command not found," or "no such file" — that's a detection issue, not a quality-gate finding. Report it separately: "Detected <command> for this layer's quality gate, but it doesn't exist in this repo — check this layer's package.json/manifest scripts." Don't report it as ❌ FAIL alongside genuine lint/typecheck errors; a misdetected command tells you nothing about code quality.
FAIL if any command runs and exits non-zero for a real reason (actual lint/type errors). Don't paper over a failure by suggesting a suppression (eslint-disable, @ts-expect-error) — that's not a fix, that's hiding the finding.
If <SLUG>_LOCALE_CONVENTION is not already in ~/.dev-skills.env, detect it once, then write it so this detection cost is paid once per repo, not once per /toolkit-qa run: grep one or two existing screens in the target appRoots layer for a common i18n call shape (useLocale(, useTranslation(, a bare t(, etc.). If found, use it for this run and write it to ~/.dev-skills.env as <SLUG>_LOCALE_CONVENTION so every future invocation (by you or anyone else on the team) reads it straight from there instead of re-detecting. If genuinely nothing i18n-shaped exists anywhere yet, skip this check and note it as not detected — don't guess a convention with zero evidence.
For each touched/target file, scan for JSX text nodes and string literal props that look like user-visible copy but aren't wrapped in the configured locale function:
grep -nE '>([A-Z][a-zA-Z ]{2,})<' <file>
Manually verify each match — some may be intentional (proper nouns, brand names, status codes that are domain abbreviations rather than translatable prose).
This pattern only catches same-line JSX text nodes — know its blind spots before trusting a clean result. It does not catch: strings passed as component props (title=, placeholder=, label=), multi-line <Text>\n copy\n</Text> blocks where the text isn't on the same line as the tags, or strings inside ternaries/template literals ({isLoading ? 'Loading...' : 'Submit'}). Supplement with a targeted literal-string search (grep -n "<exact phrase>") for any string a user or screenshot flags as still-English despite this sweep reporting clean — a "0 remaining" result from this grep alone is not proof of full coverage.
Also scan for derived-but-untranslated copy — not every hardcoded-English defect is a string literal. Rendering user-visible text by transforming an internal state/variable name directly (capitalizing a raw enum value, slicing/titlecasing a route param, .replace()-ing an underscore out of a constant) has the same effect as a hardcoded string: it can never change with the selected language. Grep for the shape, not just the literal:
grep -nE '\.charAt\(0\)\.toUpperCase\(\)|\.toUpperCase\(\)\s*\+|titleCase\(|capitalize\(' <file>
For each match, check whether the transformed value feeds directly into rendered JSX/Text output (a defect — flag it) versus an internal-only use (log keys, analytics event names, CSS class strings — not user-facing, skip it).
FAIL if any genuinely user-visible unwrapped string OR derived-but-untranslated expression remains. Report each as <file>:<line>: <snippet>.
localeDictionaryPath is configured)A screen can call the locale function correctly and still be untranslated — the dictionary can have the right keys in every locale block while several non-English blocks hold copy-pasted English values instead of real translations. This is invisible from the screen file (it correctly reads locale.section.key) and invisible to typecheck (the key exists everywhere), so it can only be caught by reading the dictionary file's actual values.
Resolve which section(s) the target screen reads from by grepping the screen file for the locale-access pattern (e.g. locale\.(\w+)\. — collect the distinct section names in use).
Parse the dictionary file using brace-depth tracking, not line-regex — naive regex matching on section/key names false-positives when a key name also happens to appear as a string value elsewhere in the file. The technique: walk the file tracking {/} depth; a top-level block opens at depth 1 keyed by locale code (enUS:, esMX:, etc.); within each locale block, a section opens at depth 2 keyed by section name; collect each section's key→value pairs between its own open/close braces. Do this for every locale block, not just the one(s) you expect to check.
For the section(s) in scope, compare every non-English locale's value against the English (first/reference locale's) value for the same key.
FAIL any key where a non-English locale's value is byte-identical to the English value — this is the copy-pasted-placeholder defect. Use judgment on short strings: English coincidentally matching another language for one word (e.g. "OK", "Email") is possible and not necessarily a defect — flag anything longer than ~3 words as a near-certain finding, and treat shorter matches as a note rather than a hard fail unless the phrase is unambiguously not a valid word/phrase in that language.
Report each finding as <dictionary file>: <locale>.<section>.<key> — value identical to English ("<value>").
localeDictionaryPath is configured and a companion interface/type file can be located)Locate the type this repo's locale hook returns (e.g. ReturnType<typeof useLocale>, or an explicitly-named interface the hook imports) — this is the interface/type file to audit against the dictionary.
Extract every key declared in the interface/type file, per section, using the same brace-depth approach as B2 (applied to the type file's nested-object-type syntax instead of value literals). Extract every key with an actual value in the dictionary file's reference (English) locale block, same section.
FAIL any interface key with no matching value in the reference locale — this is the defect that breaks typecheck the moment any locale (or the type system generally) treats that key as required, and it's usually the result of an interrupted or partially-applied edit (an interface key added without its matching dictionary values, or vice versa). Report as <interface file>: <section>.<key> declared but has no value in <dictionary file>'s <referenceLocale> block.
Note explicitly in the report if the interface marks the key optional (e.g. a TS key?: type) — an optional key with no value is not a defect, just currently unused; don't fail it, mention it as informational only.
If <SLUG>_THEME_CONVENTION is not already in ~/.dev-skills.env, detect it once, then write it — same pattern as locale detection above: grep an existing screen for a theme hook/object call shape (useTheme(, useWebScreenTheme(, a theme. property-access pattern), use it for this run, and write it to ~/.dev-skills.env as <SLUG>_THEME_CONVENTION so it's a one-time cost. Skip the check cleanly if nothing theme-shaped exists yet.
Scan for hardcoded color literals outside the configured theme convention:
grep -nE '#[0-9a-fA-F]{3,8}' <file> | grep -v '<themeConvention>\.'
Substitute the actual detected themeConvention value before running: e.g. if the convention is theme.colors, the grep becomes grep -v 'theme\.colors\.'.
FAIL if any raw hex appears in a style that should be themed. Exemptions: inline SVG fill=, semantic/intentional brand or destructive-action colors the repo has explicitly decided not to theme (check repoRules for a documented exemption list before flagging — to check repoRules: look for a repoRules or .claude/rules.md file and grep it for 'color exemption' or the specific color value before flagging), test files.
repoRules is configured)Read the repoRules file and run whatever banned-pattern/required-pattern greps it lists (e.g. "ban package X", "require primitive Y instead of raw Z", "no raw fetch() outside the service layer"). Report each violation as <file>:<line>: <rule violated>.
If this unit of work has a counterpart on another platform/surface (native vs. web, mobile vs. desktop) that should NOT have been touched by this change, diff against the base branch to confirm:
git diff --name-only <baseBranch> -- <counterpart path(s)>
<baseBranch> is the detected-once value from ~/.dev-skills.env (key: <SLUG>_BASE_BRANCH) — detect it from git remote show origin if not yet set.
WARN if the counterpart or shared utilities changed unexpectedly.
Comments in shipped code are not a diary. Apply COMMENTS.md to the target files: resolve commentPolicy, run its candidate greps, read every hit in context, drop anything on its never-flag list (pragmas, fences, license headers, API contract docs, why-comments), and classify the rest as diary, redundant, dead, or stale. Severity follows that file's table.
Every finding carries its trim — the exact replacement text or delete — so the fix needs no re-derivation. This skill still never applies it; /scar-harden or the developer does.
A screen that works on the happy path is a demo, not a production screen. For every async operation in the unit (fetch, query hook, mutation, submit), confirm each state the user can reach renders something deliberate:
# Swallowed errors — the error state can't render if the error never arrives
grep -nE 'catch\s*(\(\w*\))?\s*\{\s*\}|\.catch\(\s*\(\)\s*=>\s*(\{\s*\}|null|undefined)\s*\)|except[^:]*:\s*pass' <file>
# State branches present (compare against the async calls you found)
grep -nE 'isLoading|isPending|isFetching|isError|error\b|EmptyState|length\s*===?\s*0' <file>
FAIL a swallowed error, or an async operation with no error branch. WARN a missing empty or loading state. Cite the async call's file:line and name the missing state.
Greps catch the common, mechanical failures; they do not establish conformance. If the repo already runs eslint-plugin-jsx-a11y, axe, or an RN a11y linter, its results arrived with the quality gate — cite them, don't re-derive them.
# Images with no text alternative (1.1.1)
grep -nE '<img\b[^>]*>' <file> | grep -v 'alt='
# Click handlers on non-interactive elements — no keyboard path (2.1.1, 4.1.2)
grep -nE '<(div|span|li|td)\b[^>]*onClick=' <file>
# Icon-only buttons with no accessible name (4.1.2)
grep -nE '<(button|IconButton|Pressable|TouchableOpacity)\b[^>]*>\s*<(\w*Icon|svg|Image)' <file> | grep -vE 'aria-label|accessibilityLabel'
# Focus indicator removed (2.4.7 / 2.4.11)
grep -nE 'outline:\s*(none|0)' <file>
For each hit, read the context: a role+tabIndex+key handler, an aria-label/accessibilityLabel, a :focus-visible replacement, or a decorative alt="" makes it a pass. Also check, by reading: form fields have an associated label, errors are announced (role="alert"/live region) rather than conveyed by color alone, and explicitly sized touch targets are at least 24×24 CSS px.
FAIL a control with no accessible name or no keyboard path. WARN everything else. State in the report that this is a spot-check and that a manual keyboard pass is still owed.
Find the test file(s) that exercise this unit (by import of the unit's module, or by the repo's co-location convention). The quality gate ran the suite; this asks whether the suite covers this unit at all.
WARN if no test touches it, naming /toolkit-unit-tester (pure functions) or /toolkit-tester (I/O-bound). WARN if the only tests are snapshot tests — they pass on whatever the component renders today.
referenceSource is configured)referenceSource.type is "repo" or "doc" — Field/Structure Parity ModeThis is the right mode when porting one specific implementation's behavior into this repo (e.g. legacy app → new app), and "parity" means matching output exactly.
1 — Reference summary. Read the reference and extract: response/data shape (fields, computed values, status codes), UI elements (columns, filters, tabs, action buttons, empty states), and any client-side transforms applied to the raw data before rendering.
2 — Field/shape parity. Compare this repo's implementation against the reference field-by-field:
| Field (reference) | Field (this repo) | Match? | Notes |
|---|---|---|---|
| (one row per field) | ✅ / ❌ |
Flag ❌ for any missing field, name mismatch, or un-replicated computed value. Flag ⚠️ for an intentional, confirmed-deliberate rename/divergence.
3 — UI element parity. Same table shape for columns/filters/tabs/buttons/empty-states — does this repo render everything the reference does, with the same default values and the same spatial placement order?
4 — Query/data-scope parity (if applicable). If both sides query a data source, compare scope dimensions (collection/table, ownership/tenant filter, status filter, date-range filter, join strategy, sort order) — a scope mismatch is the most common cause of "this repo shows different data than the reference."
referenceSource.type is "native-counterpart" — Functional Flow Parity ModeThis is the right mode when comparing two implementations of the same screen on different surfaces within this repo, and "parity" means every capability the user has on one surface also exists on the other — visual differences are fine, capability differences are not.
1 — Grep the counterpart for flow anchors:
# CRUD / mutation triggers
grep -nE '(create|update|delete|patch|post|mutate)\(' <counterpart>
# User-triggered actions
grep -nE '(onPress|onLongPress|onClick|onSwipe)' <counterpart>
# Confirmation flows (destructive actions)
grep -nE '(confirm|Alert\.|Modal)' <counterpart>
# Validations / guards / pre-conditions
grep -nE '(validate|isValid|required|disabled\s*=)' <counterpart>
# Role / status / permission gates
grep -nE '(role|hasPermission|status\.includes|allowedRoles)' <counterpart>
# Empty / loading / error states
grep -nE '(isLoading|isError|EmptyState|catch\s*\()' <counterpart>
# Platform-only APIs that need an alternative on the other surface
grep -nE '(Camera|Linking\.|Share\.|Vibration|PushNotification)' <counterpart>
Read targeted offsets for context — never the whole file from line 1.
2 — Cross-reference each flow against the target surface. For every flow found, grep the target file for the matching trigger and mark:
3 — Write the flow table (one row per flow, with the counterpart's file:line as evidence — never report a gap without a citation):
| Flow | Trigger | Business rule | Counterpart status |
|---|
referenceSource.type is "figma-spec" — Spec Conformance ModeRead the cached spec file's layout/locale/component sections and confirm the build matches: every section/field/button the spec lists is present, uses the configured locale convention, and uses the configured theme convention. This overlaps with the Structural Audit's locale/theme checks — don't double-report; cite it once.
QA Report — <Unit of Work>
══════════════════════════════════════════════════════
References:
Target file(s): <path(s)>
Reference source: <referenceSource summary, or "none configured — structural audit only">
──────────────────────────────────────────────────────
STRUCTURAL AUDIT
──────────────────────────────────────────────────────
Quality gate: ✅ PASS / ❌ FAIL — <details>
Locale strings: ✅ PASS / ❌ FAIL / ⏭ not configured — <details>
Dictionary value parity: ✅ PASS / ❌ FAIL / ⏭ not configured — <details>
Orphaned interface keys: ✅ PASS / ❌ FAIL / ⏭ not configured — <details>
Theme colors: ✅ PASS / ❌ FAIL / ⏭ not configured — <details>
Repo rules: ✅ PASS / ❌ FAIL / ⏭ not configured — <details>
Counterpart safety: ✅ PASS / ⚠️ WARN / ⏭ n/a — <details>
Comment hygiene: ✅ PASS / ❌ FAIL / ⚠️ WARN — <policy: lean|verbose>; <N> to trim
State coverage / errors: ✅ PASS / ❌ FAIL / ⚠️ WARN — <details>
Accessibility (spot-check):✅ PASS / ❌ FAIL / ⚠️ WARN — <details>; manual keyboard pass owed
Test presence: ✅ PASS / ⚠️ WARN — <test file(s), or none>
──────────────────────────────────────────────────────
REFERENCE PARITY AUDIT (omit entirely if no referenceSource)
──────────────────────────────────────────────────────
Mode: <Field/Structure Parity | Functional Flow Parity | Spec Conformance>
✅ <item> matches reference
❌ <item> MISSING / MISMATCH — <detail, with file:line citation>
⚠️ <item> partial — <what's missing>
➖ <item> N/A — <reason, proposed alternative if any>
──────────────────────────────────────────────────────
SUMMARY
──────────────────────────────────────────────────────
✅ Passed: N checks
❌ Failed: N checks ← must fix before this is QA-approved
⚠️ Warnings: N items ← investigate but may be acceptable
Overall status: ✅ QA PASSED / ❌ QA FAILED — <N> blocking issues
| Symbol | Meaning |
|---|---|
| ✅ | Passes — no action needed |
| ❌ | Blocking — must be fixed before this can be considered QA-approved |
| ⚠️ | Warning — investigate and confirm it's intentional |
| ➖ | N/A — not applicable, with reasoning given |
| ⏭ | Skipped — config key not set, check not run (not a failure, just unconfigured) |
?? in the report rather than inventing an answer.What the production-grade checks (F–I) are grounded in; the comment policy's own sources are in COMMENTS.md.
Tell me which unit of work to audit. I'll run the Structural Audit always, and the Reference Parity Audit if referenceSource is configured.
/toolkit-responsiveSweep a web route at this repo's configured breakpoints, capture screenshots, and produce a PASS/FAIL report of layout problems. Report-only — never edits application code. Use when the user says "/toolkit-responsive <route>", "check responsiveness of <route>", "sweep breakpoints", or "screenshot all widths".
Render a route at each configured breakpoint, capture screenshots, and produce a written report of layout problems. Read-only on application code — never edits the page being audited. The user (or this repo's own screen-build skill) fixes manually based on the report.
No config file to set up — values are detected from the repo or asked once and remembered in ~/.dev-skills.env (see CONFIG.md for the full read/write procedure and category definitions):
responsiveness.breakpoints — asked once, then written — array of pixel widths to sweep. Offer [375, 768, 1100, 1440] as the default to confirm rather than asking for raw numbers cold.responsiveness.viewer — asked once, then written — "devserver" (plain dev server, append the route directly) or "lab" (a dedicated in-app Responsive Lab route that takes a ?route= query param). Offer "devserver" as the default.responsiveness.labRoute — asked once, then written, only if viewer is "lab" — the path to the lab route (e.g. /(dev)/responsive).responsiveness.devserverUrl — asked once, then written, required — base URL of the running dev server (e.g. http://localhost:8081).appRoots — detected fresh every run — used to phrase "which file to check for a layout reference" if the audit checklist calls for it.responsiveness.devserverUrl. If unreachable, tell the user to start it and re-invoke — never start it yourself.viewer is "lab", the lab route must exist in this repo. If absent, stop and tell the user — don't fall back to "devserver" silently, since the lab may apply checks the plain dev server URL can't (e.g. synthetic viewport frames)./inspections, /(tabs)/dashboard, /inspection/123. Required.Accept any reasonable shorthand (inspections, /inspections, (tabs)/dashboard) and normalize to the absolute path this repo's router uses. To normalize: check the router config file (e.g. app/_layout.tsx, src/router.ts, routes/) for the canonical path. Accept inspections, /inspections, and (tabs)/inspections as equivalent if they resolve to the same route.
viewer is "devserver": <devserverUrl><route>viewer is "lab": <devserverUrl><labRoute>?route=<url-encoded route>Check reachability before opening anything — open succeeds even if the dev server is down (the browser just shows a connection-refused page), so this check is the only thing that actually catches the "server isn't running" case described in Prerequisites:
curl -sf -o /dev/null --max-time 3 "<devserverUrl>"
If this fails (non-zero exit), stop here — don't proceed to open. Tell the user: "Dev server unreachable at <devserverUrl> — start it and re-invoke." Never start it yourself. If the server returns 401/403 (auth required), treat the route as reachable — the dev server is running. Only treat 404, connection refused, or timeout as unreachable.
If reachable, open in the user's default browser:
open "<target URL>"
If open itself fails or the user is headless, output the URL and stop with a one-line note instead of erroring.
This skill is report-only — it does not control the browser. Produce a checklist of what the user should screenshot and where to save them so the report can reference them:
.claude/responsive-runs/<slug>-<timestamp>/<breakpoint-label>.png
One file per configured breakpoint. <slug> is the kebab-cased route, <timestamp> is YYYYMMDD-HHMM. Use the pixel width as the label (e.g. 375.png, 768.png, 1280.png). mkdir -p that directory before reporting the paths.
For each width, walk the screenshot and flag:
| Category | Check |
|---|---|
| Horizontal scroll | Is the page wider than the viewport? (Tables that don't collapse on small widths are the #1 offender) |
| Overflowing text | Any text clipped or hidden by a parent's overflow: hidden? |
| Touch targets | At the smallest configured width, are buttons reasonably tappable (≥ ~44×44 px)? |
| Card / grid stacking | At the smallest width, do multi-column grids collapse to a single column? |
| Layout-reference parity (only if this repo has a sibling layout reference for the same route — e.g. a native screen, a Figma frame) | Does the smallest-width render mirror that reference's structure? |
| Chrome collapse | Does persistent navigation chrome (sidebar/topbar) collapse or hide appropriately at smaller widths? |
| Empty state | At the smallest width, does empty-state content fit without crop? |
| Modals / panels | Do overlays become full-screen at small widths? Do they trap focus? |
| Typography scale | Is text legible at all widths? |
Print to chat AND save to .claude/responsive-runs/<slug>-<timestamp>/REPORT.md (same directory as the per-breakpoint screenshots — not in a subdirectory):
# Responsive sweep — <Route>
Captured: <YYYY-MM-DD HH:MM>
Screenshots: .claude/responsive-runs/<slug>-<timestamp>/
## Summary
- <breakpoint label> (<width>px): PASS | FAIL — <one-line>
(one row per configured breakpoint)
Overall: PASS | FAIL
## Findings
### <breakpoint label> (<width>px) — <PASS|FAIL>
- <category>: <observation>. Likely fix: <hint>. File: <path, if known>
(repeat per breakpoint)
## Suggested fixes (do not auto-apply)
- <prioritized list>
## Re-run
Once fixes applied: re-invoke `/ai-skills:responsive <route>` to verify.
Print the summary table inline plus the report path. Do not modify any application file.
REPORT.md file and mkdir -p of the screenshot directory under .claude/responsive-runs/.Tell me which route to sweep. I'll resolve the URL from your configured viewer mode, give you the screenshot checklist, and write the report once screenshots are saved.
/toolkit-seoMakes a public web screen SEO-ready by adding per-page meta tags that exactly match a configured reference's titles and robots directives. Run this whenever a public (unauthenticated) screen is added or modified.
You are the seo agent. Your job is to look up the matching page in this repo's configured reference, copy its exact meta tag values, and wire them into this repo's screen via whatever per-page head-metadata mechanism this repo uses (a hook, a head component, framework-native metadata export, etc.).
This skill is invoked with /ai-skills:seo <screen name or route>.
No config file to set up — values are detected from the repo or asked once and remembered in ~/.dev-skills.env (see CONFIG.md for the full read/write procedure and category definitions):
referenceSource — asked once, then written, required for this skill — type: "repo" is the only mode this skill supports meaningfully. If nothing's been asked/answered yet, or the answer isn't a "repo"-type reference, ask the user where the reference page lives — this skill needs one to function.appRoots — detected fresh every run — used to locate this repo's screen file.seoMechanism — detected once, then written — the name of this repo's head-metadata mechanism (e.g. a hook name like useSeoMeta, or "framework-native metadata export" if the framework has one built in). Step 3 inspects an existing screen's title/meta-tag setup the first time this skill runs and uses whatever it finds.qualityGate — detected fresh every run — lint/typecheck commands to run after wiring.<title> and meta tags.Look under the relevant appRoots layer for the named route. If this repo has more than one variant of a screen per platform/surface (e.g. a platform-specific override and a shared file), use whichever variant actually renders for the public web surface.
Some "screens" are rendered as a modal inside another screen rather than owning their own URL (the URL never changes when the modal opens). Check whether the named screen is actually a modal hosted inside another screen before wiring SEO metadata directly onto it:
grep -rn "<screen-component-name>" <appRoots layer> --include="*.tsx" -l
If the screen is a modal with no URL of its own:
Grep first — never read from line 1.
find <referenceSource.path>/views -iname "*<screen>*" -type f 2>/dev/null
grep -rln "<screen>" <referenceSource.path>/views 2>/dev/null
If the reference uses a different directory structure (e.g. /app for Next.js, /src/pages for a SPA), adapt the search path to match — do not assume /views.
If the reference uses a different templating convention than server-rendered views (a JS framework's own per-page metadata, for instance), adapt the search accordingly — the goal is the page's head/title-equivalent content, however the reference repo expresses it.
Once found, extract only the head/metadata block:
grep -n "<title>\|meta name\|meta property\|canonical\|robots" <file>
Record these values (use exactly what is in the file — do not normalize, correct, or improve):
| Tag | Value |
|---|---|
<title> |
|
<meta name="description"> |
(if present) |
<link rel="canonical" href="..."> |
(if present) |
<meta name="robots"> |
(if explicitly set; otherwise check the reference's global robots rules/robots.txt for a route-specific default) |
If the reference does not have a matching page (this screen doesn't exist there), derive a reasonable title following this repo's existing public-screen title pattern, and default robots to index, follow for a genuinely public screen. Look at 3–5 existing public-screen titles in this repo to identify the naming convention before creating a new title — do not invent a pattern.
If <SLUG>_SEO_MECHANISM isn't already in ~/.dev-skills.env, detect it once, then write it: grep a couple of existing screens for how they already set document.title/meta tags (a hook call, a head component, a framework-native metadata export), use whatever convention you find for this run, and write it to ~/.dev-skills.env as <SLUG>_SEO_MECHANISM so every future /ai-skills:seo run reads it from there instead of re-detecting — this is a one-time token cost, not a per-run one. If multiple competing patterns exist, ask once which is canonical rather than guessing. If it's already set, just confirm it still matches what's actually in the code before trusting it. (<SLUG> is the repo slug detected from git remote get-url origin — the same convention used by all other skills.)
If nothing exists yet (genuinely the first SEO-able screen in this repo), scaffold a minimal version following this repo's existing hook/utility conventions, then save that as seoMechanism too.
Illustrative pattern (a React Native + web hook pair, native no-op / web DOM implementation — adapt entirely to this repo's actual stack, do not assume this is universal):
// Native no-op variant — head metadata is a web-only concept in a cross-platform app
export function useSeoMeta(_meta: {
title: string
description?: string
ogTitle?: string
ogDescription?: string
robots?: string
}) {}
// Web variant
import { useEffect } from 'react'
function setMeta(name: string, content: string, isProperty = false) {
const attr = isProperty ? 'property' : 'name'
let el = document.querySelector(`meta[${attr}="${name}"]`) as HTMLMetaElement | null
if (!el) {
el = document.createElement('meta')
el.setAttribute(attr, name)
document.head.appendChild(el)
}
el.content = content
}
export function useSeoMeta({ title, description, ogTitle, ogDescription, robots }: {
title: string
description?: string
ogTitle?: string
ogDescription?: string
robots?: string
}) {
useEffect(() => {
const prev = document.title
document.title = title
if (description) setMeta('description', description)
if (ogTitle) setMeta('og:title', ogTitle, true)
if (ogDescription) setMeta('og:description', ogDescription, true)
if (robots) setMeta('robots', robots)
return () => { document.title = prev }
}, [title, description, ogTitle, ogDescription, robots])
}
If this repo's framework has a built-in per-page metadata mechanism (a metadata export, a head component, etc.), use that instead of hand-rolling a hook — check the framework's own documentation/convention before scaffolding anything. E.g. Next.js's metadata export or generateMetadata(), Remix's <Meta> component, Astro's <Head> component. Check the framework's docs if uncertain.
useSeoMeta({
title: '<exact title from reference>',
// description: '<if reference has a description meta>',
// ogTitle: '<if reference has og:title>',
robots: 'index, follow',
})
Run this layer's qualityGate commands (lint + typecheck). Fix all warnings/errors before reporting done.
SEO Audit — <Screen Name>
══════════════════════════════════════════════════════
Screen: <path>
Reference page: <referenceSource path>/<matching file>
──────────────────────────────────────────────────────
META TAG VALUES (from reference)
──────────────────────────────────────────────────────
title: <exact string>
description: <exact string, or "— not present">
canonical: <URL, or "— not present">
robots: <directive>
──────────────────────────────────────────────────────
CHANGES APPLIED
──────────────────────────────────────────────────────
<bullet list>
──────────────────────────────────────────────────────
QUALITY GATE
──────────────────────────────────────────────────────
<command>: ✅ pass / ❌ N issues
──────────────────────────────────────────────────────
RESULT: ✅ SEO READY / ❌ ISSUES REMAIN
robots: 'index, follow' when this repo's baseline default is noindex, nofollow for authenticated/app routes — public screens must explicitly opt in./toolkit-summaryProduce a session work-log — list tasks done grouped by logical change, NEVER enumerate file diffs or stats. Optionally post as a Jira comment. Use only when explicitly triggered ("/toolkit-summary", "summarize this session", "what did we do today"). Do not offer proactively.
Recap a Claude Code session as a list of completed tasks.
git status echo. The user can read the diff themselves.Only when the user explicitly asks. Trigger phrases: /ai-skills:summary, "summarize", "what did we do today", "session summary", "summarize the session".
Do not offer proactively at the end of every session. Most sessions end without a summary; that's fine.
Default scope: the current conversation. If the user says "today", expand to all sessions in the day for this project — but warn them this requires reading other transcripts and ask before doing it.
For the current conversation, walk the messages and identify:
Distinguish "completed" from "discussed." A 30-message research loop that ended in no code change is one task: "Researched X — concluded Y" — not 30 tasks.
A theme is a coherent unit of work a reviewer would treat as one PR or one commit. Examples:
Within a theme, every task is one line.
If the session covered exactly one theme, skip the theme heading.
One line per task. ≤ 32 words. Single statement. Lead with verb.
Pattern: - <verb> <what> — <why or where>
Good:
- Converted AI-Skills to a Claude Code plugin — added .claude-plugin/plugin.json, namespaced all 15 skills as /ai-skills:*.- Fixed baseBranch undefined variable in qa/SKILL.md — added detection from ~/.dev-skills.env.- Updated CONTRIBUTING.md with version bump policy and PR checklist.Bad:
- Modified 3 files in .claude/skills/. (count, not content)- Spent some time thinking about how skills should be structured and decided to go with a phase-aligned approach because it... (paragraph, not task)- 247 insertions, 12 deletions across 8 files. (diff stats — banned)A short "Open / not done" section. Use the same one-line format. Include:
This is the most useful part for the user — they pick up here next session.
After rendering the summary, ask: "Want me to post this as a Jira comment?"
If yes:
PROJ-1234) or ask the user.mcp__claude_ai_Atlassian_Rovo__addCommentToJiraIssue to post the summary body as a comment.<TICKET-KEY>."If Atlassian MCP is unavailable, tell the user and offer to copy the text for manual pasting instead.
## Session summary
### <Theme 1> (if more than one theme)
- <task>
- <task>
### <Theme 2>
- <task>
- <task>
### Open / not done
- <task> — blocked on <X>
- <task> — parked per user
### Suggested next
- <one or two pointers — only if obvious>
If only one theme: drop the heading, emit the bullets directly.
/toolkit-test-allRuns the full test suite for every configured layer and reports results. No code changes, no test generation. Use when the user says "/toolkit-test-all" for an overall repo health check.
You are a thin runner. When invoked, run the full test suite for every layer that has one and report results — no code changes, no test generation.
No config file to set up — values are detected from the repo (see CONFIG.md for the full procedure):
appRoots — detected fresh every run — the layers to check.qualityGate — detected fresh every run — the test-runner command per layer, read from each layer's manifest scripts.For every appRoots layer that has a test command configured under qualityGate, run it. If a layer's test command is missing or the layer has no test setup yet, skip it and note that in the report rather than erroring.
Test Suite Report
══════════════════════════════════════════
<layer> ✅ <N> passed, <N> total (<N> test files)
<layer> ✅ <N> passed, <N> total (<sub-project breakdown, if the runner has multiple projects/environments>)
<layer> ⚠️ no test command configured yet
Overall: ✅ ALL PASSING / ❌ <N> FAILING — <list failing test names/files>
If anything fails, list the specific failing test file(s) and test name(s) — don't just say "tests failed." Do not attempt to fix failures — that's the developer's call, possibly via /ai-skills:tester if the failure reveals a test that needs updating after an intentional behavior change.
/ai-skills:test-all Does NOT Do/ai-skills:tester's job, scoped to one unit of work at a time./ai-skills:qa (per-unit) or the quality gate in /ai-skills:create-pr for that.Running now — no input needed.
/toolkit-testerGenerates regression tests for a named screen/feature's I/O-bound, stateful functions — locking in current, already-verified behavior so future changes don't silently break it. Use when the user says "/ai-skills:tester <name>" after a QA pass is clean. For pure, side-effect-free functions, use /ai-skills:unit-tester instead.
You are the tester agent. Your job is to generate regression tests for a named unit of work — specifically its async/I-O-bound, stateful functions — that lock in their current, already-verified behavior so future changes don't silently break it.
/ai-skills:tester vs /ai-skills:unit-tester vs /ai-skills:qa — Read This First/ai-skills:qa checks parity: does this screen match its configured reference right now? Read-only audit, no code written./ai-skills:tester (this skill) checks regression for stateful/I-O-bound functions: async functions that hit a database, network, or other external dependency. Tests them against a mocked dependency./ai-skills:unit-tester checks regression for pure functions: sync, no I/O, deterministic output from input. Tests them directly with input/output pairs, no mocking.Run Step 0.5 below in whichever skill you invoke first — it classifies every function in scope and tells you which skill each one actually belongs to. If a function classifies as the other skill's shape, don't force it into this skill's mocking pattern — hand it off — tell the user which functions belong to /ai-skills:unit-tester instead and stop.
/ai-skills:tester treats the current implementation as the source of truth. It does not re-derive parity logic — that would duplicate /ai-skills:qa and the two would drift out of sync. If /ai-skills:qa later finds and fixes a parity bug, regenerate the affected tests afterward; don't have /ai-skills:tester second-guess /ai-skills:qa's findings.
Test-after, not test-first. The reference (or design spec) defines correct behavior — /ai-skills:qa is what verifies a match. The order is always: build → verify (/ai-skills:qa clean) → lock in (/ai-skills:tester / /ai-skills:unit-tester). Writing tests before the screen exists, or before /ai-skills:qa passes, locks in a guess instead of verified behavior.
Gate — do not skip: Before generating any test, confirm /ai-skills:qa <name> has been run for this unit of work and reported zero ❌ this session. If you cannot confirm this, tell the user explicitly:
Generating tests now will lock in the current (unverified) behavior as correct. Run
/ai-skills:qa <name>first, or confirm you want to proceed anyway.
Proceed only after the user confirms a clean /ai-skills:qa pass or explicitly overrides.
No config file to set up — values are detected from the repo (see CONFIG.md for the full procedure):
appRoots — detected fresh every run — where to locate source files and where to write test files, per layer.qualityGate — detected fresh every run — the test-runner command per layer, read from each layer's manifest scripts (used to run the suite after generating tests).If this repo has its own testing conventions doc (check for one under its .claude/ directory or root — naming varies per repo), read it before writing any test file. It will tell you the test runner, mocking conventions, and any non-obvious environment gotchas (e.g. multi-project test configs, polyfills, version pins) specific to this repo's stack. Do not assume a particular test framework (Vitest, Jest, etc.) — detect it from the repo's existing config/package files, or ask if ambiguous.
Locate the source files for this unit of work under the relevant appRoots layer(s). Check for a cached spec file under .claude/ if this repo uses one — it may already list the functions/procedures this unit of work depends on, saving you from re-deriving that from scratch.
Check for a platform/surface sibling. Some repos have more than one implementation of the same screen sharing a route/name (e.g. a native-extension convention, or a desktop/mobile shadow pair). If both exist, both need coverage — they're different components with potentially different rendering logic, even when they call the same underlying functions. Don't assume testing one covers the other.
Not every function is a good test target for this skill. Before writing tests, classify each function this unit of work touches:
| Shape | Belongs to | Why |
|---|---|---|
| Async function that hits a database/external service | /ai-skills:tester (this skill) — regression test, mocked dependency |
Needs a mocked data layer; tests the function's branching logic, not the real I/O |
| Pure function (sync, no I/O, deterministic output from input) | /ai-skills:unit-tester — unit test, no mocking |
Cheapest, fastest, most precise signal — test it directly with input/output pairs. Hand off to /ai-skills:unit-tester rather than writing it here. |
| Logic embedded inline inside a larger aggregation/pipeline or a framework lifecycle method, not extracted to its own function | Neither feasible as-is | Flag it — extracting it into a testable function is a separate refactor decision, out of scope for either skill to do unprompted |
Mechanically: grep the function signatures in the target file(s). async + calls a data layer → this skill's shape. Sync, no I/O, takes input and returns output → /ai-skills:unit-tester's shape, not this skill's — report it but don't write a regression-style test around it. Logic with no name of its own, embedded in something else → note it, don't force a test around it.
If the classification turns up unit-test candidates, tell the user: "These N functions are pure and better suited to /ai-skills:unit-tester: <list>. Run /ai-skills:unit-tester <name> to cover them — continuing here with just the regression-shaped candidates."
If classification turns up zero regression-shaped candidates for this unit of work — every function in scope is pure, embedded in a pipeline, or otherwise not regression-testable in isolation — stop here. Do not write a regression test file (and don't write UI component tests either, if the screen itself has no I/O-bound logic of its own worth locking in beyond what /ai-skills:unit-tester already covers). Report plainly:
No regression-test candidates found for
<unit of work>. Every function in scope is either pure (unit-test shape — see/ai-skills:unit-tester) or not extracted into a standalone testable function. No regression tests were generated.
List which category each function fell into so the user knows this was a deliberate finding, not a skipped step.
Before writing any frontend test, audit every interactive element in the unit of work's component files — pressable cards, buttons, chips, action rows, detail-panel triggers — and check whether each has a stable testID prop.
Why this matters: E2E tests (Playwright, Detox, etc.) that target elements without a testID must fall back to fragile role/text/position selectors. These break silently when copy changes or layout shifts. A testID is the only stable, intent-expressing hook for E2E targeting.
Grep the component files for TouchableOpacity, Pressable, and Button usages. For each one, check whether a testID prop is present.
grep -n "TouchableOpacity\|Pressable\|Button" <component-file> | grep -v "//"
Then check which of those already have testID:
grep -n "testID=" <component-file>
| Element | Has testID? |
Action |
|---|---|---|
| Pressable/button that E2E needs to target | ✅ Yes | Use it as-is in the test |
| Pressable/button that E2E needs to target | ❌ No | Add testID to the component first, then write the test against it |
| Decorative / non-interactive element | N/A | Skip — no testID needed |
Use kebab-case, screen-scoped IDs. Format: <screen>-<element>. Examples:
calendar-event-blockreservations-rowfilter-type-chipmaintenance-block-cardnotice-create-buttonAfter auditing, output this block before writing any test code:
testID Audit — <unit of work>
──────────────────────────────────────────
✅ Already present:
• <component-file>: testID="<id>" (line N)
🔧 Added now (component updated):
• <component-file>: testID="<id>" added to <element description> (line N)
⏭️ Skipped (decorative / non-interactive):
• <element description> — no testID needed
If you added any testIDs to component files, include those file changes in the final commit alongside the test files. Do not write a test that targets an element by testID before confirming that testID is actually in the component.
Check whether this repo's backend test runner is already installed — a test framework in package.json/manifest devDependencies, a config file present, and at least one existing test file to use as a mocking-pattern reference. Do not assume it's there. A repo with zero existing tests has none of this, and writing the first test file is a different task than adding to an established suite.
If a test runner exists: check the existing test suite for a reference pattern (a file that already mocks this repo's data layer) before inventing a new mocking approach — reuse the existing convention rather than introducing a second one.
If no test runner exists yet: this is a one-time setup task, not something to skip past silently.
Identify what this repo's stack needs (the test framework that matches its language/runtime — e.g. Vitest for a Node/TS service, the equivalent for whatever this repo actually runs).
Check for an existing testing-conventions doc in this repo first — it may already specify a chosen framework even if it's not installed yet.
Ask the user before adding new devDependencies or config files. Tell them what you'd install and why, and wait for confirmation — this is a one-time, repo-wide decision (test framework choice, config shape), not something this skill should decide unilaterally on a single invocation.
Once confirmed and installed, write the first test file using a minimal, idiomatic mocking setup for this repo's data layer — this becomes the reference pattern future runs reuse.
For each regression-shaped (async, I/O-bound) service function/procedure this unit of work calls, write a test block in the appropriate test file for that domain.
Build (or reuse) a mock data-layer factory matching this repo's existing pattern.
Import the function under test directly from its module — don't go through a network/transport layer if the repo's convention is to test the function directly.
Write cases covering:
If a test file for this domain already exists, add new test blocks rather than duplicating coverage already present.
Run this layer's test command from qualityGate and report pass/fail.
Scope is the function/service layer when one exists, separately from thin pass-through layers (e.g. a router/controller that just validates input and calls a service) — don't write redundant tests for a pass-through layer that has no logic of its own.
Exception — no separate service layer for this domain, logic lives directly in the pass-through layer. Per this skill's own rule (test the current implementation, don't refactor it), test that layer directly using whatever direct-invocation mechanism the framework provides, rather than moving the logic into a new function as a side effect of adding tests — that's a refactor, out of scope here.
UI component tests are regression-shaped almost by definition (they render against mocked data-fetching state) — they belong here, not in /ai-skills:unit-tester, even though the screen itself may call some pure helper functions internally (those pure helpers go through /ai-skills:unit-tester separately).
Check whether this repo's frontend test stack is already set up (test runner present in package.json devDependencies, a config file present). If it is, do not repeat the bootstrap — just write tests using the existing setup.
If it is not yet set up, this is a non-trivial one-time task: research what this repo's framework needs (e.g. a dual test-environment setup if the repo ships both a DOM-rendered and a native-rendered variant from the same codebase), check for an existing testing-conventions doc in this repo first, and ask the user before introducing new devDependencies or config files if the shape isn't obvious from how the repo already builds/runs.
If the unit of work has more than one platform/surface variant (Step 0), generate a test file for each variant — they are separate components and need separate coverage.
For each variant:
.test.tsx vs. .web.test.tsx)./ai-skills:qa already confirmed correct for this unit of work:
Run this layer's test command, scoped to the new file, and report pass/fail.
After generating tests, run the full test command for every layer touched (from qualityGate in config). Report exact pass/fail counts. Do not mark a unit of work "tested" if any suite has failures — fix the test (not the implementation, unless the test caught a real bug) until it passes.
/ai-skills:tester Does NOT Do/ai-skills:unit-tester instead./ai-skills:qa's job exclusively.testID. Step 0.75 must run first — add the testID to the component, then write the test. Never fall back to position/text/role selectors for pressable elements that QA needs to target deterministically.Tell me which unit of work to write regression tests for. I will confirm /ai-skills:qa status, classify candidates (Step 0.5) and flag any pure-function candidates for /ai-skills:unit-tester, then generate tests and report pass/fail counts.
/toolkit-unit-testerGenerates unit tests for a named screen/feature's pure, side-effect-free functions — direct input/output assertions, no mocking. Use when the user says "/ai-skills:unit-tester <name>" after a QA pass is clean, or when /ai-skills:tester's classification step flags pure-function candidates. For async/I-O-bound functions, use /ai-skills:tester instead.
You are the unit-tester agent. Your job is to generate unit tests for a named unit of work's pure functions — functions with no I/O, no side effects, and deterministic output from input — asserted directly with input/output pairs, no mocking.
/ai-skills:unit-tester vs /ai-skills:tester vs /ai-skills:qa — Read This First/ai-skills:qa checks parity: does this screen match its configured reference right now? Read-only audit, no code written./ai-skills:tester checks regression for stateful/I-O-bound functions: async functions hitting a database/network/external service. Needs a mocked dependency./ai-skills:unit-tester (this skill) checks regression for pure functions: sync, no I/O, deterministic. No mocking — call the function, assert the output.Run Step 0.5 below (shared with /ai-skills:tester) in whichever skill you invoke first — it classifies every function in scope and tells you which skill each one actually belongs to. If a function classifies as the other skill's shape, don't force it into this skill's no-mocking pattern — hand it off — tell the user which functions belong to /ai-skills:tester instead and stop.
/ai-skills:unit-tester treats the current implementation as the source of truth, same as /ai-skills:tester. It does not re-derive parity logic. If /ai-skills:qa later finds and fixes a parity bug in a pure function, regenerate the affected unit test afterward.
Test-after, not test-first. The order is always: build → verify (/ai-skills:qa clean) → lock in (/ai-skills:unit-tester / /ai-skills:tester). Writing tests before the function's behavior is verified locks in a guess instead of verified behavior.
Gate — do not skip: Before generating any test, confirm /ai-skills:qa <name> has been run for this unit of work and reported zero ❌ this session. If you cannot confirm this, tell the user explicitly:
Generating tests now will lock in the current (unverified) behavior as correct. Run
/ai-skills:qa <name>first, or confirm you want to proceed anyway.
Proceed only after the user confirms a clean /ai-skills:qa pass or explicitly overrides.
No config file to set up — values are detected from the repo (see CONFIG.md for the full procedure):
appRoots — detected fresh every run — where to locate source files and where to write test files, per layer.qualityGate — detected fresh every run — the test-runner command per layer, read from each layer's manifest scripts (used to run the suite after generating tests).If this repo has its own testing conventions doc, read it before writing any test file — it will tell you the test runner and any non-obvious gotchas. Do not assume a particular test framework — detect it from the repo's existing config/package files, or ask if ambiguous.
Locate the source files for this unit of work under the relevant appRoots layer(s). Check for a cached spec file under .claude/ if this repo uses one.
Identical classification step to /ai-skills:tester's Step 0.5 — run it once per unit of work, not twice, if both skills are being invoked back to back in the same session.
| Shape | Belongs to | Why |
|---|---|---|
| Pure function (sync, no I/O, deterministic output from input) | /ai-skills:unit-tester (this skill) |
Cheapest, fastest, most precise signal — test it directly with input/output pairs |
| Async function that hits a database/external service | /ai-skills:tester — regression test, mocked dependency |
Out of scope here — hand off rather than force a mock-free test onto an I/O-bound function |
| Logic embedded inline inside a larger aggregation/pipeline or a framework lifecycle method, not extracted to its own function | Neither feasible as-is | Flag it — extracting it into a testable function is a separate refactor decision, out of scope for either skill to do unprompted |
Mechanically: grep the function signatures in the target file(s). Sync, no I/O, takes input and returns output, no calls to a data/network layer → this skill's shape. async + calls a data layer → /ai-skills:tester's shape, not this skill's. Logic with no name of its own, embedded in something else → note it, don't force a test around it.
If the classification turns up regression-test candidates, tell the user: "These N functions are I/O-bound and better suited to /ai-skills:tester: <list>. Run /ai-skills:tester <name> to cover them — continuing here with just the pure-function candidates."
A function only belongs in this skill if it is genuinely pure: same input always produces the same output, no reads/writes to a database, network, filesystem, global mutable state, or wall-clock/random source. If a function reads the current date/time or a random value internally without it being passed as an argument, it is not pure — flag it as a "neither feasible as-is" candidate (or suggest refactoring it to accept that value as a parameter, as an explicit out-of-scope recommendation, not something this skill does unprompted).
If classification turns up zero pure-function candidates for this unit of work — every function in scope is I/O-bound, embedded in a pipeline, or otherwise not unit-testable — stop here. Do not write a unit test file. Report plainly:
No unit-test candidates found for
<unit of work>. Every function in scope is either I/O-bound (regression-test shape — see/ai-skills:tester) or not extracted into a standalone testable function. No unit tests were generated.
List which category each function fell into so the user knows this was a deliberate finding, not a skipped step.
Check whether this repo's test runner is already installed (a test framework in package.json/manifest devDependencies, a config file present). Do not assume it's there — a repo with zero existing tests has none of this, and unit tests for pure functions are sometimes the very first tests a repo gets, since they need no mocking infrastructure to start.
If a test runner exists: check the existing test suite for a unit-test reference pattern (a file with no mocking, just direct calls) before inventing a new file structure.
If no test runner exists yet: this is a one-time setup task. Identify what this repo's stack needs, check for an existing testing-conventions doc first, and ask the user before adding new devDependencies or config files — wait for confirmation before installing anything. Once confirmed, write the first test file using a minimal setup; this becomes the reference pattern future runs reuse. (If /ai-skills:tester has already bootstrapped this repo's backend test runner in a prior run, reuse that setup rather than asking again — check before assuming nothing exists.)
For each pure-function candidate:
null/undefined handling if the function's signature allows them, malformed input if the function is expected to validate rather than assume well-formed input.toBe/toEqual/toMatchObject (or this repo's test framework's equivalent) — prefer exact-value assertions over loose truthy checks, since the entire value of a unit test is precision.qualityGate and report pass/fail.Do not mock anything in this skill. If you find yourself reaching for a mock, the function under test isn't actually pure — stop and re-classify it as a /ai-skills:tester candidate instead of mocking around the issue here.
After generating tests, run the full test command for every layer touched (from qualityGate in config). Report exact pass/fail counts. Do not mark a unit of work "tested" if any suite has failures — fix the test (not the implementation, unless the test caught a real bug) until it passes.
/ai-skills:unit-tester Does NOT Do/ai-skills:tester's candidate, not this skill's./ai-skills:tester instead./ai-skills:qa's job exclusively.Tell me which unit of work to write unit tests for. I will confirm /ai-skills:qa status, classify candidates (Step 0.5) and flag any I/O-bound candidates for /ai-skills:tester, then generate tests and report pass/fail counts.
/toolkit-verify-testsFinds and runs the existing test file(s) for one named unit of work — nothing else. Use when the user says "/toolkit-verify-tests <name>" to quickly re-run a unit's existing coverage.
You are a thin runner, scoped to one unit of work. When invoked, you find its existing test file(s) and run them — nothing else.
/ai-skills:verify-tests vs /ai-skills:tester vs /ai-skills:test-all — Read This First/ai-skills:tester generates new regression tests for a unit of work that doesn't have them yet (or extends existing ones)./ai-skills:test-all runs everything, every test file, every layer — for overall repo health./ai-skills:verify-tests runs only the test file(s) for one unit of work that already has them — fast, scoped, no new tests, no code changes.If the unit of work has no test file yet, say so and suggest /ai-skills:tester <name> — do not generate tests yourself.
No config file to set up — values are detected from the repo (see CONFIG.md for the full procedure):
appRoots — detected fresh every run — used to locate test files if a coverage tracker isn't available.qualityGate — detected fresh every run — the test-runner invocation shape per layer, read from each layer's manifest scripts.Check whether this repo has a coverage tracker doc (a file that lists which units of work have test coverage and where). If one exists, match the requested name against it first — faster and more reliable than guessing paths.
If no tracker exists or the unit isn't listed, locate test files the way /ai-skills:qa//ai-skills:tester do: search under the relevant appRoots layer for a colocated or domain-named test file matching this unit of work.
If you still find nothing, tell the user no test file exists for this unit of work and stop — suggest /ai-skills:tester <name> instead of writing anything.
Run only the test runner invocation scoped to the specific file(s) found — not the full suite. Use this layer's qualityGate test command as the base invocation, with the specific file path appended per that test runner's own syntax for scoping to one file.
Path-matching gotcha: if this repo uses a file-based router with route-group syntax (literal parentheses in folder names — e.g. Expo Router's or Next.js App Router's (group)/ convention), most test-runner path matchers treat those parens as regex groups and fail to match. Use a parenthesis-free substring of the path instead of the full path when scoping to one file. Repos without this routing convention won't hit this at all.
Verify Tests — <Unit of Work>
══════════════════════════════════════════
<layer> ✅ <N> passed, <N> total (<test file>)
<layer> ✅ <N> passed, <N> total (<test file>)
Overall: ✅ ALL PASSING / ❌ <N> FAILING — <failing test names>
If a unit of work only has one layer covered, state that plainly rather than reporting a phantom 0/0 for the missing layer.
If anything fails, list the specific failing test name(s) — don't just say "tests failed."
/ai-skills:verify-tests Does NOT Do/ai-skills:tester's job./ai-skills:test-all. This is scoped to one unit's existing test file(s) only./ai-skills:qa or /ai-skills:create-pr's quality gate for that.Tell me which unit of work to verify. I will look up its existing test file(s), run them, and report pass/fail.
This is the Codex-only compatibility boundary for Scar. It links skills from `core/skills` into `~/.agents/skills`, merges the shared hook roster only into the selected project's `.codex/hooks.json`, updates a managed block in `~/.codex/AGENTS.md`, and can register MCP through the Codex vendor CLI. Claude configuration, Claude skills, shared installers, trust controls, and model/provider settings are outside this adapter.
Turns raw incident records into transferable lessons — the L0 → L1 compiler. Scar holds ~1,000 records at roughly 2.3M tokens, a layer no agent can ever load; the ~99 lessons distilled out of it cost about 3k. That ratio is the whole design: distillation is paid once, at write time, instead of every session forever.
Arm delays and failures inside your own dev server, at runtime, scoped to the requests you're testing. Zero dependencies. Fastify and Express adapters; the core is framework-agnostic.
The inbox where agents report on Scar itself, so it can be judged as a system rather than document by document. The per-document counters answer "was this lesson good"; they structurally cannot answer "was Scar good", because every such judgement is a *ratio over queries* ("two thirds were noise", "0 of 7 useful hits came from the project layer") and the query text was discarded the moment recall returned. That gap is why this exists.
Two of Scar's rules — "consult before you edit" and "stop reasoning and observe after 3 disproven theories" — were stated in prose in files an agent reads once and can silently fail to act on for an entire session. Both were observed failing exactly that way in real use. These hooks move the rule out of the agent's judgment and into the Claude Code harness, which cannot forget to check.
Scar, exposed as MCP tools — so an agent working in *any* project can search entries, read rules, and pull up skills without a human copy-pasting context in. Read-only: the one write path stays the agent skills + the editor (`core/AGENTS.md`, "Contract with the tooling").
A legacy app is being replaced. The legacy app kept receiving fixes while the replacement was being built. **Which of those fixes never made it across?**
The retrieval side: given a sentence describing what you are about to do, return the few lessons worth reading and nothing else. BM25 over the distillate, zero dependencies — the searched corpus is ~134 documents, not 1,000, small enough that keyword ranking beats embeddings on staleness, dependencies and explainability.
brain-index.mjsCLI entry point for the folder-index regeneration. The work itself is `lib/brain-index.mjs`, because `tools/distill/lib/lesson.mjs` calls it after writing a lesson — and a `lib/` reaching into another package's top-level script is the inversion `scripts/layering-lint.mjs` refuses. The fail-CLOSED im
brain-view.mjsbrain-view — renders the whole brain into one self-contained HTML page. Zero dependencies, no assets, no network. Frontmatter IS the database: this reads the same fields brain-index and frontmatter-lint use, so there is exactly one source of truth and the viewer can never disagree with the markdown
build-leak-check-selftest.mjsPlanted-violation coverage for build-leak-check.mjs — the gate's gate. WHY THIS EXISTS. The leak gate grew structural handlers (PNG text chunks on 2026-08-20, font name/metadata tables the same day), and every handler is a place a leak could slip through if it rots. AGENTS.md: "a gate that has only
build-leak-check.mjsGrep the BUILT public output against the denylist — the publishing gate, as code. WHY THIS EXISTS. AGENTS.md carries the constraint in prose: "A new publishing surface inherits none of the old surface's filters. Two real leaks were found this way. Before adding one, check it against the denylist by
citation-lint.mjsA lesson or rule id written into source must resolve to something that exists. WHY THIS EXISTS (2026-08-25). `frontmatter-lint` fails on a dangling `[[wikilink]]` inside the corpus. Nothing checked the other direction — an id cited FROM `tools/`, `scripts/` or `ui/`, in the comment explaining why th
cli-selftest.mjsPlanted cases for cli.mjs — the router, proved without running anything it routes to. WHY IT TESTS `route()` AND NOT THE COMMANDS. The public surface includes a web server that runs until interrupted, a distiller that spends money on the Anthropic API, and a verify pass that takes a minute. A suite
constraints-lint.mjsEvery standing constraint in the always-loaded files must name the mechanism behind it. WHY THIS EXISTS (2026-09-11). Root AGENTS.md sat 15 tokens over its advisory cap, core/AGENTS.md 2 tokens under its own, and every lesson or decision that touched either arrived as the same question for the owner
contrast-lint.mjscontrast-lint — the gate that lets `ui/` carry two palettes. WHY THIS EXISTS. `ui/app/globals.css` once said "ONE THEME. There is no light mode and no toggle", and the reason it gave was correct: a second palette is a second thing to keep right on every change, and the one that was there had shipped
cutover-readiness.mjsTHE PRE-CUTOVER GATE. Proves the runtime that will exist AFTER the authority switch, without touching a single live project config. The path that matters after cutover is not the desktop window. It is: Claude/Codex hook → packaged launcher → bundled core → ~/.scar → gate/recall → evidence written ba
decisions-lint.mjsDECISIONS.md's numbering is the citation scheme — keep it one sequence. WHY THIS EXISTS (2026-09-05). `STATUS.md` and `AGENTS.md` cite decisions by bare number — "(85)", "(104)", "DECISIONS.md 97" — and the file's own header says to grep it rather than read it. That only works while the numbers form
denylist.mjsThe identifier denylist, from the terminal — the same `editDenylist` the app's Sanitization card calls, so the two paths share one gate (`AGENTS.md`, "two paths, one gate"). npm run denylist list the terms npm run denylist -- add <term> [--word] preview: wh
desktop-boundary-lint.mjsThe Electron shell at `desktop/` is allowed to be a UI and nothing else. This gate enforces that at the source, mechanically, on every run — not as a design note someone reads once and a later change quietly violates. WHY THIS EXISTS. `desktop/` talks to the brain by spawning `tools/dashboard/serve.
electron-leak-check.mjsDoes the packaged desktop app carry anything from the private tier? WHY A SECOND LEAK CHECK. `scripts/build-leak-check.mjs` walks `DEPLOYED = ['server','static']` under `ui/.next`. An Electron bundle is nowhere near that path, so it inherits none of those filters — precisely the standing constraint
engine-parity.mjsDoes the published engine still match the source it was extracted from? WHY THIS EXISTS. `brain/tools/` and `brain/scripts/` were copied into the `brain-engine` plugin's `engine/` tree on 2026-08-22 (see the marketplace repo's EXTRACTION.md). That file says "re-run this scan after any further copy f
firing-record-lint.mjsEvery component that puts text into somebody's context must record that it did. WHY THIS EXISTS (2026-09-18 audit). `signature-scan/scan.mjs` called `markRan` and never `markFired`. It worked perfectly and had, by its own census, 1,828 lifetime matches and 168 dismissals — and contributed **zero** r
fix-yaml-quoting.mjsRepairs frontmatter scalars that this repo's permissive parser accepts but real YAML rejects. node scripts/fix-yaml-quoting.mjs # dry run — shows every line it would change node scripts/fix-yaml-quoting.mjs --write # applies Why this exists: the migration wrote titles like `BUG-001:
frontmatter-lint.mjsValidates every entry's frontmatter against core/SCHEMA.md. Keeps grep-scoping reliable as the brain grows — an entry with malformed or missing frontmatter is invisible to brain-index and to every /scar-* skill that searches by status/type/tags.
gate-lint.mjsWeak-check lint — catches a test SHAPED like the failures core/rules/verification's own scars record, before it ships and joins them. WHY THIS EXISTS (2026-08-27). `core/rules/verification` carries 13 scars, several of which are the same failure: a check that reports success without having verified
index-fresh.mjsFail when `.recall/index.json` is older than the corpus it claims to index. WHY THIS EXISTS (2026-09-04). `recall-benchmark` and `reachability` are the only checks in `verify` that grade RETRIEVAL, and both read `.recall/index.json` — derived, gitignored, and rebuilt only when somebody runs `npm run
init-selftest.mjsWhere a standalone installation would be created, and what `scar init` refuses. WHY THE COLLISION CASE IS THE FIRST ONE. The default data directory is a ONE-WAY DOOR: a later release cannot move it without stranding a corpus and a ledger that regenerate from nothing. The first version of `defaultDat
init.mjs`scar init` — create a standalone SCAR installation (STANDALONE-PLAN.md Phase 8). WHAT THIS IS FOR. Until now the only place SCAR's data could live was the checkout it was cloned into. A packaged application bundle is read-only, so it needs somewhere else to write — and that somewhere has to exist,
install-hooks.mjsCopies the pre-commit and pre-push hooks into .git/hooks/. Not tracked by git itself (hooks never are), so this script is what makes the gate real after a fresh clone — run once after git init/clone.
install-skills.mjsSymlinks core/skills/* into ~/.claude/skills/ so the /scar-* commands are available in every project, not just this one. core/skills/ stays the single source of truth — re-run this after any skill is added or edited (in particular, /scar-skill runs this automatically).
layering-lint.mjsTHE LIBRARY LAYER MUST STAY A LAYER — caught mechanically. WHY THIS EXISTS. On 2026-09-22 a review asked whether this codebase had gone to spaghetti. The answer was no, and the evidence was strong: ONE `parseFrontmatter`, ONE `loadCorpus`, one denylist matcher, and ~150 of ~180 cross-package import
lesson-tail.mjs`npm run lesson-tail -- career/lessons/<file>` — did a lesson's verify tail ACTUALLY land? WHY THIS EXISTS (2026-09-16). `/scar-lesson` tells a product session to hand everything after `sanitize-check` to a background subagent, so the tail's ~105s-per-iteration fix-and-rerun loop costs the user's cl
lessons.mjs`npm run lessons` — the human browsing surface, computed at read time. WHY THIS EXISTS (2026-08-31). The generated INDEX.md files were derived state that was COMMITTED — the only such artifact in the repo, where `.recall/` and `.distill/` are gitignored and rebuilt. That combination bought the worst
new-entry.mjsScaffolds a new entry from its template with pre-filled frontmatter, so entries are born schema-valid instead of relying on the author to remember every field. Usage: node scripts/new-entry.mjs <type> "<title>" --dir <path> [--slug custom-slug] Types: bug | ticket | decision | snippet | concept
pause.mjs`npm run pause` — stand the consultation gates down for a stated, expiring window. The mechanism, the four properties that keep it honest, and the incident that forced it are in `tools/hooks/lib/pause.mjs`. This is only the operator's end of it. npm run pause pause for 60 minutes
repo-about.mjsThe GitHub repo's "About" field is hand-typed text living outside this repo entirely — GitHub settings, not a file — so nothing has ever regenerated it. It was found stale: "993 project records (~2.2M tokens)" while the real corpus (2026-08-12) is 1,057 records, and README.md's own headline number h
retrain.mjs`npm run retrain` — refit the judge-side models from whatever labels have arrived, and say only what changed. SCAR Model v1 tools/judge/train.mjs gate-verdict judge (upheld / dismissed) worker escalation tools/judge/escalation-train.mjs escalate a session to a stronger mod
root-lint.mjsONE IDENTIFIER, TWO CONSUMERS WITH OPPOSITE NEEDS — caught mechanically. WHY THIS EXISTS. `BRAIN_ROOT` is the CODE root. `DATA_ROOT` owns `.recall/` and `.distill/`; `CORPUS_ROOT` owns `career/ core/ projects/ feedback/` and the private lists. In a checkout all three are the same string, so using th
sanitize-check.mjsDenylist gate. Two tiers, deliberately different in scope: secrets-denylist.txt checked on EVERY path — there is no context where a live key belongs. sanitize-denylist.txt client/employer identifiers; checked on career/ and core/ only (the tiers promotable to public)
shadow.mjsSeed a SHADOW installation — the same code, pointed at a copy of the state and corpus. WHY THIS EXISTS (STANDALONE-PLAN.md Phase 6). The standalone successor must be developed against the real system without being able to disturb it, and the obvious way — copy the whole checkout — is the one this re
site-claims.mjssite-claims — the gate that keeps ui/ from describing a system that no longer exists. WHY THIS EXISTS. On 2026-09-28 the site said "No embeddings" on the day an encoder shipped (decision 250), "Six gates" beside a roster of twelve, "gives up after two denials" where the caps run one to four, and lis
skill-roster.mjsThe `/scar-*` command surface, pinned. AGENTS.md has said "9 `/scar-*` skills is the ceiling, not the baseline" since the surface was nine. It was carried entirely as prose, and prose lost twice: /scar-harden took it to ten (2026-09-02) and /scar-diagnose to eleven (2026-09-04), each time with a wr
snapshot-configs.mjsTHE ROLLBACK PRIMITIVE (STANDALONE-PLAN.md §14). Snapshot every agent config file this brain is registered in, so that undoing a cutover is a FILE COPY rather than a reconstruction. WHY IT IS A SCRIPT AND NOT A SHELL ONE-LINER. The first attempt was an inline `while read` loop with `cp`. Something c
static-lint.mjsDoes anything in this repo READ the code? Until 2026-09-03, no. WHY THIS EXISTS. `npm run verify` ran 23 checks and every one of them was behavioural: it loaded the corpus, ran the ranker, planted a violation in a gate, priced the ledger. Not one of them parsed a source file. So 149 JS/TS files and
stores-selftest.mjsDoes `scripts/lib/stores.mjs` still describe the stores that actually exist? WHY THIS EXISTS. A roster is a hand-maintained list, and this repo's own history says a hand-maintained list drifts and then certifies something incomplete — `wiring.mjs` and `harnesses.mjs` both carry that scar. A roster o
sync.mjs`scar sync` — refresh the standalone CANDIDATE from the live repo brain. One-way, always. THE MODEL THIS ENFORCES (owner, 2026-09-21): repo SCAR + its brain AUTHORITATIVE — the only canonical writer │ one-way import ▼ Scar.app + ~/.scar CANDIDATE — current, never canonical "Not
system-map.mjsA generated, human-readable map of the whole brain — what each piece is, and why it exists. WHY THIS EXISTS. The system grew to ~9 skills, 8 rules, 4 gates, a 4-layer pipeline, and a growing corpus, and the person who has to keep it all in their head asked for a standing explanation that can't go s
token-budget.mjsGuards the standing per-session token cost — the number this whole system exists to keep small. The brain's pitch is "2.2M tokens of experience for ~4k tokens a session." That claim has two components, and only one of them was enforced: the kernel is capped by `--kernel-budget` in tools/recall/buil
upkeep.mjs`npm run upkeep -- --decisions N --repairs M [--note "…"]` — record what one /scar-update run found. `npm run upkeep` alone reads the log back, tallied by release. WHY THIS EXISTS (2026-09-11). Asked whether the system is stable, nobody could answer, because "stable" had never been defined. The defi
verify.mjsThe verify chain, quiet on pass and complete on failure. WHY THIS EXISTS. Measured 2026-08-14 on a real run: one `npm run verify` emits ~2,713 tokens, of which **42% is selftest `ok` lines and 17% is the benchmark's per-query reject/noisy enumeration** — ~59% carrying no information whatsoever when