Evolving CLAUDE.md
Wires CLAUDE.md to maintain a living Decisions & Learnings log that the assistant updates as the project evolves — without letting the file bloat. CLAUDE.md is read into context every turn, so the goal is a small, well-pruned set of durable decisions, with everything else linked out or archived. Three hooks keep the log healthy automatically. It complements the built-in init skill: init describes the codebase, this adds the mechanism that lets the file grow safely.
Trigger it
/evolving-claude-md:evolving-claude-md
This skill is model-invokable — most often you just ask in natural language: "make CLAUDE.md evolve", "add a self-updating decisions log to CLAUDE.md", or "CLAUDE.md is getting too big, compact it".
To compact the log on demand, run the dedicated command:
/evolving-claude-md:compact # plan, propose, apply on approval
/evolving-claude-md:compact plan # just show the plan
/evolving-claude-md:compact yes # apply without waiting
When to use it
-
"add learning mechanism to CLAUDE.md", "make CLAUDE.md evolve", "self-updating CLAUDE.md"
-
"decisions log", "ADR-style notes in CLAUDE.md", "project memory in CLAUDE.md"
-
"compact CLAUDE.md" / "CLAUDE.md is getting too big"
-
A
CLAUDE.mdexists but has no mechanism to keep itself current -
Not for plain codebase documentation — use
initfor that
What it does
Automatic hooks
The plugin ships hooks/hooks.json (pathed via ${CLAUDE_PLUGIN_ROOT}), so the hooks register on enable with nothing to add to settings:
-
SessionStart —
audit-claude-md.pychecks the D&L section + whole-file size + staleness and injects a recommendation when something needs attention; silent when healthy. Since 1.4.0 staleness covers three rot classes, grounded against the tree rather than reasoned about: a cited artifactgit grepcan no longer find, a version pin every build file now contradicts ("Spring Boot 4.0.5" vs a pom that says 4.0.7 — prefix-compatible pins and examples stay silent; since 1.9.0 the build files read are Maven, Gradle, npm, Go, Cargo, Python, .NET, Bundler and Composer, plus toolchain pin files such as.nvmrcand.tool-versions), and a sequence fact the tree has moved past ("latest isV27`" once `V28exists). The predicates live in a vendorable stdlib core,freshness.py, calibrated across 29 real repos to zero false positives. -
PreToolUse on
Write|EditofCLAUDE.md—lint-claude-md.pyrejects new entries that lack a date ortopic-tag, or whose body runs past 500 characters (200 is the target; past that, the detail belongs indocs/decisions/). -
InstructionsLoaded —
record-loads.pyrecords which instruction files actually loaded, so a rule whose globs never match, or a CLAUDE.md that never loads, becomes visible. -
PostCompact — re-runs the audit so the assistant sees the current state without paying to re-read the whole file.
-
PostToolUse and Stop — capture triggers that nominate a possible learning from a commit message or a session’s end. Off by default.
What’s new in 1.8.0 — compact as a command
/evolving-claude-md:compact runs the four downward pressures as one reviewed edit, in the order that keeps lessons safe: supersede gaps, release mirrors (keep what a release taught, drop what it shipped), same-day merges, graduation into Conventions/Gotchas, mega-entry splits, stale checks — and the archive last. Archiving first is the mistake an ad-hoc compaction makes: lessons that should be standing rules end up in the quarterly file instead of the always-loaded one.
The plan for this repository’s own CLAUDE.md — a real run of the planner the command reads.
Every edit is drafted with its exact text and shown before anything is written; plan stops there, yes skips the wait. The data comes from a read-only planner, compact-claude-md.py. Its graduation list finds a tag repeated three or more times — which misses lessons that recur by theme under different tags — so it also lists every entry about to age out, for the session to review first.
Fixed along the way: archive-decisions.py was unsafe to run twice. Rebuilding the section from dated entries only, it dropped the section’s format note, and an entry absorbed any undated line below it — so a second run archived the first run’s teaser and left a new one reporting only its own count. It now parses the section line by line, keeps everything it doesn’t archive in place, and merges into an existing quarter teaser.
What’s new in 1.7.0 — evidence and invalidation
1.6.0 gated what gets written. 1.7.0 borrows the two mechanisms an 80-entry survey of instruction-file upkeep tools showed were missing here — chosen because exactly half of those 80 document no staleness handling at all.
Load evidence. A new InstructionsLoaded hook records which instruction file
actually loaded, in which session, and why (session_start, nested_traversal,
path_glob_match, include, compact), appending to
.claude/evolving-claude-md/loads.jsonl. Every other check in the plugin
reasons about what the file says; this one records what the harness did, and
that makes two failures checkable which no content audit can see:
-
a
.claude/rules/*.mdwhosepaths:globs never match anything that gets read — the rule is well-written and simply never consulted, so its content is not the defect; -
a
CLAUDE.mdthat never loads at all, because Claude Code loads one up to 4 MiB and silently skips a larger file, and a file in an unread path is skipped just as quietly. The repo looks configured and isn’t.
Both wait for load_min_sessions (default 5) of recorded evidence: "hasn’t
loaded yet" and "never loads" are indistinguishable on day one.
Supersede links. An entry may end Supersedes: YYYY-MM-DD topic-tag. The
audit then reports a target that is still unstruck — a replaced rule still
reading as current — or one that doesn’t exist, catching a typo’d date or tag.
The lint exempts the link from the body cap, because charging bookkeeping
against the prose budget would price the link out of use. The idea comes from
temporal knowledge graphs, which record when a fact stopped being true instead
of leaving both versions standing.
Also corrected: the docs claimed in three places that the lint hook enforces a 200-character body cap. It enforces 500, so entries between the two had been passing silently against a documented rule. Both numbers are now stated, with the reason they differ — 200 keeps the log scannable, 500 is where an entry is provably a design doc. Enforcement was deliberately left at 500; tightening it would start rejecting entries in every consuming repo.
What’s new in 1.6.0 — the inlet
1.5.0 gave the log a write path. 1.6.0 asks whether a given write has earned its place, because every mechanism before it was downstream of a write nothing gated: four pruning pressures, zero inlet control.
The bar decides; the six triggers only nominate. Write the candidate as a sentence that is true about the repo tomorrow, then name what a future session does differently knowing it. If the only true sentence is "we did X", there is no entry — the changelog, the tickets and git log already hold it. A delivery earns an entry only when it taught something: a constraint, a trap, a reversal. Most sessions nominate zero, and zero is the right answer to a session that merely shipped. The session-end capture prompt carries the bar, and what NOT to log gains the durable-artifact rule: where a repo keeps a changelog, ADRs or tickets, an entry is earned only by the part that isn’t in them.
New audit check — CHANGELOG MIRROR. Measures that duplication directly: entries
naming a version CHANGELOG.md already documents. It found 63% on the repo that
motivated the change and stays silent on logs of real learnings.
Parallel lanes spool instead of colliding. Measured on one 2,400-commit repo
running worktree lanes: 86 of 257 non-merge CLAUDE.md edits were made on lane
branches, and all 50 merge commits touching the file were resolving them. A lane
— a linked git worktree, or any non-default branch with "lane_spool": "branch"
— now routes its nominations to .claude/evolving-claude-md/incoming/<branch>.md.
One file per writer, so the conflict class disappears rather than being
resolved. Then capture-triggers.py fold prints every lane’s nominations
together with cross-lane near-duplicates flagged, for one writer to curate on the
default branch — five lanes on one epic learn the same lesson five times, and
only a reader holding all of them can dedup it. An ordinary feature branch is
deliberately not a lane: one writer at a time merges cleanly.
Obsolescence findings name the edit. MISSING from the layout and OBSOLETE in
the layout now prescribe the change to make — add the line, or strike it — because
a flag a reader must translate into an action is a flag that gets skimmed past.
A fifth check was written and cut before shipping: tag-uniqueness looked like a clean symptom until calibration showed the fleet’s healthiest log runs 100% distinct tags and the changelog-shaped one 98%. It measured naming style, not quality, and would have flagged the best repo in the fleet.
What’s new in 1.5.0 — the capture side
Every earlier mechanism was pruning pressure on entries that already exist; the blind spot was a log that was never written to (issue #37 — installed for weeks on a 127-commit repo, zero output, while sessions rediscovered the same gotchas).
New tree-grounded audit states, fleet-calibrated on 33 repos:
-
Unadopted — a hand-rolled gotchas/learnings heading with bullets but zero parseable D&L entries: the repo already tried to solve the problem by hand; the audit offers the migration.
-
Empty log — 20+ commits,
docs/markdown at least 5x CLAUDE.md, and no entries: the learning is going somewhere that doesn’t load every turn. -
Docs recurrence — "the third time…" self-counting language in
docs/; a rule that proves itself in prose belongs in the file that loads every session. -
Layout drift, both directions — a git-tracked top-level directory CLAUDE.md never mentions, and a mentioned
dir/gone from the tree (filtered by git history so foreign paths from skill docs never fire).
New opt-in capture triggers (both default off until real-use noise is
measured): a session-end prompt when a session committed and edited files but
never touched CLAUDE.md, and commit-message mining that offers to promote
gotcha-shaped commit prose while the context is hot. Both route entries:
repo-durable → the D&L log; machine-personal → the learn-on-failure skill.
New structure review: analyze the file’s internal shape against a shipped,
researched rulebook (references/claude-md-best-practices.md — 96 cited
findings from official docs, ~40 real repo files, and measured studies;
adversarially verified). Recommendations must cite a rule or a tree-provable
check, and the measured evidence hierarchy caps cosmetic rearrangement.
What’s new in 1.1.1
-
Merge as the fourth downward pressure — SKILL.md now documents merging same-session same-area clusters (phased rollouts
e2a/e2b/e3/e4; multi-aspect bursts within ~48h) as the compaction lever that works pre-14-days. Graduation and quarterly archive both require old entries; merge does not. Discovered while compacting a sibling repo’s 37-entry log when the audit fired "RECOMMENDED" but every documented pressure was blocked.
What’s new in 1.1.0
-
Whole-file size check — audit now reports total
CLAUDE.mdKB independently (warn 25 KB, recommend 40 KB). Catches bloat that lives outside the# Decisions & Learningssection (Conventions / Architecture / Gotchas) which the entry/line counters never saw. -
Self-report when D&L is missing — a repo without the
# Decisions & Learningssection no longer exits silently; the audit reports the file size + total dated-bullet count and nudges to wireevolving-claude-md. -
Staleness trigger — for each D&L entry, the audit extracts backticked artifact-looking tokens (paths,
ClassName.method,--cli-flags,<xml-tags>, filenames with known extensions) and runsgit grep -qFagainst the tree. Entries citing ≥2 missing tokens surface as "may be wrong now" candidates. Budget-capped (~2.5s) so the hook stays fast. Closes the only signal the bloat-only audit lacked: a delete lever, not just a compress lever.
Tuning per repo
"Concise" is not a universal number — 40 KB is bloat in a library and reasonable in a monorepo. Every threshold is overridable, resolved defaults → ~/.claude/evolving-claude-md/config.json → <repo>/.claude/evolving-claude-md/config.json.
{
"file_warn_kb": 25, "file_recommend_kb": 40,
"lines_warn": 200, "lines_recommend": 300,
"entries_warn": 25, "entries_recommend": 35,
"mega_entry_chars": 800, "topic_cluster": 3,
"layout_min_dirs": 5,
"coverage": true, "nested": true
}
Name only what you want changed. Unknown keys are ignored and a corrupt file falls back to defaults — this runs on SessionStart, so a bad config must never be why a session starts badly.
Companion context files
.claude.local.md and nested CLAUDE.md files load into context exactly like the root file, so bloat in them was previously the same problem measured nowhere. The audit now size-checks both, walking up to three levels deep and skipping node_modules, target, build and friends.
Only the root file gets the full treatment — Decisions & Learnings parsing, staleness, coverage — because that is where the log lives, and reporting on five files at every session start would be its own kind of noise.
The routing rule for which file a learning belongs in is one question: would this still be true on a teammate’s laptop, in CI, and in a fresh clone? No means .claude.local.md. An absolute path containing your username is the most common way a personal detail gets committed — ~/ is fine, /home/alex/… is not.
Coverage — the one upward check
Every other check pushes content down: bloat, staleness, clustering, archiving. A file can pass all of them and still be useless — well under every threshold, perfectly formatted, and never saying how to run the tests.
Two gaps, both grounded in the tree rather than in a generic checklist:
| Gap | Fires only when |
|---|---|
no build/test command |
a build file exists ( |
nothing on layout |
the repo has 5+ meaningful top-level directories — generated ones like |
The grounding is the design. A docs repo has no build command and a three-directory repo needs no layout section; a generic checklist nags both. Measured across a 29-repo sample these fired zero times, because every one already covered what its tree justified.
Decline a gap by putting the decision in the file itself:
<!-- audit-skip: commands, layout -->
A gotchas check was built and then removed: on the same sample it fired on 17 of 29 repos, and those 17 were exactly the ones with no Decisions & Learnings log — so it was re-detecting "hasn’t adopted this skill", which the audit already reports. Gotchas arrive by graduation from the log, and the topic-cluster check is the grounded way to prompt for them.
Notes
-
Manual (non-plugin) install requires copying the three scripts into
.claude/skills/evolving-claude-md/and adding the hooks to.claude/settings.json, then one session restart. -
The quarterly archive is the one operational step nothing automates — set a calendar reminder.
-
Disable a hook by removing its entry; disable all via
disableAllHooks: true.
Claude Code vs Codex
Tier 3 — lifecycle hooks on both clients.
| Claude Code | Codex | |
|---|---|---|
Hooks |
Six events: SessionStart audit, InstructionsLoaded load-recording, PreToolUse format lint, PostToolUse commit mining, Stop capture, PostCompact audit. |
Five of the six, through |
Instructions file |
|
|
Format lint |
Gates |
Gates |
Capture |
Stop and commit detection read the Claude transcript. |
The same capture reads Codex rollout calls; settings and lane spools stay in |
Compaction |
|
No plugin slash command. Ask to compact the instructions file: the same procedure runs on the same scripts ( |
The difference that matters: An instruction file edited from the shell (sed, a script) is not a matched tool path, so the lint never sees it. PostCompact findings arrive as a systemMessage rather than injected context.
Install, hook trust and verified limits: Codex compatibility.