Evolving CLAUDE.md

Wires CLAUDE.md to maintain a living Decisions & Learnings log that the assistant updates as the project evolves — without letting the file bloat. CLAUDE.md is read into context every turn, so the goal is a small, well-pruned set of durable decisions, with everything else linked out or archived. Three hooks keep the log healthy automatically. It complements the built-in init skill: init describes the codebase, this adds the mechanism that lets the file grow safely.

Install

/plugin install evolving-claude-md@alexmskills

Trigger it

/evolving-claude-md:evolving-claude-md

This skill is model-invokable — most often you just ask in natural language: "make CLAUDE.md evolve", "add a self-updating decisions log to CLAUDE.md", or "CLAUDE.md is getting too big, compact it".

To compact the log on demand, run the dedicated command:

/evolving-claude-md:compact          # plan, propose, apply on approval
/evolving-claude-md:compact plan     # just show the plan
/evolving-claude-md:compact yes      # apply without waiting

When to use it

  • "add learning mechanism to CLAUDE.md", "make CLAUDE.md evolve", "self-updating CLAUDE.md"

  • "decisions log", "ADR-style notes in CLAUDE.md", "project memory in CLAUDE.md"

  • "compact CLAUDE.md" / "CLAUDE.md is getting too big"

  • A CLAUDE.md exists but has no mechanism to keep itself current

  • Not for plain codebase documentation — use init for that

What it does

Automatic hooks

The plugin ships hooks/hooks.json (pathed via ${CLAUDE_PLUGIN_ROOT}), so the hooks register on enable with nothing to add to settings:

  • SessionStart — audit-claude-md.py checks the D&L section + whole-file size + staleness and injects a recommendation when something needs attention; silent when healthy. Since 1.4.0 staleness covers three rot classes, grounded against the tree rather than reasoned about: a cited artifact git grep can no longer find, a version pin every build file now contradicts ("Spring Boot 4.0.5" vs a pom that says 4.0.7 — prefix-compatible pins and examples stay silent; since 1.9.0 the build files read are Maven, Gradle, npm, Go, Cargo, Python, .NET, Bundler and Composer, plus toolchain pin files such as .nvmrc and .tool-versions), and a sequence fact the tree has moved past ("latest is V27`" once `V28 exists). The predicates live in a vendorable stdlib core, freshness.py, calibrated across 29 real repos to zero false positives.

  • PreToolUse on Write|Edit of CLAUDE.md — lint-claude-md.py rejects new entries that lack a date or topic-tag, or whose body runs past 500 characters (200 is the target; past that, the detail belongs in docs/decisions/).

  • InstructionsLoaded — record-loads.py records which instruction files actually loaded, so a rule whose globs never match, or a CLAUDE.md that never loads, becomes visible.

  • PostCompact — re-runs the audit so the assistant sees the current state without paying to re-read the whole file.

  • PostToolUse and Stop — capture triggers that nominate a possible learning from a commit message or a session’s end. Off by default.

One command

/evolving-claude-md:compact acts on what the audit reports — see What’s new in 1.8.0 below.

What’s new in 1.8.0 — compact as a command

/evolving-claude-md:compact runs the four downward pressures as one reviewed edit, in the order that keeps lessons safe: supersede gaps, release mirrors (keep what a release taught, drop what it shipped), same-day merges, graduation into Conventions/Gotchas, mega-entry splits, stale checks — and the archive last. Archiving first is the mistake an ad-hoc compaction makes: lessons that should be standing rules end up in the quarterly file instead of the always-loaded one.

The compaction plan for this repository’s own CLAUDE.md: 51 entries

The plan for this repository’s own CLAUDE.md — a real run of the planner the command reads.

Every edit is drafted with its exact text and shown before anything is written; plan stops there, yes skips the wait. The data comes from a read-only planner, compact-claude-md.py. Its graduation list finds a tag repeated three or more times — which misses lessons that recur by theme under different tags — so it also lists every entry about to age out, for the session to review first.

Fixed along the way: archive-decisions.py was unsafe to run twice. Rebuilding the section from dated entries only, it dropped the section’s format note, and an entry absorbed any undated line below it — so a second run archived the first run’s teaser and left a new one reporting only its own count. It now parses the section line by line, keeps everything it doesn’t archive in place, and merges into an existing quarter teaser.

What’s new in 1.7.0 — evidence and invalidation

1.6.0 gated what gets written. 1.7.0 borrows the two mechanisms an 80-entry survey of instruction-file upkeep tools showed were missing here — chosen because exactly half of those 80 document no staleness handling at all.

Load evidence. A new InstructionsLoaded hook records which instruction file actually loaded, in which session, and why (session_start, nested_traversal, path_glob_match, include, compact), appending to .claude/evolving-claude-md/loads.jsonl. Every other check in the plugin reasons about what the file says; this one records what the harness did, and that makes two failures checkable which no content audit can see:

  • a .claude/rules/*.md whose paths: globs never match anything that gets read — the rule is well-written and simply never consulted, so its content is not the defect;

  • a CLAUDE.md that never loads at all, because Claude Code loads one up to 4 MiB and silently skips a larger file, and a file in an unread path is skipped just as quietly. The repo looks configured and isn’t.

Both wait for load_min_sessions (default 5) of recorded evidence: "hasn’t loaded yet" and "never loads" are indistinguishable on day one.

Supersede links. An entry may end Supersedes: YYYY-MM-DD topic-tag. The audit then reports a target that is still unstruck — a replaced rule still reading as current — or one that doesn’t exist, catching a typo’d date or tag. The lint exempts the link from the body cap, because charging bookkeeping against the prose budget would price the link out of use. The idea comes from temporal knowledge graphs, which record when a fact stopped being true instead of leaving both versions standing.

Also corrected: the docs claimed in three places that the lint hook enforces a 200-character body cap. It enforces 500, so entries between the two had been passing silently against a documented rule. Both numbers are now stated, with the reason they differ — 200 keeps the log scannable, 500 is where an entry is provably a design doc. Enforcement was deliberately left at 500; tightening it would start rejecting entries in every consuming repo.

What’s new in 1.6.0 — the inlet

1.5.0 gave the log a write path. 1.6.0 asks whether a given write has earned its place, because every mechanism before it was downstream of a write nothing gated: four pruning pressures, zero inlet control.

The bar decides; the six triggers only nominate. Write the candidate as a sentence that is true about the repo tomorrow, then name what a future session does differently knowing it. If the only true sentence is "we did X", there is no entry — the changelog, the tickets and git log already hold it. A delivery earns an entry only when it taught something: a constraint, a trap, a reversal. Most sessions nominate zero, and zero is the right answer to a session that merely shipped. The session-end capture prompt carries the bar, and what NOT to log gains the durable-artifact rule: where a repo keeps a changelog, ADRs or tickets, an entry is earned only by the part that isn’t in them.

New audit check — CHANGELOG MIRROR. Measures that duplication directly: entries naming a version CHANGELOG.md already documents. It found 63% on the repo that motivated the change and stays silent on logs of real learnings.

Parallel lanes spool instead of colliding. Measured on one 2,400-commit repo running worktree lanes: 86 of 257 non-merge CLAUDE.md edits were made on lane branches, and all 50 merge commits touching the file were resolving them. A lane — a linked git worktree, or any non-default branch with "lane_spool": "branch" — now routes its nominations to .claude/evolving-claude-md/incoming/<branch>.md. One file per writer, so the conflict class disappears rather than being resolved. Then capture-triggers.py fold prints every lane’s nominations together with cross-lane near-duplicates flagged, for one writer to curate on the default branch — five lanes on one epic learn the same lesson five times, and only a reader holding all of them can dedup it. An ordinary feature branch is deliberately not a lane: one writer at a time merges cleanly.

Obsolescence findings name the edit. MISSING from the layout and OBSOLETE in the layout now prescribe the change to make — add the line, or strike it — because a flag a reader must translate into an action is a flag that gets skimmed past.

A fifth check was written and cut before shipping: tag-uniqueness looked like a clean symptom until calibration showed the fleet’s healthiest log runs 100% distinct tags and the changelog-shaped one 98%. It measured naming style, not quality, and would have flagged the best repo in the fleet.

What’s new in 1.5.0 — the capture side

Every earlier mechanism was pruning pressure on entries that already exist; the blind spot was a log that was never written to (issue #37 — installed for weeks on a 127-commit repo, zero output, while sessions rediscovered the same gotchas).

New tree-grounded audit states, fleet-calibrated on 33 repos:

  • Unadopted — a hand-rolled gotchas/learnings heading with bullets but zero parseable D&L entries: the repo already tried to solve the problem by hand; the audit offers the migration.

  • Empty log — 20+ commits, docs/ markdown at least 5x CLAUDE.md, and no entries: the learning is going somewhere that doesn’t load every turn.

  • Docs recurrence — "the third time…​" self-counting language in docs/; a rule that proves itself in prose belongs in the file that loads every session.

  • Layout drift, both directions — a git-tracked top-level directory CLAUDE.md never mentions, and a mentioned dir/ gone from the tree (filtered by git history so foreign paths from skill docs never fire).

New opt-in capture triggers (both default off until real-use noise is measured): a session-end prompt when a session committed and edited files but never touched CLAUDE.md, and commit-message mining that offers to promote gotcha-shaped commit prose while the context is hot. Both route entries: repo-durable → the D&L log; machine-personal → the learn-on-failure skill.

New structure review: analyze the file’s internal shape against a shipped, researched rulebook (references/claude-md-best-practices.md — 96 cited findings from official docs, ~40 real repo files, and measured studies; adversarially verified). Recommendations must cite a rule or a tree-provable check, and the measured evidence hierarchy caps cosmetic rearrangement.

What’s new in 1.1.1

  • Merge as the fourth downward pressure — SKILL.md now documents merging same-session same-area clusters (phased rollouts e2a/e2b/e3/e4; multi-aspect bursts within ~48h) as the compaction lever that works pre-14-days. Graduation and quarterly archive both require old entries; merge does not. Discovered while compacting a sibling repo’s 37-entry log when the audit fired "RECOMMENDED" but every documented pressure was blocked.

What’s new in 1.1.0

  • Whole-file size check — audit now reports total CLAUDE.md KB independently (warn 25 KB, recommend 40 KB). Catches bloat that lives outside the # Decisions & Learnings section (Conventions / Architecture / Gotchas) which the entry/line counters never saw.

  • Self-report when D&L is missing — a repo without the # Decisions & Learnings section no longer exits silently; the audit reports the file size + total dated-bullet count and nudges to wire evolving-claude-md.

  • Staleness trigger — for each D&L entry, the audit extracts backticked artifact-looking tokens (paths, ClassName.method, --cli-flags, <xml-tags>, filenames with known extensions) and runs git grep -qF against the tree. Entries citing ≥2 missing tokens surface as "may be wrong now" candidates. Budget-capped (~2.5s) so the hook stays fast. Closes the only signal the bloat-only audit lacked: a delete lever, not just a compress lever.

Entry format contract

Every entry follows - YYYY-MM-DD — topic-tag — what. Why: reason. [see → docs/decisions/…​]. Absolute dates only, a mandatory kebab-case topic tag, and a hard 200-char body cap (bigger detail moves to docs/decisions/).

Downward pressure on size

A Recent (last 14 days) / Historic split keeps the actively-scanned section small. Reversals are struck through, stable patterns graduate into Conventions/Gotchas, and a quarterly archive-decisions.py run moves old entries into docs/decisions/{YYYY-Q}.md.

Tuning per repo

"Concise" is not a universal number — 40 KB is bloat in a library and reasonable in a monorepo. Every threshold is overridable, resolved defaults → ~/.claude/evolving-claude-md/config.json → <repo>/.claude/evolving-claude-md/config.json.

{
  "file_warn_kb": 25,      "file_recommend_kb": 40,
  "lines_warn": 200,       "lines_recommend": 300,
  "entries_warn": 25,      "entries_recommend": 35,
  "mega_entry_chars": 800, "topic_cluster": 3,
  "layout_min_dirs": 5,
  "coverage": true,        "nested": true
}

Name only what you want changed. Unknown keys are ignored and a corrupt file falls back to defaults — this runs on SessionStart, so a bad config must never be why a session starts badly.

Companion context files

.claude.local.md and nested CLAUDE.md files load into context exactly like the root file, so bloat in them was previously the same problem measured nowhere. The audit now size-checks both, walking up to three levels deep and skipping node_modules, target, build and friends.

Only the root file gets the full treatment — Decisions & Learnings parsing, staleness, coverage — because that is where the log lives, and reporting on five files at every session start would be its own kind of noise.

The routing rule for which file a learning belongs in is one question: would this still be true on a teammate’s laptop, in CI, and in a fresh clone? No means .claude.local.md. An absolute path containing your username is the most common way a personal detail gets committed — ~/ is fine, /home/alex/… is not.

Coverage — the one upward check

Every other check pushes content down: bloat, staleness, clustering, archiving. A file can pass all of them and still be useless — well under every threshold, perfectly formatted, and never saying how to run the tests.

Two gaps, both grounded in the tree rather than in a generic checklist:

Gap Fires only when

no build/test command

a build file exists (pom.xml, package.json, Cargo.toml, go.mod, Makefile, … 12 supported) and CLAUDE.md never mentions its command

nothing on layout

the repo has 5+ meaningful top-level directories — generated ones like target/ and node_modules/ don’t count — and CLAUDE.md never says where anything lives

The grounding is the design. A docs repo has no build command and a three-directory repo needs no layout section; a generic checklist nags both. Measured across a 29-repo sample these fired zero times, because every one already covered what its tree justified.

Decline a gap by putting the decision in the file itself:

<!-- audit-skip: commands, layout -->

A gotchas check was built and then removed: on the same sample it fired on 17 of 29 repos, and those 17 were exactly the ones with no Decisions & Learnings log — so it was re-detecting "hasn’t adopted this skill", which the audit already reports. Gotchas arrive by graduation from the log, and the topic-cluster check is the grounded way to prompt for them.

Notes

  • Manual (non-plugin) install requires copying the three scripts into .claude/skills/evolving-claude-md/ and adding the hooks to .claude/settings.json, then one session restart.

  • The quarterly archive is the one operational step nothing automates — set a calendar reminder.

  • Disable a hook by removing its entry; disable all via disableAllHooks: true.

Claude Code vs Codex

Tier 3 — lifecycle hooks on both clients.

Claude Code Codex

Hooks

Six events: SessionStart audit, InstructionsLoaded load-recording, PreToolUse format lint, PostToolUse commit mining, Stop capture, PostCompact audit.

Five of the six, through hooks/codex.json and the shared adapter (trust once with /hooks). InstructionsLoaded is Claude Code-only, so the dead-rule and never-loaded checks have no evidence to read on Codex and stay silent rather than guessing.

Instructions file

CLAUDE.md.

AGENTS.override.md, then AGENTS.md, then an existing CLAUDE.md. Audit prose names whichever is selected.

Format lint

Gates Write/Edit of CLAUDE.md.

Gates apply_patch edits — the adapter splits the patch per file, and one malformed file rejects the whole patch.

Capture

Stop and commit detection read the Claude transcript.

The same capture reads Codex rollout calls; settings and lane spools stay in .claude/evolving-claude-md/.

Compaction

/evolving-claude-md:compact — plan, propose, apply on approval.

No plugin slash command. Ask to compact the instructions file: the same procedure runs on the same scripts (compact-claude-md.py, archive-decisions.py), which both honour SKILL_INSTRUCTIONS_FILE, so they act on AGENTS.md when that is the selected file.

The difference that matters: An instruction file edited from the shell (sed, a script) is not a matched tool path, so the lint never sees it. PostCompact findings arrive as a systemMessage rather than injected context.

Install, hook trust and verified limits: Codex compatibility.