Claude Code가 실제로 내 CLAUDE.md를 따르는지 측정했습니다.

작성자

카테고리:

← 피드로
DEV Community · rulereceipt · 2026-09-21 개발(SW)

rulereceipt

I’ve been building a small thing that reads a Claude Code session afterward and checks it against your rules file — did it follow “ask before deleting,” “run the tests before claiming done,” that kind of thing.

Simple idea. The naive version took an afternoon.

Then I ran it against real data and it was a liar.

Not a little. Here’s the number that mattered: I ran 559 real rules files against 5 real session transcripts — about 2,795 reports, ~18,000 individual verdicts. The stat I cared about was “what fraction of reports show the user at least one FAIL,” because that’s what a person actually sees on one run.

15.8% of reports had at least one FAIL.

If one in six runs accuses you of breaking a rule you didn’t break, nobody trusts the tool again. It’s worse than useless — it’s a smoke alarm that goes off when you make toast. So I spent the bulk of the work not on catching more, but on catching wrongly less. It’s now at 2.9%, and the bugs I found getting there are more interesting than the tool.

A comma is not a rule. One rules file listed , as a forbidden pattern. My matcher dutifully flagged every file that contained a comma.

rm fired inside “performance”. My pattern matcher had a word boundary on the end but not the start, so the rule “don’t run rm” matched the letters r-m sitting inside the word performance. Every session that used the word got accused of running a delete command.

The worst one: guessing, and always guessing guilty. When my classifier couldn’t tell which direction a rule pointed — “you must do X” or “you must never do X” — it defaulted to the accusing direction, and I’d justified that as “preserving existing behavior.” It then failed a real session for following a rule that said “update the closest CLAUDE.md,” because the session had updated one. It saw the right thing happen and called it a violation.

The fix that stuck wasn’t a smarter heuristic. It was making the bad case impossible to build: a FAIL verdict can now only be constructed through one function that requires the rule’s direction to be typed as”forbid.” Code that’s checking a “you must do X” rule literally cannot emit an accusation — the build fails if it tries.

The general lesson, if there is one: for anything that judges other work — a linter, a test oracle, an AI-output checker — a false positive costs far more than a miss, because it spends the user’s trust, and trust doesn’t refill.

The tool is RuleReceipt (https://rulereceipt.dev) if you want the shape of it (npx rulereceipt check, runs locally, no account) — but honestly I mostly wanted to write down the toast-alarm bugs, because I suspect anyone who’s built something that grades other things has their own version of the comma.

원문에서 계속 ↗