I’ve been building a small thing that reads a Claude Code session afterward and checks it against your rules file — did it follow “ask before deleting,” “run the tests before claiming done,” that kind of thing.
Simple idea. The naive version took an afternoon.
Then I ran it against real data and it was a liar.
Not a little. Here’s the number that mattered: I ran 559 real rules files against 5 real session transcripts — about 2,795 reports, ~18,000 individual verdicts. The stat I cared about was “what fraction of reports show the user at least one FAIL,” because that’s what a person actually sees on one run.
15.8% of reports had at least one FAIL.
If one in six runs accuses you of breaking a rule you didn’t break, nobody trusts the tool again. It’s worse than useless — it’s a smoke alarm that goes off when you make toast. So I spent the bulk of the work not on catching more, but on catching wrongly less. It’s now at 2.9%, and the bugs I found getting there are more interesting than the tool.
A comma is not a rule. One rules file listed , as a forbidden pattern. My matcher dutifully flagged every file that contained a comma.
rm fired inside “performance”. My pattern matcher had a word boundary on the end but not the start, so the rule “don’t run rm” matched the letters r-m sitting inside the word performance. Every session that used the word got accused of running a delete command.
The worst one: guessing, and always guessing guilty. When my classifier couldn’t tell which direction a rule pointed — “you must do X” or “you must never do X” — it defaulted to the accusing direction, and I’d justified that as “preserving existing behavior.” It then failed a real session for following a rule that said “update the closest CLAUDE.md,” because the session had updated one. It saw the right thing happen and called it a violation.
The fix that stuck wasn’t a smarter heuristic. It was making the bad case impossible to build: a FAIL verdict can now only be constructed through one function that requires the rule’s direction to be typed as”forbid.” Code that’s checking a “you must do X” rule literally cannot emit an accusation — the build fails if it tries.
The general lesson, if there is one: for anything that judges other work — a linter, a test oracle, an AI-output checker — a false positive costs far more than a miss, because it spends the user’s trust, and trust doesn’t refill.
The tool is RuleReceipt (https://rulereceipt.dev) if you want the shape of it (npx rulereceipt check, runs locally, no account) — but honestly I mostly wanted to write down the toast-alarm bugs, because I suspect anyone who’s built something that grades other things has their own version of the comma.