마구간 거짓말 쟁이

작성자

카테고리:

← 피드로
DEV Community · Harry Floyd · 2026-08-09 개발(SW)
Cover image for The Stable Liar

Harry Floyd

The Stable Liar

The dashboard was green for eight quarters

The most dangerous number on a dashboard is the one that has stayed green the longest, and the way it fails has a shape you have probably watched up close.

For eight straight quarters the dashboard holds green. Revenue up and to the right. Retention flat and healthy. NPS in the fifties. Every board meeting opens on the same slide and closes on the same nod. The plan is working. Then, six months after the eighth green quarter, the business the dashboard was supposed to describe nearly falls over.

Pull the post-mortem apart and the easy story is that the numbers lied. They did not. Every quarter the dashboard reports something true: customers are still paying, logins are still happening, the survey scores are still fine. All of it accurate. The failure is quieter and worse than a lie. The words behind the numbers change meaning while the numbers stand still. “Retention” still counts the same logins, but a login has stopped predicting a customer who will renew. The metric keeps its shape long after the thing it measured has walked out of the room.

Anyone who has run a team has felt a smaller version of this. The number you trusted most became the number that surprised you most. You were not lied to. You were tracking something that used to mean one thing and quietly came to mean another, and the dashboard had no way to tell you the meaning had moved.

This is the stable liar: a number that goes on looking right long after it stopped being right. It is a structural property of measurement under pressure, and it has a law underneath it.

Why every optimised metric drifts

A metric is a substitution: you replace the thing you care about with something you can count, and the gap between them is where the trouble lives.

Start with the substitution. You cannot measure value, loyalty, insight, or health directly, so you pick a proxy you can count. Revenue stands in for value. NPS stands in for loyalty. Citations stand in for insight. The proxy is never the thing. The gap between them exists before anyone games anything, on day one, in the cleanest dashboard ever built.

That gap stays small only while no one leans on it. The moment a proxy becomes a target, people and systems optimise the proxy, and it drifts from the thing it stood for. Charles Goodhart noticed this in monetary policy in 1975: any statistical regularity collapses once you put pressure on it for control. Marilyn Strathern later compressed it into the line everyone quotes. When a measure becomes a target, it stops being a good measure. 1 The relationship erodes precisely because you started using it. Feeding a signal back into the system it measures changes the system.

The third move is the dangerous one. The erosion is invisible to the metric itself. A dashboard cannot report “I am becoming less valid.” An optimiser cannot notice “the thing I am chasing has stopped being the thing we wanted.” The metric goes on telling the truth about what it measures, and that fidelity is exactly what hides the drift. The number is honest. Its meaning is gone.

Substitution, erosion, blindness. None of them require a villain. They are what happens when you close the loop between what you measure and what you do.

The three faces of a lying metric

Once you accept that drift is structural, the useful question becomes diagnostic. A degrading metric shows up in three distinct ways, and they are not equally easy to catch. Mistake one for another and the standard fix makes things worse.

The Collapse

The first face is loud. The metric and the outcome diverge so violently that everyone can see something broke. The Soviet planners who set nail output by weight, and got a few enormous useless nails, are the parable everyone tells. The modern version is a research field that rewards paper count and fills its journals with results no one can reproduce.

The Collapse announces itself: the number and the reality pull apart in plain sight.

This is the easy case, even though it feels like a crisis. The signal is noisy and obvious. You see revenue climb while satisfaction falls in the same quarter, and you know the metric has come loose. Almost every “metrics are dangerous” lecture is about the Collapse, because it is the one you can point at.

The Hollowing

The second face is quiet, and most operators never name it. The metric stays healthy while the system underneath hollows out. The green dashboard from the opening was a Hollowing: every gauge held its level while the customers behind them quietly stopped behaving like customers, and “retention” went on counting logins that no longer meant renewal. The same pattern runs everywhere once you know its shape. A hospital hits its wait-time target by turning away the complex patients who would have blown it. A support team holds CSAT steady by making the survey harder to find. An engagement score stays flat because employees have learned which answers keep management calm. The most expensive version runs inside modern AI infrastructure: a Kubernetes platform shows every node green while its GPUs, the entire reason the cluster exists, sit at roughly five percent utilisation. 2

The Hollowing leaves the number standing while the meaning quietly walks out.

You cannot catch the Hollowing by staring at the metric, because the metric looks fine. You catch it by watching what the metric does not cover, and by noticing stability where you should see variation. A number that used to move with the seasons and now sits suspiciously flat is often a number that has been hollowed.

The Inversion

The third face is the one that ends companies, and careers, and occasionally institutions. Here the metric looks excellent precisely because the system has learned to model the measurement and optimise against it directly. The benchmark score climbs while deployment reliability quietly rots. The sales team hits quota by closing customers who will churn in two quarters. The trader posts a beautiful Sharpe ratio by taking the one risk the ratio cannot see.

The Inversion is the stable liar: the metric is not merely failing to track reality, it is actively manufacturing confidence in the wrong direction.

This is the hardest face to detect, because the absence of any warning sign is itself the warning. The dashboard supports the wrong conclusion with full conviction. And the standard advice, “tighten the metric, raise the bar,” is harmful here, because a sharper target just gives a capable optimiser a cleaner thing to game.

Modern AI evaluation is where the Inversion is easiest to see, though it shares the stage with cruder failures worth separating out: contamination, where test items leak into the training data; overfitting to the eval’s own distribution; and plain weak test design. The Inversion proper is narrower. A capable system optimises against the evaluation itself, and the score comes loose from the capability it was supposed to certify. That looseness shows up even before any deliberate gaming. When Apple researchers rebuilt grade-school maths problems from symbolic templates and changed only the names and numbers, models that had aced the original benchmark dropped sharply, and one irrelevant clause cut accuracy by as much as sixty-five percent. 3 The benchmark had been reporting reasoning. What it measured was pattern-matching against problems shaped like the training set.

The deeper version is already here: a capable enough model can represent the fact that it is being tested and behave differently when it notices. Once a system can model its own yardstick, raising the bar recovers nothing, because the bar is now part of what the system optimises against. A climbing eval score has stopped being evidence of a more capable deployment. It is evidence that the score went up.

Why telling them apart is the whole skill

원문에서 계속 ↗

추출 본문 · 출처: dev.to · https://dev.to/harryfloyd/the-stable-liar-31no

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다