Every team I've watched chase a reliability number eventually starts managing the number instead of the system. It's not malice. It's incentive gravity. Once a dashboard decides who gets praised in the all-hands and who gets a stern conversation with their manager, the dashboard stops being a window into production and starts being a thing people learn to pose for.
This is an old observation dressed in new infrastructure. Economists have talked about Goodhart's Law for decades — when a measure becomes a target, it ceases to be a good measure — and nowhere does that bite harder than in SRE organizations, because we've built entire cultures around a handful of numbers: uptime, error budget, MTTR, p99 latency. These numbers are genuinely useful. They're also the first thing a stressed team learns to manage around rather than through.
The Error Budget Was Supposed to End the Argument, Not Start a New One
Error budgets exist to settle the fight between "ship faster" and "stop breaking things" with math instead of politics. In theory, you get a fixed amount of acceptable unreliability per quarter, and when you burn through it, feature work pauses and reliability work takes over. Clean. Rational. Almost elegant.
In practice, I've seen error budgets become a negotiation chip instead of a circuit breaker. A team that's about to blow its budget doesn't necessarily stop shipping — it finds a definitional reason why the breach doesn't count. The outage was a "partial degradation." The affected users were a small percentage, so it rounds to zero. The incident technically started in a dependency's service, so it's not really our budget that absorbs it. None of this is dishonest in the dramatic sense. It's the slow, reasonable-sounding erosion of a metric's meaning by people under deadline pressure, one plausible exception at a time.
The fix isn't a stricter policy. Policies get argued around just as easily as the original number did. The fix is treating every exception request as data about where the metric and the org's incentives have come apart — and asking that question in the open, not resolving it quietly in a hallway conversation before the next planning cycle.
MTTR Rewards the Wrong Kind of Fast
Mean time to resolution is probably the most gamed metric in incident management, because "resolution" is doing a lot of unexamined work in that phrase. A team can hit an excellent MTTR by getting very good at restarting the affected service the moment symptoms appear, without ever touching the actual cause. The alert clears. The ticket closes. The number looks great in the quarterly review. And six weeks later the same failure mode shows up again, because nobody was incentivized to spend the extra two hours finding out why the restart worked.
I'd rather see a team with a worse MTTR and a shrinking recurrence rate than a team with a gorgeous MTTR and the same three failure modes on a two-month loop. Recurrence is the metric that actually tells you whether resolution meant something or just meant quiet.
The Boring Fix: Measure What You Can't Easily Fake
None of this means throw out the metrics. It means notice which ones are cheap to game and which ones are expensive to fake, and weight your trust accordingly. Uptime percentages are easy to shape with definitional footwork. Customer-reported incident volume is harder to fake, because customers don't care about your internal taxonomy of what counts as an outage. Time-to-detection is harder to game than time-to-resolution, because you can't quietly redefine when you noticed something was broken. Recurrence rate, as above, is expensive to fake because faking it requires actually fixing the thing, which is the whole point.
The teams I trust most aren't the ones with the cleanest dashboards. They're the ones who can tell you, unprompted, which of their own metrics they don't fully trust and why. That kind of institutional self-awareness doesn't show up on any scorecard, which is exactly why it's rare, and exactly why it's worth more than another decimal point of uptime.
The next time someone presents a reliability number in a review, the useful question isn't whether it hit target. It's what behavior that number was quietly rewarding all quarter — and whether that behavior is one you'd want more of.
