---
url: https://fgelbal.com/the-tyranny-of-metrics/
title: The Tyranny of Metrics
author:
  name: Firat Gelbal
  url: https://fgelbal.com/about
date: '2026-09-13'
category: essay
tags:
- Measurement
- Books
excerpt: My notes on Jerry Z. Muller’s book, and what a Soviet nail factory, Deming’s
  red beads, and NBA statistics say about Goodhart’s law. Instead of being metrics-driven,
  let’s make data-informed judgments.
---

Hello there! In this post I would like to share my notes of the book [The Tyranny of Metrics](https://press.princeton.edu/books/hardcover/9780691174952/the-tyranny-of-metrics) from Jerry Z. Muller. I also share my own take on the topics of the book.

The essence of the book boils down to:

> When a measure becomes a target, it ceases to be a good measure
>
> [Goodhart’s law](https://en.wikipedia.org/wiki/Goodhart%27s_law)

Instead, what I suggest is approaching with intention, exercising judgment, and focusing on outcome metrics.

One of the famous stories illustrating this law in action comes from the Soviet factories. When the factories were given targets on the basis of numbers of nails produced, they produced many tiny nails. Next, when they were given targets on the basis of weight, they produced a few giant nails. Numbers and weight are valuable to measure from the planning perspective. But once these measurements become the targets themselves, they lose their value.

<figure style="margin: 1.5rem 0;">
  <img src="https://fgelbal.com/assets/images/2026-09-13-the-tyranny-of-metrics/soviet-nail-factory-krokodil.jpg" alt="Soviet cartoon of two factory workers looking up at a single giant nail hanging from a crane" style="width:45%; display:block; margin:0 auto;" />
  <figcaption style="text-align:center; font-style:italic; font-size:0.9em; margin-top:0.25rem;">The worker asks: “Who needs this nail?” The factory bureaucrat answers: “This is irrelevant. It’s important that we fulfilled the plan immediately.” Cartoon from the Soviet satirical magazine Krokodil (1954), via <a href="https://skeptics.stackexchange.com/a/22386">Skeptics StackExchange</a>.</figcaption>
</figure>

## A widespread pattern

The author explains that the tyranny of metrics is a widespread pattern in contemporary organizational life that runs across everything from business through medicine through academia. It’s based on several beliefs:

* The first is the emphasis on standardized measurement: the notion that our judgment is unreliable. Experience and talent don’t really matter so much. What matters is measuring performance.
* The second notion is that the best way to motivate people within organizations is by attaching rewards and penalties to their measured performance.
* The third notion is connected with the idea of transparency and accountability: to make professional organizations accountable to the public is by making the standardized measures of their performance public.

Each of these ideas sounds plausible when taken on their own. You measure, reward, and punish. You make the whole process public. But the combination of these often end up producing unintended negative consequences. The book expands on these notions by providing examples from various sectors.

The book is not about the evils of measurement or the evils of rewarding people through remuneration. Measurement is often desirable. But, the author makes strong points against using standardized metrics to replace judgment and experience. In a related lesson, Dr. W. Edwards Deming illustrates this phenomenon with the [Red Bead Experiment](https://deming.org/explore/red-bead-experiment/):

> Even though a willing worker wants to do a good job, their success is directly tied to and limited by the nature of the system they are working within. Real and sustainable improvement on the part of the willing worker is achieved only when management is able to improve the system, starting small and then expanding the scope of the improvement efforts.

I see similar examples around me every day. Basketball analysts introduce statistics to measure [player performance and efficiency](https://en.wikipedia.org/wiki/Player_efficiency_rating) and there is [good criticism](https://thezscore.com/2016/02/17/the-definitive-per-criticism/) against it. It’d be tragicomic if the basketball players changed their gameplay behaviors to look good in these weird statistics.

## Don’t leave judgment out

Measurement is not an alternative to judgment. Measurement requires judgment on

* whether to measure
* what to measure
* how to evaluate what’s been measured
* whether rewards and penalties are attached to the results
* to whom to make the measurements available

## Defining good metrics

The book also provides advice on when and how to use metrics. The top one for me is to focus on outcome metrics. “Outcome metrics track the actual and perceived effect of our actions on the population.”

It’s not always possible to measure outcome directly. You might end up defining only output metrics. In this case, the key is to focus on the leading metric to influence the lagging metrics. The following illustration captures the idea:

<figure style="margin: 1.5rem 0;">
  <img src="https://fgelbal.com/assets/images/2026-09-13-the-tyranny-of-metrics/leading-lagging-metrics-sketchplanations.png" alt="Sketch contrasting leading metrics (diet in calories per day, exercise in workouts per week) with the lagging metric they influence (weight loss)" style="width:70%; display:block; margin:0 auto;" />
  <figcaption style="text-align:center; font-style:italic; font-size:0.9em; margin-top:0.25rem;">Leading and lagging metrics, by <a href="https://sketchplanations.com/leading-lagging-metrics">Sketchplanations</a>.</figcaption>
</figure>

The important parts are picking the metrics in your control, and then observing the impact of adjusting these leading metrics.

Next, it’s essential to be aware if you are measuring a proxy instead of what you really want to know. If the information is not very useful or not a good proxy for what you’re really aiming at, you’re probably better off not measuring it.

The book and the above mentioned studies advocate for approaching with intention and exercising clarity.

## Input, output, and outcome

When I first shared these notes, [Beau Lebens](https://beau.blog/), who was reading [Working Backwards](https://www.amazon.com/Working-Backwards-Insights-Stories-Secrets-ebook/dp/B08BYCQBZN/), pointed out that Amazon specifically does not focus on outcome metrics. Instead, it relentlessly focuses on input metrics. One of their examples is the range of book titles available in a category as a key driver of overall sales.

Amazon’s approach is quite similar to what The Tyranny of Metrics and I are advocating here. Input metrics are the controllable ones, known as leading indicators, whereas output metrics are the lagging indicators. The authors of Working Backwards explain how to identify controllable input metrics, like the number of new detail pages or page views, in order to influence the output metric: sales volume.

It’s important to highlight the difference between an output metric and an outcome metric. One could argue that the outcome metric in this case is profitability. At the end, the desired outcome of a business is to make a profit and increase it. Precise calculation of the profit requires the knowledge of additional factors like costs, which go beyond measuring sales. The authors mention stock management and shipping priority as they define the input metrics. I’m sure there are other costs like human resources, servers, and the opportunity cost of prioritization. There are also qualitative factors like the strategy of the company in play here: “How does this play with our ad and marketing direction?” “What’s the impression of customers on our category expansion?”

So the outcome metric calculation is more involved, and it requires a holistic view of the system. It gets challenging to map all the contributing factors to a specific product feature. I like Amazon’s methodology of continuously iterating on the identification and definition of controllable input metrics.

## Data-informed judgment

I also like metrics that let people measure their own relative change. I find it similar to how doctors ask patients to rate their pain (say, from 0 for no pain to 10). Doctors don’t criticize the patient who picks a very small or large number. It’s all relative, and the aim is to compare the difference measured by the patient before and after treatment.

What I take away from this book is: instead of being metrics-driven, let’s focus on making data-informed judgments.
