01

Evaluation

A model metric is not an operating result.

The useful metric is attached to a decision, an error cost, and an accountable owner.

Published: 2026-07-23

Accuracy can rise while the exception queue gets worse. Forecast error can fall while the planning horizon remains unusable. Evaluation starts by naming the action that changes when the output is right or wrong.

A classifier can move from 91 to 94 percent accuracy and leave the operation worse off. If the gain lands on the easy cases and the remaining errors concentrate in the category that triggers a manual investigation, the headline number improved and the work got harder.

So the first question in an evaluation is not which metric. It is which action. Something happens when the output is right, and something different happens when it is wrong. Name both, price both, and the metric usually picks itself.

This is also what decides the threshold. A missed defect that reaches a customer and a false alarm that stops a line are not symmetric costs, and no F1 score knows that. The operations team does.

Ask for the error breakdown by the segments you actually care about rather than the aggregate. If nobody can produce it, the evaluation has not really been done yet.

All notes

Disagree with any of this?

These positions come from projects that went badly before they went well. If your situation contradicts one, we would rather hear it than defend it.

Tell us where we are wrong