Classifier Assessment: Performance Metrics for Binary Decisions
Binary classifiers assign one of two labels to an observation: a training run has failed or is normal, a document is relevant or irrelevant, a patient has a condition or does not. Assessing these decisions requires several measures because missed cases and false alarms answer different questions. Medicine, information retrieval, and statistics often use different names for the same measures.
Rates such as sensitivity and precision divide one of the four outcome counts by a relevant total; summaries such as $F_1$ and correlation combine the counts differently. The measures fall into four groups:
- Outcome counts. The $2 \times 2$ table records true positives, false negatives, false positives, and true negatives ($\mathrm{TP}, \mathrm{FN}, \mathrm{FP}, \mathrm{TN}$).
- Conditional rates. Conditioning on the actual class gives $P(\hat{Y} \mid Y)$, including sensitivity and specificity. Conditioning on the predicted class gives $P(Y \mid \hat{Y})$, including precision and negative predictive value.
- Summary metrics. Accuracy measures overall agreement. $F_1$ combines precision and recall. Informedness ($\mathrm{BM}$), markedness ($\mathrm{MK}$), and Matthews correlation coefficient ($\mathrm{MCC}$) summarize other relationships between actual and predicted labels.
- Decision context. Prevalence, error costs, and the decision threshold affect how these measures guide a choice of classifier.
1. Actual and Predicted Classes: Counts and Marginals
Each classification pairs an actual label $Y \in {0, 1}$ with a predicted label $\hat{Y} \in {0, 1}$. Dividing the outcome counts by the total number of cases $N_\mathrm{total}$ gives an empirical joint distribution $P(Y=y, \hat{Y}=\hat{y})$. Row and column sums give the marginal counts for actual and predicted classes:
| Reality ($Y$) ↓ \ Decision ($\hat{Y}$) → | Predicted Positive ($\hat{Y} = 1$) | Predicted Negative ($\hat{Y} = 0$) | REALITY MARGINALS Ground Truth Population |
|---|---|---|---|
| Actually Positive ($Y = 1$) | True Positive (TP) Correct Alarm / Hit | False Negative (FN) Missed Case | Total Positives ($P$) $P = \mathrm{TP} + \mathrm{FN}$ |
| Actually Negative ($Y = 0$) | False Positive (FP) False Alarm | True Negative (TN) Correct Rejection | Total Negatives ($N$) $N = \mathrm{TN} + \mathrm{FP}$ |
| DECISION MARGINALS Classifier Outputs | Total Positive Alerts ($\hat{P}$) $\hat{P} = \mathrm{TP} + \mathrm{FP}$ | Total Negative Reports ($\hat{N}$) $\hat{N} = \mathrm{TN} + \mathrm{FN}$ | Grand Total ($N_\mathrm{total}$) $N_\mathrm{total} = P + N = \hat{P} + \hat{N}$ |
The row totals $P$ and $N$ count actual positive and negative cases. The prevalence is the positive fraction, $\pi = P / N_\mathrm{total}$. The column totals count positive and negative predictions; $\hat{P} / N_\mathrm{total}$ is the fraction classified as positive.
When only prevalence changes, sensitivity and specificity stay fixed if the class-conditional distributions of observations and the decision rule stay fixed. The prevalence comparisons below assume this setting. Changes within either class, or changes to the threshold, can change these rates too. A count-based rate is defined only when its denominator is nonzero.
Example: Monitoring LLM Training Runs
Consider a monitor that checks 1,000 LLM training runs for a failure defined by a sharp increase in loss:
- Actual outcomes: 100 runs fail ($P = 100$), while 900 runs remain normal ($N = 900$). The failure prevalence is $\pi = 10\%$.
- Catching Failures: The monitor catches 80 of the 100 failing runs ($\mathrm{TP} = 80$) and misses 20 ($\mathrm{FN} = 20$).
- False alarms: Among the 900 normal runs, the monitor flags 90 ($\mathrm{FP} = 90$) and correctly labels 810 as normal ($\mathrm{TN} = 810$).
| Reality ($Y$) ↓ \ Decision ($\hat{Y}$) → | Flagged as Failure ($\hat{Y} = 1$) | Cleared as Normal ($\hat{Y} = 0$) | REALITY MARGINALS Ground Truth Population |
|---|---|---|---|
| Actually Failed ($Y = 1$) | TP = 80 Caught Failures | FN = 20 Missed Failures | P = 100 10% Prevalence (π) |
| Actually Normal ($Y = 0$) | FP = 90 False Alarms | TN = 810 Correct Negative Predictions | N = 900 90% Negative Cases (1 − π) |
| DECISION MARGINALS Operational Output | P̂ = 170 Total Positive Predictions | N̂ = 830 Total Negative Predictions | $N_\mathrm{total} = 1{,}000$ Total Trials |
Conditioning on a row or a column gives four different questions:
- Actual positives (row 1): Of the runs that failed, what fraction did the monitor detect? $$\mathrm{TPR} = \frac{\mathrm{TP}}{P} = \frac{80}{100} = \mathbf{80.0\%} \quad (\text{Sensitivity / Recall})$$
- Actual negatives (row 2): Of the normal runs, what fraction did the monitor label as normal? $$\mathrm{TNR} = \frac{\mathrm{TN}}{N} = \frac{810}{900} = \mathbf{90.0\%} \quad (\text{Specificity})$$
- Positive predictions (column 1): Of the flagged runs, what fraction actually failed? $$\mathrm{PPV} = \frac{\mathrm{TP}}{\hat{P}} = \frac{80}{80 + 90} = \frac{80}{170} = \mathbf{47.1\%} \quad (\text{Precision})$$
- Negative predictions (column 2): Of the runs labeled normal, what fraction were actually normal? $$\mathrm{NPV} = \frac{\mathrm{TN}}{\hat{N}} = \frac{810}{810 + 20} = \frac{810}{830} = \mathbf{97.6\%} \quad (\text{Negative Predictive Value})$$
The monitor has 80% sensitivity and 90% specificity, but 52.9% of its alerts are false alarms. The 10% false positive rate applies to 900 normal runs, producing 90 false alarms. These outnumber the 80 detected failures, so precision is $80/170$.
Which Cases Are in the Denominator?
Each conditional rate refers to one row or column: Which cases are included in that reference population?
Within each nonempty row or column, a case is either correctly or incorrectly classified. The corresponding rates therefore sum to $1$:
| Conditioning Denominator | Correct Classification Rate | Complementary Error Rate | Sum | Effect of Changing Prevalence Alone |
|---|---|---|---|---|
| Condition Positive ($P = \mathrm{TP} + \mathrm{FN}$) Reality: Actually Failed | Sensitivity / Recall ($\mathrm{TPR}$) $\frac{\mathrm{TP}}{P}$ (Catch Rate) | Miss Rate ($\mathrm{FNR}$) $\frac{\mathrm{FN}}{P}$ (Type II Error) | $\mathrm{TPR} + \mathrm{FNR} = 1.0$ | Prevalence-Invariant |
| Condition Negative ($N = \mathrm{TN} + \mathrm{FP}$) Reality: Actually Normal | Specificity / Selectivity ($\mathrm{TNR}$) $\frac{\mathrm{TN}}{N}$ (Clearance Rate) | Fall-out / False Alarm ($\mathrm{FPR}$) $\frac{\mathrm{FP}}{N}$ (Type I Error) | $\mathrm{TNR} + \mathrm{FPR} = 1.0$ | Prevalence-Invariant |
| Predicted Positive ($\hat{P} = \mathrm{TP} + \mathrm{FP}$) Decision: Sounded Alarm | Precision / PPV $\frac{\mathrm{TP}}{\hat{P}}$ (Correct Alert Fraction) | False Discovery Rate ($\mathrm{FDR}$) $\frac{\mathrm{FP}}{\hat{P}}$ (False Alert Fraction) | $\mathrm{PPV} + \mathrm{FDR} = 1.0$ | Prevalence-Dependent |
| Predicted Negative ($\hat{N} = \mathrm{TN} + \mathrm{FN}$) Decision: Issued All-Clear | Negative Predictive Value ($\mathrm{NPV}$) $\frac{\mathrm{TN}}{\hat{N}}$ (Correct Negative Fraction) | False Omission Rate ($\mathrm{FOR}$) $\frac{\mathrm{FN}}{\hat{N}}$ (Missed Case Fraction Among Negative Predictions) | $\mathrm{NPV} + \mathrm{FOR} = 1.0$ | Prevalence-Dependent |
Fall-out ($\mathrm{FPR}$) is the false alarm rate among actual negatives. Miss rate ($\mathrm{FNR}$) is the missed fraction of actual positives.
Classical Hypothesis Testing and Pattern Recognition
In Neyman–Pearson hypothesis testing, a decision rule either rejects or does not reject a null hypothesis $H_0$. A Type I error is rejection when $H_0$ is true; a Type II error is failure to reject when a specified alternative $H_1$ is true. For two simple hypotheses (each specifying a distribution), the Neyman–Pearson lemma gives a likelihood-ratio test with the greatest power among tests whose Type I error probability is at most $\alpha$. Power is $1-\beta = P(\text{reject } H_0 \mid H_1)$.
These error probabilities correspond to rates in the confusion matrix when $H_0$ denotes the negative class and $H_1$ the positive class. The distinction is between a probability constraint used to design a test and a frequency measured when evaluating it:
Test design: Choose a rejection rule subject to a bound on its Type I error probability.
- Significance level ($\alpha$): Upper bound on the probability of rejection under $H_0$.
- Type II error ($\beta$): Probability of missing a true signal under $H_1$.
- Statistical Power ($1 - \beta$): Probability of correctly rejecting $H_0$ when $H_1$ is true.
Classifier evaluation: Compare predictions with known labels in a sample.
- False Positive Rate ($\mathrm{FPR}$): Measured sample frequency $\mathrm{FP} / N$.
- False Negative Rate ($\mathrm{FNR}$): Measured miss rate $\mathrm{FN} / P$.
- Sensitivity / Recall ($\mathrm{TPR}$): Measured hit rate $\mathrm{TP} / P$.
| Reality ($Y$) ↓ \ Decision ($\hat{Y}$) → | Reject $H_0$ / Predict Positive ($\hat{Y} = 1$) | Do Not Reject $H_0$ / Predict Negative ($\hat{Y} = 0$) | Hypothesis Ground Truth |
|---|---|---|---|
| Alternative Hypothesis True ($H_1$ is true, $Y = 1$) | True Positive (TP) Neyman–Pearson: Statistical Power ($1 - \beta$) Machine Learning: Hit Rate / Recall ($\mathrm{TPR} = \mathrm{TP} / P$) | False Negative (FN) Neyman–Pearson: Type II Error ($\beta$) Machine Learning: Miss Rate ($\mathrm{FNR} = \mathrm{FN} / P$) | $P = \mathrm{TP} + \mathrm{FN}$ Signal Prevalence ($\pi$) |
| Null Hypothesis True ($H_0$ is true, $Y = 0$) | False Positive (FP) Neyman–Pearson: Type I Error ($\alpha$, Significance) Machine Learning: Fall-out / False Alarm ($\mathrm{FPR} = \mathrm{FP} / N$) | True Negative (TN) Neyman–Pearson: Correct Non-Rejection ($1 - \alpha$ for a size-$\alpha$ test) Machine Learning: Specificity ($\mathrm{TNR} = \mathrm{TN} / N$) | $N = \mathrm{TN} + \mathrm{FP}$ Noise Base Rate ($1 - \pi$) |
Accuracy and Error Rate Under Class Imbalance
Accuracy is the fraction of cases whose predicted and actual labels agree. Its numerator is the agreement diagonal ($\mathrm{TP} + \mathrm{TN}$):
In Figure 1, select the Always-Negative Model and set the negative class fraction to $99\%$. Accuracy is $99\%$, but sensitivity is $0\%$: every positive case is missed.
2. Conditioning on the Actual Class
Row-conditioning asks: "Given the actual class, how likely is each prediction?" Each denominator contains only positive cases or only negative cases. These rates stay fixed when prevalence changes alone, under the fixed-distribution and fixed-rule assumptions above.
Condition Positive ($Y = 1$): Sensitivity and Miss Rate
Condition Negative ($Y = 0$): Specificity and Fall-out
3. Conditioning on the Predicted Class
Column-conditioning asks: "Given a positive prediction, what is the probability that the case is actually positive?" The denominators are the prediction totals ($\hat{P}$ and $\hat{N}$). Each total combines actual positives and negatives, so these rates generally change with prevalence.
Sensitivity assesses $P(\hat{Y}=1 \mid Y=1)$, whereas Precision assesses $P(Y=1 \mid \hat{Y}=1)$. Bayes' theorem expresses precision in terms of prevalence, sensitivity, and the false positive rate: $$\mathrm{PPV} = \frac{\pi \cdot \mathrm{TPR}}{\pi \cdot \mathrm{TPR} + (1 - \pi) \cdot \mathrm{FPR}}$$
Predicted Positive ($\hat{Y} = 1$): Precision and False Discovery
Predicted Negative ($\hat{Y} = 0$): Negative Predictive Value and False Omission
In Figure 3, drag the vertical divider to change prevalence $\pi$. The class-conditional heights stay fixed, while the widths of the actual-positive and actual-negative columns change. At low prevalence, the true positive area (teal) becomes small relative to the false positive area (coral). Precision is the teal fraction of their combined area. The two areas need not have the same height.
Comparing Rates with Different Denominators
Several metrics share a numerator but answer different questions. Sensitivity and precision both count true positives; specificity and negative predictive value both count true negatives. Comparing their denominators explains the difference:
Denominator: Actual positive cases ($P = \mathrm{TP} + \mathrm{FN}$).
Operational Question: "Given that an anomaly occurred, what is the probability the detector flags it?"
Limitation: Predicting positive for every case gives 100% recall, but also flags every negative case.
Denominator: Positive predictions ($\hat{P} = \mathrm{TP} + \mathrm{FP}$).
Operational Question: "Given that an alarm sounded, what is the probability that the case truly failed?"
Limitation: At low prevalence, false alarms can outnumber true positives even when sensitivity and specificity are high.
Denominator: Actual negative cases ($N = \mathrm{TN} + \mathrm{FP}$).
Operational Question: "Given that a case is benign, what is the probability the system clears it?"
Limitation: Predicting negative for every case gives 100% specificity, but misses every positive case.
Denominator: Issued negative decisions ($\hat{N} = \mathrm{TN} + \mathrm{FN}$).
Operational Question: "Given a negative prediction, what is the probability that the case is actually negative?"
Limitation: An always-negative classifier has $\mathrm{NPV} = 1-\pi$. Its NPV approaches 100% as prevalence approaches zero, even though it misses every positive case.
Equation: $\mathrm{FPR} = \frac{\mathrm{FP}}{N} = 1 - \mathrm{TNR}$.
Interpretation: The probability that an actual negative case triggers an alert. It stays fixed when only prevalence changes.
Equation: $\mathrm{FDR} = \frac{\mathrm{FP}}{\hat{P}} = 1 - \mathrm{PPV}$.
Interpretation: The fraction of alerts that are false positives. Unlike FPR, its denominator is all alerts, so it generally changes with prevalence.
4. How Metrics Change with Prevalence
Figure 4 varies prevalence from $0.001$ to $0.999$ while keeping sensitivity and specificity at the values selected in Figure 3. Move those controls to compare the resulting curves:
5. Composite Summaries: Harmonic $F_1$ and Geometric $MCC$
$F_1$ and MCC each summarize several rates in one number. The constructions below show a harmonic mean and a geometric mean. The right panel uses illustrative positive values of $BM$ and $MK$ that vary with the controls; they are not derived from a confusion matrix with the selected precision and recall.
Here $P$ denotes precision and $R$ recall, rather than class counts. The harmonic mean is $F_1 = 2 \frac{P \cdot R}{P + R}$. The diagonals joining the tops and opposite bases of bars of heights $P$ and $R$ intersect at height $\frac{1}{2} F_1$. Change either control to compare $F_1$ with the arithmetic mean. The harmonic mean is lower when precision and recall differ. $F_1$ uses true positives, false positives, and false negatives; it does not use true negatives.
Matthews correlation coefficient relates two directional summaries:
- Informedness ($\mathrm{BM} = \mathrm{TPR} + \mathrm{TNR} - 1$): The difference between true positive and false positive rates.
- Markedness ($\mathrm{MK} = \mathrm{PPV} + \mathrm{NPV} - 1$): The difference between the positive-class proportions among positive and negative predictions.
$$\mathrm{MCC} = \operatorname{sign}(\mathrm{TP}\mathrm{TN}-\mathrm{FP}\mathrm{FN})\sqrt{\mathrm{BM} \cdot \mathrm{MK}}$$ For nonnegative $BM$ and $MK$, the perpendicular from their junction to a semicircle of diameter $BM+MK$ has height $\sqrt{BM \cdot MK}$. MCC is the Pearson correlation between binary actual and predicted labels. It is zero when either directional summary is zero, provided all row and column totals are nonzero.
6. Terminology Across Disciplines
Select a discipline to highlight its terminology. In the hypothesis-testing column, $\alpha$ denotes the test's actual Type I error probability; it equals the chosen significance level when the test attains that bound.
| Measure | Conditioning | Medicine / Epidemiology | Radar / Signal Detection | Search / IR | Machine Learning | Hypothesis Testing | Changing Prevalence Alone |
|---|---|---|---|---|---|---|---|
| $\mathrm{TP} / P$ | $P(\hat{Y}=1 \mid Y=1)$ | Sensitivity | Hit Rate | Recall | True Positive Rate (TPR) | Power ($1-\beta$) | Invariant (↔) |
| $\mathrm{FN} / P$ | $P(\hat{Y}=0 \mid Y=1)$ | False Negative Rate | Miss Rate | Miss Rate | False Negative Rate (FNR) | Type II Error ($\beta$) | Invariant (↔) |
| $\mathrm{TN} / N$ | $P(\hat{Y}=0 \mid Y=0)$ | Specificity | Correct Rejection | — | True Negative Rate (TNR) | Correct Non-Rejection ($1-\alpha$) | Invariant (↔) |
| $\mathrm{FP} / N$ | $P(\hat{Y}=1 \mid Y=0)$ | False Positive Rate | False Alarm Rate | Fall-out | False Positive Rate (FPR) | Type I Error / Sig. ($\alpha$) | Invariant (↔) |
| $\mathrm{TP} / \hat{P}$ | $P(Y=1 \mid \hat{Y}=1)$ | PPV | — | Precision | Precision / PPV | — | Varies (↗↘) |
| $\mathrm{TN} / \hat{N}$ | $P(Y=0 \mid \hat{Y}=0)$ | NPV | — | — | NPV | — | Varies (↗↘) |
| $\mathrm{TPR} - \mathrm{FPR}$ | Class-Conditional Difference | Youden's $J$ | Hit Rate − False Alarm Rate | — | Informedness ($BM$) | Power − Size | Invariant (↔) |
| $\mathrm{PPV} + \mathrm{NPV} - 1$ | Prediction-Conditional Difference | — | — | — | Markedness ($MK$) | — | Varies (↗↘) |
| $\operatorname{sign}(\mathrm{BM})\sqrt{BM \cdot MK}$ | Overall Correlation | — | — | — | MCC / Pearson $\phi$ | Contingency $\phi$ | Varies (↗↘) |