Oliver's Notes

Classifier Assessment: Performance Metrics for Binary Decisions

How the confusion matrix connects sensitivity, precision, and other measures of binary classification.
AI-Generated AI-Drafted Human-Authored

Binary classifiers assign one of two labels to an observation: a training run has failed or is normal, a document is relevant or irrelevant, a patient has a condition or does not. Assessing these decisions requires several measures because missed cases and false alarms answer different questions. Medicine, information retrieval, and statistics often use different names for the same measures.

Rates such as sensitivity and precision divide one of the four outcome counts by a relevant total; summaries such as $F_1$ and correlation combine the counts differently. The measures fall into four groups:

  1. Outcome counts. The $2 \times 2$ table records true positives, false negatives, false positives, and true negatives ($\mathrm{TP}, \mathrm{FN}, \mathrm{FP}, \mathrm{TN}$).
  2. Conditional rates. Conditioning on the actual class gives $P(\hat{Y} \mid Y)$, including sensitivity and specificity. Conditioning on the predicted class gives $P(Y \mid \hat{Y})$, including precision and negative predictive value.
  3. Summary metrics. Accuracy measures overall agreement. $F_1$ combines precision and recall. Informedness ($\mathrm{BM}$), markedness ($\mathrm{MK}$), and Matthews correlation coefficient ($\mathrm{MCC}$) summarize other relationships between actual and predicted labels.
  4. Decision context. Prevalence, error costs, and the decision threshold affect how these measures guide a choice of classifier.

1. Actual and Predicted Classes: Counts and Marginals

Each classification pairs an actual label $Y \in {0, 1}$ with a predicted label $\hat{Y} \in {0, 1}$. Dividing the outcome counts by the total number of cases $N_\mathrm{total}$ gives an empirical joint distribution $P(Y=y, \hat{Y}=\hat{y})$. Row and column sums give the marginal counts for actual and predicted classes:

The 2×2 Contingency Matrix and Marginals
Reality ($Y$) ↓ \ Decision ($\hat{Y}$) → Predicted Positive
($\hat{Y} = 1$)
Predicted Negative
($\hat{Y} = 0$)
REALITY MARGINALS
Ground Truth Population
Actually Positive
($Y = 1$)
True Positive (TP) Correct Alarm / Hit False Negative (FN) Missed Case Total Positives ($P$) $P = \mathrm{TP} + \mathrm{FN}$
Actually Negative
($Y = 0$)
False Positive (FP) False Alarm True Negative (TN) Correct Rejection Total Negatives ($N$) $N = \mathrm{TN} + \mathrm{FP}$
DECISION MARGINALS
Classifier Outputs
Total Positive Alerts ($\hat{P}$) $\hat{P} = \mathrm{TP} + \mathrm{FP}$ Total Negative Reports ($\hat{N}$) $\hat{N} = \mathrm{TN} + \mathrm{FN}$ Grand Total ($N_\mathrm{total}$) $N_\mathrm{total} = P + N = \hat{P} + \hat{N}$

The row totals $P$ and $N$ count actual positive and negative cases. The prevalence is the positive fraction, $\pi = P / N_\mathrm{total}$. The column totals count positive and negative predictions; $\hat{P} / N_\mathrm{total}$ is the fraction classified as positive.

When only prevalence changes, sensitivity and specificity stay fixed if the class-conditional distributions of observations and the decision rule stay fixed. The prevalence comparisons below assume this setting. Changes within either class, or changes to the threshold, can change these rates too. A count-based rate is defined only when its denominator is nonzero.

Example: Monitoring LLM Training Runs

Consider a monitor that checks 1,000 LLM training runs for a failure defined by a sharp increase in loss:

Example: 1,000 Training Runs
Reality ($Y$) ↓ \ Decision ($\hat{Y}$) → Flagged as Failure
($\hat{Y} = 1$)
Cleared as Normal
($\hat{Y} = 0$)
REALITY MARGINALS
Ground Truth Population
Actually Failed
($Y = 1$)
TP = 80 Caught Failures FN = 20 Missed Failures P = 100 10% Prevalence (π)
Actually Normal
($Y = 0$)
FP = 90 False Alarms TN = 810 Correct Negative Predictions N = 900 90% Negative Cases (1 − π)
DECISION MARGINALS
Operational Output
P̂ = 170 Total Positive Predictions N̂ = 830 Total Negative Predictions $N_\mathrm{total} = 1{,}000$ Total Trials

Conditioning on a row or a column gives four different questions:

  1. Actual positives (row 1): Of the runs that failed, what fraction did the monitor detect? $$\mathrm{TPR} = \frac{\mathrm{TP}}{P} = \frac{80}{100} = \mathbf{80.0\%} \quad (\text{Sensitivity / Recall})$$
  2. Actual negatives (row 2): Of the normal runs, what fraction did the monitor label as normal? $$\mathrm{TNR} = \frac{\mathrm{TN}}{N} = \frac{810}{900} = \mathbf{90.0\%} \quad (\text{Specificity})$$
  3. Positive predictions (column 1): Of the flagged runs, what fraction actually failed? $$\mathrm{PPV} = \frac{\mathrm{TP}}{\hat{P}} = \frac{80}{80 + 90} = \frac{80}{170} = \mathbf{47.1\%} \quad (\text{Precision})$$
  4. Negative predictions (column 2): Of the runs labeled normal, what fraction were actually normal? $$\mathrm{NPV} = \frac{\mathrm{TN}}{\hat{N}} = \frac{810}{810 + 20} = \frac{810}{830} = \mathbf{97.6\%} \quad (\text{Negative Predictive Value})$$

The monitor has 80% sensitivity and 90% specificity, but 52.9% of its alerts are false alarms. The 10% false positive rate applies to 900 normal runs, producing 90 false alarms. These outnumber the 80 detected failures, so precision is $80/170$.

Which Cases Are in the Denominator?

Each conditional rate refers to one row or column: Which cases are included in that reference population?

Within each nonempty row or column, a case is either correctly or incorrectly classified. The corresponding rates therefore sum to $1$:

Conditional Rates and Their Complements
Conditioning Denominator Correct Classification Rate Complementary Error Rate Sum Effect of Changing Prevalence Alone
Condition Positive ($P = \mathrm{TP} + \mathrm{FN}$)
Reality: Actually Failed
Sensitivity / Recall ($\mathrm{TPR}$)
$\frac{\mathrm{TP}}{P}$ (Catch Rate)
Miss Rate ($\mathrm{FNR}$)
$\frac{\mathrm{FN}}{P}$ (Type II Error)
$\mathrm{TPR} + \mathrm{FNR} = 1.0$ Prevalence-Invariant
Condition Negative ($N = \mathrm{TN} + \mathrm{FP}$)
Reality: Actually Normal
Specificity / Selectivity ($\mathrm{TNR}$)
$\frac{\mathrm{TN}}{N}$ (Clearance Rate)
Fall-out / False Alarm ($\mathrm{FPR}$)
$\frac{\mathrm{FP}}{N}$ (Type I Error)
$\mathrm{TNR} + \mathrm{FPR} = 1.0$ Prevalence-Invariant
Predicted Positive ($\hat{P} = \mathrm{TP} + \mathrm{FP}$)
Decision: Sounded Alarm
Precision / PPV
$\frac{\mathrm{TP}}{\hat{P}}$ (Correct Alert Fraction)
False Discovery Rate ($\mathrm{FDR}$)
$\frac{\mathrm{FP}}{\hat{P}}$ (False Alert Fraction)
$\mathrm{PPV} + \mathrm{FDR} = 1.0$ Prevalence-Dependent
Predicted Negative ($\hat{N} = \mathrm{TN} + \mathrm{FN}$)
Decision: Issued All-Clear
Negative Predictive Value ($\mathrm{NPV}$)
$\frac{\mathrm{TN}}{\hat{N}}$ (Correct Negative Fraction)
False Omission Rate ($\mathrm{FOR}$)
$\frac{\mathrm{FN}}{\hat{N}}$ (Missed Case Fraction Among Negative Predictions)
$\mathrm{NPV} + \mathrm{FOR} = 1.0$ Prevalence-Dependent

Fall-out ($\mathrm{FPR}$) is the false alarm rate among actual negatives. Miss rate ($\mathrm{FNR}$) is the missed fraction of actual positives.

Classical Hypothesis Testing and Pattern Recognition

In Neyman–Pearson hypothesis testing, a decision rule either rejects or does not reject a null hypothesis $H_0$. A Type I error is rejection when $H_0$ is true; a Type II error is failure to reject when a specified alternative $H_1$ is true. For two simple hypotheses (each specifying a distribution), the Neyman–Pearson lemma gives a likelihood-ratio test with the greatest power among tests whose Type I error probability is at most $\alpha$. Power is $1-\beta = P(\text{reject } H_0 \mid H_1)$.

These error probabilities correspond to rates in the confusion matrix when $H_0$ denotes the negative class and $H_1$ the positive class. The distinction is between a probability constraint used to design a test and a frequency measured when evaluating it:

Neyman–Pearson Hypothesis Testing

Test design: Choose a rejection rule subject to a bound on its Type I error probability.

  • Significance level ($\alpha$): Upper bound on the probability of rejection under $H_0$.
  • Type II error ($\beta$): Probability of missing a true signal under $H_1$.
  • Statistical Power ($1 - \beta$): Probability of correctly rejecting $H_0$ when $H_1$ is true.
Pattern Recognition & Machine Learning

Classifier evaluation: Compare predictions with known labels in a sample.

  • False Positive Rate ($\mathrm{FPR}$): Measured sample frequency $\mathrm{FP} / N$.
  • False Negative Rate ($\mathrm{FNR}$): Measured miss rate $\mathrm{FN} / P$.
  • Sensitivity / Recall ($\mathrm{TPR}$): Measured hit rate $\mathrm{TP} / P$.
Illustration 2A · Hypothesis Tests and Classification Outcomes
Reality ($Y$) ↓ \ Decision ($\hat{Y}$) →
Reject $H_0$ / Predict Positive
($\hat{Y} = 1$)
Do Not Reject $H_0$ / Predict Negative
($\hat{Y} = 0$)
Hypothesis Ground Truth
Alternative Hypothesis True
($H_1$ is true, $Y = 1$)
True Positive (TP)
Neyman–Pearson: Statistical Power ($1 - \beta$) Machine Learning: Hit Rate / Recall ($\mathrm{TPR} = \mathrm{TP} / P$)
False Negative (FN)
Neyman–Pearson: Type II Error ($\beta$) Machine Learning: Miss Rate ($\mathrm{FNR} = \mathrm{FN} / P$)
$P = \mathrm{TP} + \mathrm{FN}$
Signal Prevalence ($\pi$)
Null Hypothesis True
($H_0$ is true, $Y = 0$)
False Positive (FP)
Neyman–Pearson: Type I Error ($\alpha$, Significance) Machine Learning: Fall-out / False Alarm ($\mathrm{FPR} = \mathrm{FP} / N$)
True Negative (TN)
Neyman–Pearson: Correct Non-Rejection ($1 - \alpha$ for a size-$\alpha$ test) Machine Learning: Specificity ($\mathrm{TNR} = \mathrm{TN} / N$)
$N = \mathrm{TN} + \mathrm{FP}$
Noise Base Rate ($1 - \pi$)
Illustration 2B · Signal Detection Curves & The Neyman–Pearson Lemma

Accuracy and Error Rate Under Class Imbalance

Accuracy is the fraction of cases whose predicted and actual labels agree. Its numerator is the agreement diagonal ($\mathrm{TP} + \mathrm{TN}$):

Accuracy ($\mathrm{ACC}$)
Prevalence-Dependent Overall Agreement
Question What fraction of all decisions were correct?
Equation $$\mathrm{ACC} = \frac{\mathrm{TP} + \mathrm{TN}}{N_\mathrm{total}}$$
Denominator Total population $N_\mathrm{total} = P + N$ (all four cells).
Context Weighted average $\mathrm{ACC} = \pi \cdot \mathrm{TPR} + (1 - \pi) \cdot \mathrm{TNR}$. Can be high even when no minority-class cases are detected.
Complement Error Rate ($\mathrm{ERR} = 1 - \mathrm{ACC}$).
Error Rate ($\mathrm{ERR}$)
Prevalence-Dependent Misclassification
Question What fraction of all decisions were wrong?
Equation $$\mathrm{ERR} = \frac{\mathrm{FP} + \mathrm{FN}}{N_\mathrm{total}} = 1 - \mathrm{ACC}$$
Denominator Total population $N_\mathrm{total} = P + N$ (all four cells).
Limitation Counts false positives and false negatives equally, regardless of their consequences.
Complement Accuracy ($\mathrm{ACC} = 1 - \mathrm{ERR}$).
Figure 1 · Accuracy Under Class Imbalance

In Figure 1, select the Always-Negative Model and set the negative class fraction to $99\%$. Accuracy is $99\%$, but sensitivity is $0\%$: every positive case is missed.

2. Conditioning on the Actual Class

Row-conditioning asks: "Given the actual class, how likely is each prediction?" Each denominator contains only positive cases or only negative cases. These rates stay fixed when prevalence changes alone, under the fixed-distribution and fixed-rule assumptions above.

Condition Positive ($Y = 1$): Sensitivity and Miss Rate

Sensitivity ($\mathrm{TPR}$)
Prevalence-Invariant Recall / Hit Rate / Power
Question Given an actual positive case, what is the probability the classifier detects it? $P(\hat{Y}=1 \mid Y=1)$
Equation $$\mathrm{TPR} = \frac{\mathrm{TP}}{P} = \frac{\mathrm{TP}}{\mathrm{TP} + \mathrm{FN}}$$
Denominator Condition Positive ($P$, top row).
Context Measures the fraction of positive cases detected. Sensitivity below 100% leaves some positive cases undetected.
Complement Miss Rate ($\mathrm{FNR} = 1 - \mathrm{TPR}$).
Miss Rate ($\mathrm{FNR}$)
Prevalence-Invariant Type II Error ($\beta$)
Question Given an actual positive case, what is the probability the classifier misses it? $P(\hat{Y}=0 \mid Y=1)$
Equation $$\mathrm{FNR} = \frac{\mathrm{FN}}{P} = \frac{\mathrm{FN}}{\mathrm{TP} + \mathrm{FN}}$$
Denominator Condition Positive ($P$, top row).
Context Measures missed positive cases, which may carry a high cost in screening or fault detection.
Complement Sensitivity ($\mathrm{TPR} = 1 - \mathrm{FNR}$).

Condition Negative ($Y = 0$): Specificity and Fall-out

Specificity ($\mathrm{TNR}$)
Prevalence-Invariant Selectivity / True Negative Rate
Question Given an actual negative case, what is the probability the classifier correctly rejects it? $P(\hat{Y}=0 \mid Y=0)$
Equation $$\mathrm{TNR} = \frac{\mathrm{TN}}{N} = \frac{\mathrm{TN}}{\mathrm{TN} + \mathrm{FP}}$$
Denominator Condition Negative ($N$, bottom row).
Context Measures correct negative predictions. Higher specificity reduces unnecessary follow-up among actual negatives.
Complement Fall-out / False Alarm Rate ($\mathrm{FPR} = 1 - \mathrm{TNR}$).
Fall-out ($\mathrm{FPR}$)
Prevalence-Invariant Type I Error ($\alpha$) / False Alarm Rate
Question Given an actual negative case, what is the probability of a false alarm? $P(\hat{Y}=1 \mid Y=0)$
Equation $$\mathrm{FPR} = \frac{\mathrm{FP}}{N} = \frac{\mathrm{FP}}{\mathrm{TN} + \mathrm{FP}}$$
Denominator Condition Negative ($N$, bottom row).
Context Measures false alarms per actual negative case. Forms the horizontal axis of ROC curves.
Complement Specificity ($\mathrm{TNR} = 1 - \mathrm{FPR}$).
Figure 2 · Class-Conditional Rates and Informedness

3. Conditioning on the Predicted Class

Column-conditioning asks: "Given a positive prediction, what is the probability that the case is actually positive?" The denominators are the prediction totals ($\hat{P}$ and $\hat{N}$). Each total combines actual positives and negatives, so these rates generally change with prevalence.

Sensitivity assesses $P(\hat{Y}=1 \mid Y=1)$, whereas Precision assesses $P(Y=1 \mid \hat{Y}=1)$. Bayes' theorem expresses precision in terms of prevalence, sensitivity, and the false positive rate: $$\mathrm{PPV} = \frac{\pi \cdot \mathrm{TPR}}{\pi \cdot \mathrm{TPR} + (1 - \pi) \cdot \mathrm{FPR}}$$

Predicted Positive ($\hat{Y} = 1$): Precision and False Discovery

Precision ($\mathrm{PPV}$)
Prevalence-Dependent Positive Predictive Value
Question Given that an alarm sounded, what is the probability the case is truly positive? $P(Y=1 \mid \hat{Y}=1)$
Equation $$\mathrm{PPV} = \frac{\mathrm{TP}}{\hat{P}} = \frac{\mathrm{TP}}{\mathrm{TP} + \mathrm{FP}}$$
Denominator Predicted Positives ($\hat{P}$, left column).
Context Measures the fraction of alerts that identify true cases. For a fixed nonzero false positive rate, it approaches zero as prevalence approaches zero.
Complement False Discovery Rate ($\mathrm{FDR} = 1 - \mathrm{PPV}$).
False Discovery ($\mathrm{FDR}$)
Prevalence-Dependent False Positive Fraction Among Alerts
Question Given that an alarm sounded, what is the probability it was a false alarm? $P(Y=0 \mid \hat{Y}=1)$
Equation $$\mathrm{FDR} = \frac{\mathrm{FP}}{\hat{P}} = \frac{\mathrm{FP}}{\mathrm{TP} + \mathrm{FP}}$$
Denominator Predicted Positives ($\hat{P}$, left column).
Context Measures false positives among positive predictions. In multiple hypothesis testing, FDR instead denotes the expected false discovery proportion, with a convention for zero discoveries.
Complement Precision ($\mathrm{PPV} = 1 - \mathrm{FDR}$).

Predicted Negative ($\hat{Y} = 0$): Negative Predictive Value and False Omission

Negative Predictive Value ($\mathrm{NPV}$)
Prevalence-Dependent Correct Negative Prediction Fraction
Question Given that a negative report was issued, what is the probability the case is truly negative? $P(Y=0 \mid \hat{Y}=0)$
Equation $$\mathrm{NPV} = \frac{\mathrm{TN}}{\hat{N}} = \frac{\mathrm{TN}}{\mathrm{TN} + \mathrm{FN}}$$
Denominator Predicted Negatives ($\hat{N}$, right column).
Context Measures the fraction of negative predictions that are correct. With fixed positive specificity, it approaches 100% as prevalence approaches zero.
Complement False Omission Rate ($\mathrm{FOR} = 1 - \mathrm{NPV}$).
False Omission ($\mathrm{FOR}$)
Prevalence-Dependent Positive Fraction Among Negative Predictions
Question Given that a negative report was issued, what is the probability a true case was missed? $P(Y=1 \mid \hat{Y}=0)$
Equation $$\mathrm{FOR} = \frac{\mathrm{FN}}{\hat{N}} = \frac{\mathrm{FN}}{\mathrm{TN} + \mathrm{FN}}$$
Denominator Predicted Negatives ($\hat{N}$, right column).
Context Estimates the probability of the condition among cases with negative results, a relevant quantity when using a test to rule out a condition.
Complement Negative Predictive Value ($\mathrm{NPV} = 1 - \mathrm{FOR}$).
Figure 3 · The Proportional Unit Square (Mosaic Plot / Eikosogram)
True Positive (TP) False Positive (FP) False Negative (FN) True Negative (TN) Prevalence Divider (π)

In Figure 3, drag the vertical divider to change prevalence $\pi$. The class-conditional heights stay fixed, while the widths of the actual-positive and actual-negative columns change. At low prevalence, the true positive area (teal) becomes small relative to the false positive area (coral). Precision is the teal fraction of their combined area. The two areas need not have the same height.

Comparing Rates with Different Denominators

Several metrics share a numerator but answer different questions. Sensitivity and precision both count true positives; specificity and negative predictive value both count true negatives. Comparing their denominators explains the difference:

Recall / Sensitivity (TPR) · Conditioned on Reality

Denominator: Actual positive cases ($P = \mathrm{TP} + \mathrm{FN}$).

Operational Question: "Given that an anomaly occurred, what is the probability the detector flags it?"

Limitation: Predicting positive for every case gives 100% recall, but also flags every negative case.

Precision / PPV · Conditioned on Decisions

Denominator: Positive predictions ($\hat{P} = \mathrm{TP} + \mathrm{FP}$).

Operational Question: "Given that an alarm sounded, what is the probability that the case truly failed?"

Limitation: At low prevalence, false alarms can outnumber true positives even when sensitivity and specificity are high.

Specificity / Selectivity (TNR) · Conditioned on Reality

Denominator: Actual negative cases ($N = \mathrm{TN} + \mathrm{FP}$).

Operational Question: "Given that a case is benign, what is the probability the system clears it?"

Limitation: Predicting negative for every case gives 100% specificity, but misses every positive case.

Negative Predictive Value (NPV) · Conditioned on Decisions

Denominator: Issued negative decisions ($\hat{N} = \mathrm{TN} + \mathrm{FN}$).

Operational Question: "Given a negative prediction, what is the probability that the case is actually negative?"

Limitation: An always-negative classifier has $\mathrm{NPV} = 1-\pi$. Its NPV approaches 100% as prevalence approaches zero, even though it misses every positive case.

Fall-out / FPR · False Alarms Among Actual Negatives

Equation: $\mathrm{FPR} = \frac{\mathrm{FP}}{N} = 1 - \mathrm{TNR}$.

Interpretation: The probability that an actual negative case triggers an alert. It stays fixed when only prevalence changes.

False Discovery Rate / FDR · False Alarms Among Positive Predictions

Equation: $\mathrm{FDR} = \frac{\mathrm{FP}}{\hat{P}} = 1 - \mathrm{PPV}$.

Interpretation: The fraction of alerts that are false positives. Unlike FPR, its denominator is all alerts, so it generally changes with prevalence.

4. How Metrics Change with Prevalence

Figure 4 varies prevalence from $0.001$ to $0.999$ while keeping sensitivity and specificity at the values selected in Figure 3. Move those controls to compare the resulting curves:

Figure 4 · Metrics Across Prevalence
Sensitivity (TPR) (Invariant) Specificity (TNR) (Invariant) Precision (PPV) (Dependent) NPV (Dependent) Accuracy (Linear Mix) F₁-Score (Dependent) MCC (Dependent)

5. Composite Summaries: Harmonic $F_1$ and Geometric $MCC$

$F_1$ and MCC each summarize several rates in one number. The constructions below show a harmonic mean and a geometric mean. The right panel uses illustrative positive values of $BM$ and $MK$ that vary with the controls; they are not derived from a confusion matrix with the selected precision and recall.

Figure 5 · Composite Summaries: F₁ Crossing Diagonals & MCC Semicircle
The F₁ Harmonic Mean (Left Panel)

Here $P$ denotes precision and $R$ recall, rather than class counts. The harmonic mean is $F_1 = 2 \frac{P \cdot R}{P + R}$. The diagonals joining the tops and opposite bases of bars of heights $P$ and $R$ intersect at height $\frac{1}{2} F_1$. Change either control to compare $F_1$ with the arithmetic mean. The harmonic mean is lower when precision and recall differ. $F_1$ uses true positives, false positives, and false negatives; it does not use true negatives.

MCC and the Geometric Mean (Right Panel)

Matthews correlation coefficient relates two directional summaries:

  • Informedness ($\mathrm{BM} = \mathrm{TPR} + \mathrm{TNR} - 1$): The difference between true positive and false positive rates.
  • Markedness ($\mathrm{MK} = \mathrm{PPV} + \mathrm{NPV} - 1$): The difference between the positive-class proportions among positive and negative predictions.

$$\mathrm{MCC} = \operatorname{sign}(\mathrm{TP}\mathrm{TN}-\mathrm{FP}\mathrm{FN})\sqrt{\mathrm{BM} \cdot \mathrm{MK}}$$ For nonnegative $BM$ and $MK$, the perpendicular from their junction to a semicircle of diameter $BM+MK$ has height $\sqrt{BM \cdot MK}$. MCC is the Pearson correlation between binary actual and predicted labels. It is zero when either directional summary is zero, provided all row and column totals are nonzero.

6. Terminology Across Disciplines

Select a discipline to highlight its terminology. In the hypothesis-testing column, $\alpha$ denotes the test's actual Type I error probability; it equals the chosen significance level when the test attains that bound.

Showing terminology from all five fields. Hover over a row to highlight the corresponding measure.
The Terminology Matrix Across Disciplines
Measure Conditioning Medicine / Epidemiology Radar / Signal Detection Search / IR Machine Learning Hypothesis Testing Changing Prevalence Alone
$\mathrm{TP} / P$ $P(\hat{Y}=1 \mid Y=1)$ Sensitivity Hit Rate Recall True Positive Rate (TPR) Power ($1-\beta$) Invariant (↔)
$\mathrm{FN} / P$ $P(\hat{Y}=0 \mid Y=1)$ False Negative Rate Miss Rate Miss Rate False Negative Rate (FNR) Type II Error ($\beta$) Invariant (↔)
$\mathrm{TN} / N$ $P(\hat{Y}=0 \mid Y=0)$ Specificity Correct Rejection — True Negative Rate (TNR) Correct Non-Rejection ($1-\alpha$) Invariant (↔)
$\mathrm{FP} / N$ $P(\hat{Y}=1 \mid Y=0)$ False Positive Rate False Alarm Rate Fall-out False Positive Rate (FPR) Type I Error / Sig. ($\alpha$) Invariant (↔)
$\mathrm{TP} / \hat{P}$ $P(Y=1 \mid \hat{Y}=1)$ PPV — Precision Precision / PPV — Varies (↗↘)
$\mathrm{TN} / \hat{N}$ $P(Y=0 \mid \hat{Y}=0)$ NPV — — NPV — Varies (↗↘)
$\mathrm{TPR} - \mathrm{FPR}$ Class-Conditional Difference Youden's $J$ Hit Rate − False Alarm Rate — Informedness ($BM$) Power − Size Invariant (↔)
$\mathrm{PPV} + \mathrm{NPV} - 1$ Prediction-Conditional Difference — — — Markedness ($MK$) — Varies (↗↘)
$\operatorname{sign}(\mathrm{BM})\sqrt{BM \cdot MK}$ Overall Correlation — — — MCC / Pearson $\phi$ Contingency $\phi$ Varies (↗↘)

Related Study Notes