Classifier Decisions: Prevalence and Expected Cost
1. Choosing a Model for a Population
Two classifiers can rank differently by expected cost in different populations. A model with higher sensitivity misses fewer actual positives. A model with higher specificity produces fewer false positives. Their relative costs depend on the prevalence and on how the two errors are weighted.
Consider two illustrative settings for a system that flags cases for review:
- A selected population: Suppose that 25% of cases are positive, after an initial screening step.
- A broader population: Suppose that 0.5% of cases are positive.
These prevalence values are examples, not estimates for a particular application. Which model has lower expected cost in each setting if its sensitivity and specificity remain unchanged?
2. Two Models and Their Error Costs
Each model predicts either positive ($\hat{Y} = 1$) or negative ($\hat{Y} = 0$). The rates below are stipulated for this worked comparison.
Detects more actual positives, but also produces more false positives.
Produces fewer false positives, but misses more actual positives.
3. Figure 1: Comparing the Models
Drag the prevalence slider or select a preset to compare the models on 10,000 cases. Change either error cost to see how the lower-cost choice changes. The displayed counts are rounded to whole cases, and the table's metrics and costs use those counts. Very narrow positive-class columns are widened for visibility.
| Performance Metric | Model A | Model B | Dependence on Prevalence |
|---|---|---|---|
| Sensitivity / Recall (TPR) | 96.0% | 80.0% | Held fixed; displayed value reflects rounded counts |
| Specificity / Selectivity (TNR) | 88.0% | 98.5% | Held fixed; displayed value reflects rounded counts |
| Precision (PPV) | -- | -- | Approaches 0 as π approaches 0 for these models |
| Negative Predictive Value (NPV) | -- | -- | Approaches 100% as π approaches 0 for these models |
| F₁-Score | -- | -- | Harmonic mean of precision and recall |
| Matthews Correlation (MCC) | -- | -- | Correlation between actual and predicted binary labels |
| Overall Accuracy | -- | -- | Weights sensitivity by π and specificity by 1−π |
| Approximate Expected Cost ($) | -- | -- | Cost per 10,000 cases, using rounded counts |
| Cost Comparison | -- | -- | Compares these two models under the selected costs |
4. Deriving the Crossover Point $\pi^*$
The crossover point is the prevalence at which the two models have equal expected cost. For these models, the cost difference is linear in prevalence, so we can solve for it directly.
Let $\mathcal{L}_A(\pi)$ and $\mathcal{L}_B(\pi)$ now denote expected cost per case, without the factor $N$:
$$\mathcal{L}_A(\pi) = \pi (1 - \text{TPR}_A) C_{\text{FN}} + (1 - \pi) \text{FPR}_A \cdot C_{\text{FP}}$$ $$\mathcal{L}_B(\pi) = \pi (1 - \text{TPR}_B) C_{\text{FN}} + (1 - \pi) \text{FPR}_B \cdot C_{\text{FP}}$$Set the difference in expected cost to zero: $\Delta \mathcal{L}(\pi) = \mathcal{L}_A(\pi) - \mathcal{L}_B(\pi) = 0$.
$$\pi \left[ (1 - \text{TPR}_A) - (1 - \text{TPR}_B) \right] C_{\text{FN}} + (1 - \pi) \left[ \text{FPR}_A - \text{FPR}_B \right] C_{\text{FP}} = 0$$Using $(1 - \text{TPR}_A) - (1 - \text{TPR}_B) = -(\text{TPR}_A - \text{TPR}_B)$ gives:
$$(1 - \pi) \cdot (\text{FPR}_A - \text{FPR}_B) \cdot C_{\text{FP}} = \pi \cdot (\text{TPR}_A - \text{TPR}_B) \cdot C_{\text{FN}}$$Solving for $\pi$ gives the crossover point:
$$\boxed{\pi^* = \frac{C_{\text{FP}} (\text{FPR}_A - \text{FPR}_B)}{C_{\text{FP}} (\text{FPR}_A - \text{FPR}_B) + C_{\text{FN}} (\text{TPR}_A - \text{TPR}_B)}}$$For the stipulated rates, $\text{FPR}_A - \text{FPR}_B = 0.105$ and $\text{TPR}_A - \text{TPR}_B = 0.160$. With the default $10:1$ cost ratio:
$$\pi^* = \frac{50 \times 0.105}{50 \times 0.105 + 500 \times 0.160} = \frac{5.25}{5.25 + 80.0} = \frac{5.25}{85.25} \approx \mathbf{6.16\%}$$5. Figure 2: Expected Cost Across Prevalence
Figure 2 uses the unrounded expected counts. The curves intersect at $\pi^*$, where the expected costs are equal. Change the error costs above to move the crossover point. Use the logarithmic axis to examine low prevalence values, or the linear axis to see that each cost function is linear in $\pi$.
6. Choosing a Threshold for a Continuous Score
A classifier that produces a continuous score also requires a rule for converting scores to decisions. With fixed class-conditional score distributions, how should that rule change when prevalence changes?
The Bayes decision rule minimizes expected cost. With zero cost for correct decisions and positive costs for errors, it assigns an observation $x$ to the positive class when its likelihood ratio meets this threshold (either decision has the same conditional expected cost at equality):
$$\Lambda(x) = \frac{p(x \mid Y=1)}{p(x \mid Y=0)} \ge \tau^*(\pi) = \left( \frac{1 - \pi}{\pi} \right) \cdot \frac{C_{\text{FP}}}{C_{\text{FN}}}$$The factor $\frac{1 - \pi}{\pi}$ is the prior odds against the positive class. As prevalence decreases, the likelihood ratio needed for a positive decision increases. The following values use the default costs throughout:
| Illustrative Setting | Prevalence π | Prior Odds Factor (1−π)/π | Optimal Likelihood Ratio Threshold τ* (at C_FP/C_FN = 0.1) |
|---|---|---|---|
| Balanced classes | 50.0% | 1.0× | 0.10 |
| 20% positive | 20.0% | 4.0× | 0.40 |
| Default crossover point | 6.16% | 15.2× | 1.52 |
| 1% positive | 1.0% | 99.0× | 9.90 |
| 0.1% positive | 0.1% | 999.0× | 99.90 |
For a posterior probability $q = P(Y=1 \mid x)$ calibrated to the current population, the equivalent rule is $q \ge C_{\text{FP}} / (C_{\text{FP}} + C_{\text{FN}})$. This probability threshold depends on the costs; prevalence is already included in $q$. With equal error costs it is 0.5. When prevalence changes, probabilities calibrated to an earlier population may need adjustment even if the class-conditional distributions stay fixed.
Figure 3 uses a separate illustrative score model: $s \mid Y=0 \sim \mathcal{N}(0,1)$ and $s \mid Y=1 \sim \mathcal{N}(2.2,1)$. These densities stay fixed as you move the prevalence slider. The likelihood ratio increases with $s$, so a higher required likelihood ratio moves the score cutoff to the right. Change prevalence or either error cost above to see the cutoff move.
7. Python Calculation
This script calculates expected cost and precision across prevalence for the same two models and default costs. It requires NumPy and Matplotlib and plots the results without rounding expected counts.
#!/usr/bin/env python3
"""
Classifier Decisions: Prevalence and Expected Cost.
Calculates expected cost and precision for two fixed classifiers
and finds their crossover point under unequal error costs.
"""
import numpy as np
import matplotlib.pyplot as plt
# 1. Model Specifications
MODEL_A = {"name": "Model A (higher sensitivity)", "tpr": 0.960, "fpr": 0.120}
MODEL_B = {"name": "Model B (higher specificity)", "tpr": 0.800, "fpr": 0.015}
# 2. Cost Matrix ($)
COST_FN = 500.0 # Cost per false negative
COST_FP = 50.0 # Cost per false positive
N_COHORT = 10000
def compute_loss(model, pi, c_fn, c_fp):
"""Expected cost per individual trial."""
return pi * (1.0 - model["tpr"]) * c_fn + (1.0 - pi) * model["fpr"] * c_fp
def compute_metrics(model, pi):
"""Computes PPV, NPV, and F1 across prevalence."""
tpr, fpr = model["tpr"], model["fpr"]
tnr = 1.0 - fpr
# Joint probabilities
p_tp = pi * tpr
p_fp = (1.0 - pi) * fpr
p_fn = pi * (1.0 - tpr)
p_tn = (1.0 - pi) * tnr
ppv = p_tp / (p_tp + p_fp) if (p_tp + p_fp) > 0 else 0.0
npv = p_tn / (p_tn + p_fn) if (p_tn + p_fn) > 0 else 0.0
f1 = 2 * p_tp / (2 * p_tp + p_fp + p_fn) if (2 * p_tp + p_fp + p_fn) > 0 else 0.0
return ppv, npv, f1
# 3. Crossover Point
delta_fpr = MODEL_A["fpr"] - MODEL_B["fpr"]
delta_tpr = MODEL_A["tpr"] - MODEL_B["tpr"]
pi_star = (COST_FP * delta_fpr) / (COST_FP * delta_fpr + COST_FN * delta_tpr)
print(f"Crossover point: π* = {pi_star*100:.3f}%")
# 4. Prevalence Sweep
pi_vals = np.logspace(np.log10(0.001), np.log10(0.50), 300)
loss_a = [compute_loss(MODEL_A, p, COST_FN, COST_FP) * N_COHORT for p in pi_vals]
loss_b = [compute_loss(MODEL_B, p, COST_FN, COST_FP) * N_COHORT for p in pi_vals]
# 5. Plotting Results
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(13, 5), dpi=120)
# Panel 1: Expected Cost
ax1.plot(pi_vals * 100, loss_a, label="Model A (higher sensitivity)", color="#0d9488", lw=2.5)
ax1.plot(pi_vals * 100, loss_b, label="Model B (higher specificity)", color="#7c3aed", lw=2.5)
ax1.axvline(pi_star * 100, color="#b8412a", linestyle="--", label=f"Crossover point (π*={pi_star*100:.2f}%)")
ax1.set_xscale("log")
ax1.set_xlabel("Prevalence π (%) [Log Scale]")
ax1.set_ylabel("Expected Cost ($ / 10,000 cases)")
ax1.set_title("Expected Cost Across Prevalence")
ax1.legend()
ax1.grid(True, alpha=0.3)
# Panel 2: Precision (PPV) Trajectories
ppv_a = [compute_metrics(MODEL_A, p)[0] * 100 for p in pi_vals]
ppv_b = [compute_metrics(MODEL_B, p)[0] * 100 for p in pi_vals]
ax2.plot(pi_vals * 100, ppv_a, label="Model A Precision (PPV)", color="#0d9488", lw=2.5)
ax2.plot(pi_vals * 100, ppv_b, label="Model B Precision (PPV)", color="#7c3aed", lw=2.5)
ax2.set_xscale("log")
ax2.set_xlabel("Prevalence π (%) [Log Scale]")
ax2.set_ylabel("Positive Predictive Value (%)")
ax2.set_title("Precision (PPV) Across Prevalence")
ax2.legend()
ax2.grid(True, alpha=0.3)
plt.tight_layout()
plt.show()
8. Applying the Comparison
A comparison on balanced data may rank models differently from a comparison at the deployment prevalence. If class-conditional rates remain stable, expected cost can be recalculated for the new prevalence. Otherwise, new estimates of sensitivity and specificity are needed.
Increasing the cost of a false negative can favor the more sensitive model.