A model that's 97% accurate but misses 60% of critical cases is a liability.
A model that's 97% accurate but misses 60% of critical cases is a liability. We explain why sensitivity/specificity calibrated to clinical risk thresholds is the only metric that matters in production.
Imagine a clinical AI system tasked with flagging patients at high risk of sepsis. A dataset where 5% of patients develop sepsis. A naive model that predicts "no sepsis" for every patient achieves 95% accuracy. It also catches zero cases.
This isn't a contrived example, it's a common failure mode in clinical AI systems evaluated against headline accuracy. The metric looks good because the dataset is imbalanced. The system is useless because it misses what matters.
In clinical settings, the relevant metrics are sensitivity (true positive rate, of all patients who actually had the condition, how many did we flag?) and specificity (true negative rate, of all patients who didn't have the condition, how many did we correctly clear?).
These two metrics exist in tension. Increasing sensitivity (catching more true cases) typically decreases specificity (flagging more false alarms). The calibration of this tradeoff is a clinical decision, not an engineering one.
For a sepsis detection model, a clinician might accept 70% specificity to achieve 95% sensitivity, i.e., generate more false alerts to avoid missing any real cases, because the cost of missing sepsis vastly exceeds the cost of an unnecessary clinical review.
For a radiology AI flagging potential tumours for radiologist review, a different calibration makes sense. The cost of a false negative (missed tumour) and the cost of a false positive (unnecessary biopsy) need to be weighed against each other, and the threshold set accordingly.
The engineering team cannot make this call. The clinical team must. And the metric framework must be agreed before a single model is trained.
When we build clinical AI systems, the accuracy target is set in a joint session with the clinical team before architecture begins. The evaluation plan specifies: primary metric (e.g., sensitivity at target specificity), acceptable operating range, clinical sign-off required before production deployment, and monitoring cadence post-launch.
No system goes to production without a clinician reviewing the validation results against the agreed thresholds. This is not bureaucracy, it's the only way to build clinical AI responsibly.
We run 2-hour feasibility calls at no cost. We'll tell you what applies to your project.