Key measures
Sensitivity, specificity, PPV, NPV, and likelihood ratios with 95% CI.
Editor
Output
Values are shown with 95% confidence intervals where applicable.
Companion tool
Two raters classify the same items as positive / negative. Enter the four counts to get Cohen's κ with 95% CI (Fleiss 1969) and a Landis–Koch qualitative label. Useful for imaging or pathology cases where the "test" is a human judgement.
Quantify the test's ability to correctly identify people with and without the condition. Useful for understanding the intrinsic performance of the test. Learn more at Wikipedia.
Depend on prevalence (pre-test probability). Use them to estimate the probability that a person has or does not have the disease after a positive or negative result. See Wikipedia for further reading.
Relate the probability of a result in diseased vs. non-diseased individuals and can be applied directly in Bayes' formula to update the probability of disease. Further reference at Wikipedia.
Result from applying the LR to pre-test odds. They are crucial for clinical decisions such as initiating treatment or ordering additional tests. See also Wikipedia.
Critical appraisal
Numbers from a 2×2 only mean what the underlying study allows them to mean. The STARD 2015 checklist has 30 items; six of them dominate the interpretation of sens / spec / PPV / NPV.
If the gold standard is imperfect (e.g. expert consensus instead of biopsy), sensitivity and specificity are biased. Worse: if the reference standard incorporates the index test result, agreement is inflated (incorporation bias). Ask which standard was used and whether the index test fed into it.
Performance estimated in obvious cases ("sick" vs healthy volunteers) almost always overestimates real-world accuracy. Check whether enrolment was consecutive or random, and whether it spanned the clinical spectrum you'll apply the test to (mild, atypical, comorbid).
Partial verification bias arises when positive-index-test patients are more likely to receive the gold standard than negative ones — inflates sensitivity, deflates specificity. Look for differential or post-hoc verification protocols.
If the person interpreting the index test also knew the reference result (or vice versa), expect review bias: estimates drift toward perfect agreement. Imaging and pathology studies should explicitly report blinding between readers and from clinical context.
If indeterminate, missing, or excluded results were silently dropped, the 2×2 misrepresents real practice. Ask how many results fell outside the binary scheme and how they were handled (re-tested, excluded, counted as negative…).
Sensitivity / specificity often vary by age, severity, prior testing, or setting. A single 2×2 hides those differences. Check whether the authors pre-specified clinically meaningful subgroups and reported them separately.
Educational tool
Interactive calculator for exploring diagnostic test performance measures and applying Bayesian reasoning in clinical practice.
Sensitivity, specificity, PPV, NPV, and likelihood ratios with 95% CI.
Convert pre-test probability into positive and negative post-test probability.
Results explained to support self-directed learning.