Diagnostic Test 2 by 2 Table
Menu location: Analysis_Clinical Epidemiology_Diagnostic Test (2 by 2).
This function gives predictive values (post-test likelihood) with change, prevalence (pre-test likelihood), sensitivity, specificity and likelihood ratios with robust confidence intervals (Interpretation of diagnostic data, 1983; Sackett et al., 1991; Zhou et al., 2002).
The quality of a diagnostic test is often expressed in terms of sensitivity and specificity. Sensitivity is the ability of the test to pick up what it is testing for and specificity is the ability of the test to reject what it is not testing for.
| DISEASE | |||
| Present | Absent | ||
| TEST: | +: | a (true +ve) | b (false +ve) |
| -: | c (false -ve) | d (true -ve) | |
Sensitivity = a/(a+c)
Specificity = d/(b+d)
+ve predictive value = a/(a+b)
-ve predictive value = d/(d+c)
Likelihood ratio of a positive test = [a/(a+c)]/[b/(b+d)]
Likelihood ratio of a negative test = [c/(a+c)]/[d/(b+d)]
Likelihood ratios have become useful because they enable one to quantify the effect a particular test result has on the probability of a certain diagnosis or outcome. Using a simplified form of Bayes' theorem:
posterior odds = prior odds * likelihood ratio
where:
odds = probability/(1-probability)
probability = odds/(odds+1)
This function is not truly Bayesian because it does not use any starting/prior probability. Likelihood ratios, however, are provided and these can be used to direct the flow of probabilities in Bayesian analysis. For an excellent account of this approach in medical diagnosis, see Sackett (1991).
Another way to summarise diagnostic test performance is via the diagnostic odds ratio:
Diagnostic odds ratio = true/false = (a * d)/(b * c)
Technical validation
The confidence intervals for the likelihood ratios are constructed using the likelihood-based approach to binomial proportions of Koopman (1984) suggested by Gart and Nam (1988). The confidence intervals for all other statistics are exact binomial confidence intervals constructed using the method of Clopper and Pearson (Newcombe, 1998c).
Example
In a hypothetical example of a diagnostic test, serum levels of a biochemical marker of a particular disease were compared with the known diagnosis of the disease. 100 international units of the marker or greater was taken as an arbitrary positive test result:
| Disease | No Disease | |
| Marker ≥100: | 431 | 30 |
| Marker < 100: | 29 | 116 |
To analyse these data in StatsDirect select Diagnostic Test (2 by 2) from the Clinical Epidemiology section of the Analysis menu. Choose the default 95% confidence interval. Then enter the above frequencies into the 2 by 2 table on the screen.
For this example:
| Disease / Feature: | ||||
| Present | Absent | Totals | ||
| Test: | Positive: | 431 | 30 | 461 |
| Negative: | 29 | 116 | 145 | |
| Totals: | 460 | 146 | 606 | |
Including 95% confidence intervals:
Prevalence (pre-test likelihood of disease)
0.759076 (0.722988 to 0.792617), 75.91% (72.3% to 79.26%)
Predictive value of +ve test (post-test likelihood of disease)
0.934924 (0.908401 to 0.955666), 93.49% (90.84% to 95.57%), {change = 17%}
Predictive values of -ve test
(post-test likelihood of no disease)
0.8 (0.725563 to 0.861777), 80% (72.56% to 86.18%), {change = 56%}
(post-test disease likelihood despite -ve test)
0.2 (0.138223 to 0.274437), 20% (13.82% to 27.44%), {change = -56%}
Sensitivity (true positive rate)
0.936957 (0.910711 to 0.957376), 93.7% (91.07% to 95.74%)
Specificity (true negative rate)
0.794521 (0.719844 to 0.856862), 79.45% (71.98% to 85.69%)
Likelihood Ratio
LR (positive test) = 4.559855 (3.364957 to 6.340323)
LR (negative test) = 0.079348 (0.055211 to 0.113307)
Diagnostic Odds Ratio
Observed odds ratio = 57.466667
Conditional maximum likelihood estimate = 56.64839 (32.064885 to 103.54465)
Here we can say with 95% confidence that marker results of ≥ 100 are at least three (3.365) times more likely to come from patients with disease than those without disease. Also with 95% confidence we can say that marker results of <100 are at most only about one tenth (0.113) as likely to come from patients with disease as from those without disease.
R code
This R code reproduces the example above. It needs no packages and was checked with R 4.6.1. Paste it into R, or save it as a script and run it.
# Diagnostic test 2 by 2 table: the StatsDirect help example (a hypothetical serum
# marker of disease, positive at 100 units or more: 431 diseased and 30 disease-free
# patients tested positive, 29 and 116 tested negative) in R
tp <- 431 # a: test positive, disease present
fp <- 30 # b: test positive, disease absent
fn <- 29 # c: test negative, disease present
tn <- 116 # d: test negative, disease absent
n <- tp + fp + fn + tn
counts <- matrix(c(tp, fp, fn, tn), 2, byrow = TRUE,
dimnames = list(Test = c("Positive", "Negative"),
Disease = c("Present", "Absent")))
print(addmargins(counts))
# Each rate in the report is a binomial proportion with the exact (Clopper-Pearson)
# confidence interval that binom.test gives; the prevalence, for instance:
print(binom.test(tp + fn, n))
# The report prints each rate as a proportion and as a percentage
six <- function(x) formatC(x, digits = 6, format = "f", drop0trailing = TRUE)
pc <- function(x) {
paste0(formatC(100 * x, digits = 2, format = "f", drop0trailing = TRUE), "%")
}
rate <- function(x, n) {
ci <- binom.test(x, n)$conf.int
paste0(six(x / n), " (", six(ci[1]), " to ", six(ci[2]), "), ", pc(x / n), " (",
pc(ci[1]), " to ", pc(ci[2]), ")")
}
prevalence <- (tp + fn) / n
cat("Prevalence (pre-test likelihood of disease)\n")
cat(rate(tp + fn, n), "\n")
# The change, in braces, is from the pre-test to the post-test likelihood, in whole
# percentages (halves round to even, as the program does; the pre-test rate is
# evaluated as the program evaluates it, so that both round alike)
ppv <- tp / (tp + fp)
cat("Predictive value of +ve test (post-test likelihood of disease)\n")
cat(rate(tp, tp + fp), ", {change = ", round(100 * ppv) - round(100 * prevalence),
"%}\n", sep = "")
npv <- tn / (tn + fn)
cat("Predictive values of -ve test\n")
cat("(post-test likelihood of no disease)\n")
cat(rate(tn, tn + fn), ", {change = ", round(100 * npv) - round((fp + tn) / n * 100),
"%}\n", sep = "")
# 1 - npv, whose exact interval is that of the false negatives among the test negatives
cat("(post-test disease likelihood despite -ve test)\n")
cat(rate(fn, tn + fn), ", {change = ",
round(100 * (1 - npv)) - round(100 * prevalence), "%}\n", sep = "")
sens <- tp / (tp + fn)
cat("Sensitivity (true positive rate)\n")
cat(rate(tp, tp + fn), "\n")
spec <- tn / (fp + tn)
cat("Specificity (true negative rate)\n")
cat(rate(tn, fp + tn), "\n")
# A likelihood ratio is a ratio of two binomial proportions: the positive rate among
# the diseased over the positive rate among the disease-free, and the negative rate
# among the diseased over the negative rate among the disease-free. Its interval is
# Koopman's (1984) score interval: the ratios at which the score chi-square, with the
# proportions estimated under the constraint that they have that ratio, reaches
# the 95% critical value
koopman <- function(x1, n1, x0, n0, level = 0.95) {
chi2 <- function(theta) {
A <- (n0 + n1) * theta
B <- -((x0 + n1) * theta + x1 + n0)
p0 <- (-B - sqrt(B^2 - 4 * A * (x0 + x1))) / (2 * A) # constrained estimate
p1 <- theta * p0
(x1 - n1 * p1)^2 / (n1 * p1 * (1 - p1)) *
(1 + n1 * (theta - p1) / (n0 * (1 - p1)))
}
est <- (x1 / n1) / (x0 / n0) # both counts non-zero here
f <- function(lt) chi2(exp(lt)) - qchisq(level, 1)
lower <- exp(uniroot(f, c(log(est) - 20, log(est)), tol = 1e-12)$root)
upper <- exp(uniroot(f, c(log(est), log(est) + 20), tol = 1e-12)$root)
c(est, lower, upper)
}
cat("Likelihood Ratio\n")
lr_pos <- koopman(tp, tp + fn, fp, fp + tn)
cat("LR (positive test) = ", six(lr_pos[1]), " (", six(lr_pos[2]), " to ",
six(lr_pos[3]), ")\n", sep = "")
lr_neg <- koopman(fn, tp + fn, tn, fp + tn)
cat("LR (negative test) = ", six(lr_neg[1]), " (", six(lr_neg[2]), " to ",
six(lr_neg[3]), ")\n", sep = "")
# The diagnostic odds ratio: observed, then the conditional maximum likelihood
# estimate with its exact interval, which fisher.test gives
cat("Diagnostic Odds Ratio\n")
cat("Observed odds ratio =", six(tp * tn / (fp * fn)), "\n")
print(fisher.test(counts))
# fisher.test's estimate and limits agree with the report's to about four significant
# figures. Solving the same equations closely reproduces the report: given the
# margins, the true positive count follows the non-central hypergeometric
# distribution with the odds ratio as its parameter; the estimate makes its mean the
# observed count, and each limit puts a tail probability of 2.5% at the observed count
m1 <- tp + fp
n1 <- tp + fn
x <- max(0, m1 - (n - n1)):min(m1, n1)
logd <- dhyper(x, n1, n - n1, m1, log = TRUE)
dens <- function(log_or) {
w <- exp(logd + x * log_or - max(logd + x * log_or))
w / sum(w)
}
solve_or <- function(f) exp(uniroot(f, c(-40, 40), tol = 1e-13)$root)
cmle <- solve_or(function(lo) sum(x * dens(lo)) - tp)
cmle_lower <- solve_or(function(lo) sum(dens(lo)[x >= tp]) - 0.025)
cmle_upper <- solve_or(function(lo) sum(dens(lo)[x <= tp]) - 0.025)
cat("Conditional maximum likelihood estimate = ", six(cmle), " (", six(cmle_lower),
" to ", six(cmle_upper), ")\n", sep = "")