Sample Size for Independent Case-control Studies
Menu location: Analysis_Sample Size_Independent Case-Control.
This function gives the minimum number of case subjects required to detect a real odds ratio or case exposure rate with power POWER and two sided type I error probability ALPHA. This sample size is also given as a continuity-corrected value intended for use with corrected chi-square and Fisher's exact tests (Schlesselman, 1982; Casagrande et al. 1978; Dupont, 1990).
Information required
- POWER: probability of detecting a real effect.
- ALPHA: probability of detecting a false effect (two sided: double this if you need one sided).
- P0: probability of exposure in controls.
- *: Input either P1 or OR.
- P1: probability of exposure in case subjects.
- OR: odds ratio of exposures between cases and controls.
- M: number of control subjects per case subject.
Practical issues
- Usual values for POWER are 80%, 85% and 90%; try several in order to explore/scope.
- 5% is the usual choice for ALPHA.
- P0 can be estimated as the population prevalence of exposure.
- If possible, choose a range of odds ratios that you want to have the statistical power to detect.
Technical validation
The estimated sample size n is calculated as:
- where α = alpha, β = 1 - power, ψ = odds ratio, nc is the continuity corrected sample size, m is the number of control subjects per case, and zp is the standard normal deviate for probability p. n is rounded up to the closest integer.
Example
Suppose that 20% of the population from which the controls will be drawn are exposed to the factor under study, and that you plan a study with one control per case to detect an odds ratio of 2 with 80% power at the 5% two sided significance level. The figures are invented for this illustration.
To run this in StatsDirect select Independent Case-Control from the Sample Size section of the Analysis menu. Enter 0.2 as the probability of exposure in controls, choose Odds Ratio and enter 2 as the odds ratio, enter 1 as the number of controls per case, and accept 80% power and 5% alpha.
For this example:
Sample size for independent case-control study
Probability of exposure in controls = 0.2
Probability of exposure in cases = 0.333333
Controls per case subject = 1
Alpha = 0.05
Power = 0.8
For uncorrected chi-square test:
N = 172 case subjects and 172 controls
For corrected chi-square and Fisher's exact tests:
N = 187 case subjects and 187 controls
An odds ratio of 2 with 20% of controls exposed means that a third of the cases are expected to be exposed. 172 cases and 172 controls would be enough for the uncorrected chi-square test, or 187 of each if the analysis will use the continuity-corrected chi-square test or Fisher's exact test. With two controls per case the same study needs 126 cases and 252 controls (137 and 274 with the correction): more controls per case reduce the number of cases needed, with diminishing returns beyond about four per case.
R code
This R code reproduces the example above. It needs no packages and was checked with R 4.6.1. Paste it into R, or save it as a script and run it.
# Sample size for an independent case-control study: the StatsDirect help example
# (invented figures: 20% of the controls are exposed and an odds ratio of 2 is to be
# detected, one control per case, 80% power, 5% two sided alpha) in R
p0 <- 0.2 # probability of exposure in the controls
or <- 2 # the odds ratio to detect
p1 <- p0 * or / (1 + p0 * (or - 1)) # the probability of exposure in the cases
m <- 1 # controls per case
power <- 0.8
alpha <- 0.05
six <- function(x) formatC(x, digits = 6, format = "f", drop0trailing = TRUE)
cat("Probability of exposure in controls =", six(p0), "\n")
cat("Probability of exposure in cases =", six(p1), "\n")
cat("Controls per case subject =", m, "\n")
cat("Alpha =", alpha, "\n")
cat("Power =", power, "\n")
# power.prop.test (two sided, its default strict = FALSE) gives the same n as the
# formula below when m = 1: it takes one size per group, so it applies only to one
# control per case
print(power.prop.test(p1 = p0, p2 = p1, power = power, sig.level = alpha))
# The report's figures, from the formula in the topic, for any number m of controls per
# case: n is rounded up to a whole number, and the number of controls is m times that,
# rounded down
z_alpha <- qnorm(1 - alpha / 2)
z_beta <- qnorm(power)
pbar <- (p1 + m * p0) / (1 + m)
n <- (z_alpha * sqrt((1 + 1 / m) * pbar * (1 - pbar)) +
z_beta * sqrt(p1 * (1 - p1) + p0 * (1 - p0) / m))^2 / (p1 - p0)^2
cat("For uncorrected chi-square test:\n")
cat("N =", ceiling(n), "case subjects and", floor(m * ceiling(n)), "controls\n")
# The continuity-corrected size for the corrected chi-square and Fisher's exact tests
# (Casagrande et al. 1978), calculated from the unrounded n
nc <- n / 4 * (1 + sqrt(1 + 2 * (m + 1) / (n * m * abs(p0 - p1))))^2
cat("For corrected chi-square and Fisher's exact tests:\n")
cat("N =", ceiling(nc), "case subjects and", floor(m * ceiling(nc)), "controls\n")