Woolf Statistics for 2 by 2 Tables and Series
Menu location: Analysis_Chi-square_Woolf.
In case-control studies observed frequencies can often be represented by a series of two by two tables. Each stratum of this series represents observations taken at different times, different places or another system of sub-grouping within one large study.
A pooled odds ratio for all strata can be calculated by the method of Mantel and Haenszel or that of Woolf. The Mantel-Haenszel method is more robust when some of the strata contain small frequencies.
Results are given for individual tables and for the combined statistics both without and with the Haldane correction (0.5 added to each cell for the log odds ratio, with variance 1/(a+1) + 1/(b+1) + 1/(c+1) + 1/(d+1)), including chi-square for heterogeneity between the tables.
DATA INPUT:
Observed frequencies should be entered as multiple fourfold tables:
| feature present | feature absent | |
| outcome positive: | a | b |
| outcome negative: | c | d |
Example
From Armitage and Berry (1994, p. 516).
The following data compare the smoking status of lung cancer patients with controls. Ten different studies are combined in an attempt to improve the overall estimate of relative risk. The matching of controls has been ignored because there was not enough information about matching from each study to be sure that the matching was the same in each study.
| Lung cancer | Controls | ||
| smoker | non-smoker | smoker | non-smoker |
| 83 | 3 | 72 | 14 |
| 90 | 3 | 227 | 43 |
| 129 | 7 | 81 | 19 |
| 412 | 32 | 299 | 131 |
| 1350 | 7 | 1296 | 61 |
| 60 | 3 | 106 | 27 |
| 459 | 18 | 534 | 81 |
| 499 | 19 | 462 | 56 |
| 451 | 39 | 1729 | 636 |
| 260 | 5 | 259 | 28 |
To analyse these data in StatsDirect you must select the Woolf function from the chi-square section of the analysis menu. Then enter each row of the table above as a separate 2 by 2 contingency table:
i.e. The first row is entered as:
|
|
Smkr |
Non |
|
Lung cancer |
83 |
3 |
|
Control |
72 |
14 |
... this is then repeated for each of the ten rows.
For this example:
For combined tables without Haldane correction:
Number of tables = 10
Mean log odds ratio = 1.531495 giving odds ratio = 4.625084
Variance of mean log odds ratio = 0.009478 standard error = 0.097355
Approximate 95% CI for mean log odds ratio = 1.340683 to 1.722306
Giving odds ratio of 3.821652 to 5.597423
Chi² for expected log odds ratio = 0 is 247.466729 Chi = 15.731075 P < 0.0001
Chi² for heterogeneity = 6.62565 DF = 9 P = 0.676
For combined tables with Haldane correction:
Number of tables = 10
Mean log odds ratio = 1.506344 giving odds ratio = 4.510211
Variance of mean log odds ratio = 0.00893 standard error = 0.0945
Approximate 95% CI for mean log odds ratio = 1.321127 to 1.691561
Giving odds ratio of 3.747641 to 5.427948
Chi² for expected log odds ratio = 0 is 254.086475 Chi = 15.94009 P < 0.0001
Chi² for heterogeneity = 6.532642 DF = 9 P = 0.6857
Here we can say that there was no convincing evidence of heterogeneity between the separate estimates of relative risk from each of the different studies. The pooled estimate suggested with 95% confidence that the true population odds for being a smoker were between 3.7 and 5.4 times greater in lung cancer patients compared with controls.
The equivalent analysis using the Mantel-Haenszel method gave a confidence interval for the pooled odds ratio of 3.9 to 5.7; the difference is partly accounted for by the Haldane correction. You should use the more robust Mantel-Haenszel for most analyses of this kind. Woolf's method is included for further investigation of inter-table relationships under expert statistical guidance.
R code
This R code reproduces the example above. It needs no packages and was checked with R 4.6.1. Paste it into R, or save it as a script and run it.
# Woolf chi-square analysis of 2 by 2 series: the StatsDirect help example (Armitage
# and Berry 1994, p. 516, smoking among lung cancer patients and controls in ten
# studies; the test workbook's Meta worksheet columns Smokers total, Smokers cancer,
# Control total and Control cancer) in R
# a = Smokers cancer, b = Control cancer, c = Smokers total - a, d = Control total - b
a <- c(83, 90, 129, 412, 1350, 60, 459, 499, 451, 260) # lung cancer, smoker
b <- c(3, 3, 7, 32, 7, 3, 18, 19, 39, 5) # lung cancer, non-smoker
c <- c(72, 227, 81, 299, 1296, 106, 534, 462, 1729, 259) # control, smoker
d <- c(14, 43, 19, 131, 61, 27, 81, 56, 636, 28) # control, non-smoker
k <- length(a)
# Base R has no Woolf function, so the statistics are computed by Woolf's formulae: each
# table's log odds ratio is weighted by the reciprocal of its variance, the weighted
# mean is the pooled log odds ratio, and the weighted sum of squares about that mean
# is the chi-square for heterogeneity between the tables, as Woolf proposed
six <- function(x) formatC(x, digits = 6, format = "f", drop0trailing = TRUE)
pv <- function(p) {
if (p < 0.0001) "P < 0.0001" else
paste("P =", formatC(p, digits = 4, format = "f", drop0trailing = TRUE))
}
z <- qnorm(0.975)
woolf <- function(lor, v) {
w <- 1 / v
mean_lor <- sum(w * lor) / sum(w)
se <- sqrt(1 / sum(w))
ci <- mean_lor + c(-1, 1) * z * se
chi2 <- (mean_lor / se)^2
het <- sum(w * lor^2) - sum(w * lor)^2 / sum(w)
cat("Number of tables =", k, "\n")
cat("Mean log odds ratio =", six(mean_lor), "giving odds ratio =",
six(exp(mean_lor)), "\n")
cat("Variance of mean log odds ratio =", six(se^2), " standard error =", six(se),
"\n")
cat("Approximate 95% CI for mean log odds ratio =", six(ci[1]), "to", six(ci[2]),
"\n")
cat("Giving odds ratio of", six(exp(ci[1])), "to", six(exp(ci[2])), "\n")
cat("Chi-square for expected log odds ratio = 0 is", six(chi2), "Chi =",
six(mean_lor / se), pv(pchisq(chi2, 1, lower.tail = FALSE)), "\n")
cat("Chi-square for heterogeneity =", six(het), "DF =", k - 1,
pv(pchisq(het, k - 1, lower.tail = FALSE)), "\n")
}
# Without the Haldane correction: log(ad/bc) with variance 1/a + 1/b + 1/c + 1/d,
# which needs every cell of every table to be greater than zero
cat("For combined tables without Haldane correction:\n")
woolf(log(a * d / (b * c)), 1 / a + 1 / b + 1 / c + 1 / d)
# With the Haldane correction: 0.5 is added to each cell for the log odds ratio, and
# its variance is 1/(a+1) + 1/(b+1) + 1/(c+1) + 1/(d+1)
cat("For combined tables with Haldane correction:\n")
woolf(log((a + 0.5) * (d + 0.5) / ((b + 0.5) * (c + 0.5))),
1 / (a + 1) + 1 / (b + 1) + 1 / (c + 1) + 1 / (d + 1))
# The Mantel-Haenszel pooled odds ratio that the text compares with, from R's standard
# test, without its continuity correction
tables <- array(rbind(a, c, b, d), c(2, 2, k)) # each table filled column by column
print(mantelhaen.test(tables, correct = FALSE))