Fisher's Exact Test
Menu location: Analysis_Exact_Fisher.
Like the chi-square test for fourfold (2 by 2) tables, Fisher's exact test examines the relationship between the two dimensions of the table (classification into rows vs. classification into columns). The null hypothesis is that these two classifications are not different.
The P values in this test are computed by considering all possible tables that could give the row and column totals observed. A mathematical short cut relates these permutations to factorials; a form shown in many textbooks. StatsDirect uses the hypergeometric distribution for the calculation (Conover 1999). The test statistic that is hypergeometrically distributed is the first count A; the report shows its expected value.
This exact treatment of the fourfold table should be used instead of the chi-square test when any expected frequency is less than 1 or 20% of expected frequencies are less than or equal to 5. With StatsDirect, it is reasonable to use Fisher's exact test by default because the computational method used can cope with large numbers.
StatsDirect uses the definition of a two sided P value described by Bailey (1977) (P values for all possible tables with P less than or equal to that for the observed table are summed). Some authors prefer simply to double the one sided P value (Armitage and Berry, 1994; Bland, 2000).
Consider using mid-P values and intervals when you have several similar studies within an overall investigation (Armitage and Berry, 1994; Barnard, 1989). Mid-P values are given with the other P values, for a table of any size.
Assumptions:
· each observation is classified into exactly one cell
· the row and column totals are fixed, not random
The assumption of fixed marginal (row/column) totals is controversial and causes disagreements such as the best approach to two sided inference from this test.
DATA INPUT:
Observed frequencies should be entered as a standard fourfold table. The frequencies should be whole numbers; one that is not is rounded to the nearest whole number. A table with an empty row or an empty column cannot be analysed.
| feature present | feature absent | |
| outcome positive: | a | b |
| outcome negative: | c | d |
Example
From Armitage and Berry (1994, p. 138).
The following data compare malocclusion of teeth with method of feeding infants.
| Normal teeth | Malocclusion | |
| Breast fed: | 4 | 16 |
| Bottle fed: | 1 | 21 |
To analyse these data in StatsDirect you must select Fisher's exact test from the exact tests section of the analysis menu. Enter the frequencies into the contingency table on screen as shown above.
For this example:
Rearranged table:
| 4 | 1 | 5 |
| 16 | 21 | 37 |
| 20 | 22 | 42 |
Expectation of A = 2.380952
One sided (upper tail) P = 0.1435 (doubled one sided P = 0.2871)
Two sided (by summation) P = 0.1745
One sided mid-P = 0.0809
Two sided mid-P = 0.1618
Here we cannot reject the null hypothesis that there is no association between these two classifications, i.e. between feeding method and malocclusion.
R code
This R code reproduces the example above. It needs no packages and was checked with R 4.6.1. Paste it into R, or save it as a script and run it.
# Fisher's exact test: the StatsDirect help example (Armitage and Berry 1994, p. 138,
# malocclusion of teeth by method of feeding infants) in R
teeth <- matrix(c(4, 16,
1, 21), 2, byrow = TRUE,
dimnames = list(feeding = c("Breast fed", "Bottle fed"),
teeth = c("Normal", "Malocclusion")))
print(addmargins(teeth))
# R's standard test. Its two sided P is the report's "by summation": the probabilities
# of every table with the same margins that is no more likely than the observed one are
# added up. R also estimates the odds ratio conditional on the margins, with a
# confidence interval, which the report does not show. For a one sided P use
# alternative = "greater" (upper tail for the top left count) or "less" (lower tail).
print(fisher.test(teeth))
# The report's other lines come from the hypergeometric distribution of the top left
# count A given the margins: dhyper(x, m, n, k) is the probability that A = x when m
# is the first column total, n the second column total and k the first row total.
six <- function(x) formatC(x, digits = 6, format = "f", drop0trailing = TRUE)
pv <- function(p) {
if (p < 0.0001) "P < 0.0001" else
paste("P =", formatC(p, digits = 4, format = "f", drop0trailing = TRUE))
}
counts <- as.vector(t(teeth)) # a, b, c, d row by row
# StatsDirect turns the table so that its A is the smaller of the two diagonal counts
# (for this example it only transposes the table, which changes nothing). The P values
# are the same whichever way the table is turned; the expectation is that of A.
if (counts[1] > counts[4]) counts <- rev(counts)
a <- counts[1]
m <- counts[1] + counts[3]
n <- counts[2] + counts[4]
k <- counts[1] + counts[2]
expected <- k * m / (m + n)
cat("Expectation of A =", six(expected), "\n")
x <- max(0, k - n):min(k, m) # the values A can take with these margins
pr <- dhyper(x, m, n, k)
observed <- pr[x == a]
# One sided P: the tail in which A lies, upper when A exceeds its expectation. Turning
# the table never changes which side of its expectation A lies (A - E(A) = (ad - bc)/n).
if (a > expected) {
tail <- "(upper tail)"
one_sided <- sum(pr[x >= a])
} else {
tail <- "(lower tail)"
one_sided <- sum(pr[x <= a])
}
cat("One sided", tail, pv(one_sided), " (doubled one sided",
paste0(pv(min(1, 2 * one_sided)), ")"), "\n")
# Two sided P by summation (Bailey 1977), as fisher.test gives it: the tables no more
# probable than the observed one, allowing a little rounding in the comparison
two_sided <- sum(pr[pr <= observed * (1 + 1e-7)])
cat("Two sided (by summation)", pv(min(1, two_sided)), "\n")
# Mid-P: only half the probability of the observed table is counted in its tail
mid_p <- one_sided - observed / 2
cat("One sided mid-", pv(mid_p), "\n", sep = "")
cat("Two sided mid-", pv(min(1, 2 * mid_p)), "\n", sep = "")