Crosstabs
Menu location: Analysis_Crosstabs.
This a two or three way cross tabulation function. If you have two columns of numbers that correspond to different classifications of the same individuals then you can use this function to give a two way frequency table for the cross classification. This can be stratified by a third classification variable.
For two way crosstabs, StatsDirect offers a range of analyses appropriate to the dimensions of the contingency table. For more information see chi-square tests and exact tests.
For three way crosstabs, StatsDirect offers either odds ratio (for case-control studies) or relative risk (for cohort studies) meta-analyses for 2 by 2 by k tables, and generalised Cochran-Mantel-Haenszel tests for r by c by k tables. For 2 by 2 by k tables the second category of the first classifier (rows) is taken as the exposed group, and the second category of the second classifier (columns) as the outcome of interest; with categories coded 0 and 1, for example, 1 is exposed and 1 is the outcome.
The categories of each classifier are put in the order of their labels: labels that are numbers in order of size, and after them the other labels in the order of their characters. A record with a missing value in one of the classifiers of a table is left out of the table.
If the first and second classifiers have different numbers of categories then you are asked whether to force the table to be symmetrical. If you agree, each classifier is given the categories of the other that it lacks, as empty rows or columns; this is what you want where the same classification has been made twice, as in a study of agreement. Rows and columns without counts take no part in the tests, their degrees of freedom or the measures of association.
Example
A database of test scores contains two fields of interest, sex (M=1, F=0) and grade of skin reaction to an antigen (none = 0, weak + = 1, strong + = 2). Here is a list of those fields for 10 patients:
| Sex | Reaction |
| 0 | 0 |
| 1 | 1 |
| 1 | 2 |
| 0 | 2 |
| 1 | 2 |
| 0 | 1 |
| 0 | 0 |
| 0 | 1 |
| 1 | 2 |
| 1 | 0 |
In order to get a cross tabulation of these from StatsDirect you should enter these data in two workbook columns. Then choose crosstabs from the analysis menu.
For this example:
| Reaction | ||||
| 0 | 1 | 2 | ||
| Sex: | 0 | 2 | 2 | 1 |
| 1 | 1 | 1 | 3 | |
We could then proceed to an r by c (2 by 3) contingency table analysis to look for association between sex and reaction to this antigen:
Contingency table analysis
| Observed | 2 | 2 | 1 | 5 |
| % of row | 40% | 40% | 20% | |
| % of col | 66.67% | 66.67% | 25% | 50% |
| Observed | 1 | 1 | 3 | 5 |
| % of row | 20% | 20% | 60% | |
| % of col | 33.33% | 33.33% | 75% | 50% |
| Total | 3 | 3 | 4 | 10 |
| % of n | 30% | 30% | 40% |
TOTAL number of cells = 6
WARNING: 6 out of 6 cells have EXPECTATION < 5
NOMINAL INDEPENDENCE
Chi-square = 1.666667, DF = 2, P = 0.4346
G-square = 1.726092, DF = 2, P = 0.4219
Fisher-Freeman-Halton exact P = 0.5714
ANOVA
Chi-square for equality of mean column scores = 1.5
DF = 2, P = 0.4724
LINEAR TREND
Sample correlation (r) = 0.361158
Chi-square for linear trend (M²) = 1.173913
DF = 1, P = 0.2786
NOMINAL ASSOCIATION
Phi = 0.408248
Pearson's contingency = 0.377964
Cramér's V = 0.408248
ORDINAL
Goodman-Kruskal gamma = 0.555556
Approximate test of gamma = 0: SE = 0.384107, P = 0.1481, 95% CI = -0.197281 to 1.308392
Approximate test of independence: SE = 0.437445, P = 0.2041, 95% CI = -0.301821 to 1.412932
Kendall tau-b = 0.348155
Approximate test of tau-b = 0: SE = 0.275596, P = 0.2065, 95% CI = -0.192002 to 0.888313
Approximate test of independence: SE = 0.274138, P = 0.2041, 95% CI = -0.189145 to 0.885455
R code
This R code reproduces the example above. It needs no packages and was checked with R 4.6.1. Paste it into R, or save it as a script and run it.
# Crosstabs: the StatsDirect help example (sex and grade of skin reaction to an
# antigen in 10 patients) in R, with the contingency table analysis that follows
sex <- c(0, 1, 1, 0, 1, 0, 0, 0, 1, 1) # M = 1, F = 0
reaction <- c(0, 1, 2, 2, 2, 1, 0, 1, 2, 0) # none = 0, weak = 1, strong = 2
# R's standard cross tabulation, with the percentages of each row and column
tab <- table(Sex = sex, Reaction = reaction)
print(tab)
print(round(100 * prop.table(tab, 1), 2)) # % of row
print(round(100 * prop.table(tab, 2), 2)) # % of col
n <- sum(tab)
six <- function(x) formatC(x, digits = 6, format = "f", drop0trailing = TRUE)
pv <- function(p) {
paste("P =", formatC(p, digits = 4, format = "f", drop0trailing = TRUE))
}
# Nominal independence: Pearson's chi-square without Yates' correction (R warns
# about the small expected counts, as the report does), the likelihood ratio
# G-square, and the Fisher-Freeman-Halton exact test, which fisher.test gives
# for a table larger than 2 by 2
chi <- suppressWarnings(chisq.test(tab, correct = FALSE))
cat("Chi-square =", six(chi$statistic), " DF =", chi$parameter, " ", pv(chi$p.value),
"\n")
g2 <- 2 * sum(tab * log(tab / chi$expected), na.rm = TRUE) # 0 log 0 counts as 0
cat("G-square =", six(g2), " DF =", chi$parameter, " ",
pv(pchisq(g2, chi$parameter, lower.tail = FALSE)), "\n")
cat("Fisher-Freeman-Halton exact", pv(fisher.test(tab)$p.value), "\n")
cat("Expected counts below 5:", sum(chi$expected < 5), "of", length(tab), "\n")
# ANOVA: a chi-square comparing the mean row score (the sex code) between the
# columns, (n - 1) times the between-column sum of squares over the total sum of
# squares, with the number of columns less 1 as its degrees of freedom
between <- sum(tapply(sex, reaction, function(v) length(v) * (mean(v) - mean(sex))^2))
chi_eq <- (n - 1) * between / sum((sex - mean(sex))^2)
df_eq <- ncol(tab) - 1
cat("Chi-square for equality of mean column scores =", six(chi_eq), " DF =", df_eq, " ",
pv(pchisq(chi_eq, df_eq, lower.tail = FALSE)), "\n")
# Linear trend: the correlation between the row and column scores, and
# Mantel-Haenszel's chi-square (n - 1) r^2 with one degree of freedom
r <- cor(sex, reaction)
m2 <- (n - 1) * r^2
cat("Sample correlation (r) =", six(r), "\n")
cat("Chi-square for linear trend (M2) =", six(m2), " DF = 1 ",
pv(pchisq(m2, 1, lower.tail = FALSE)), "\n")
# Nominal association: phi, Pearson's contingency coefficient and Cramer's V
x2 <- as.numeric(chi$statistic)
cat("Phi =", six(sqrt(x2 / n)), "\n")
cat("Pearson's contingency =", six(sqrt(x2 / (x2 + n))), "\n")
cat("Cramer's V =", six(sqrt(x2 / (n * (min(dim(tab)) - 1)))), "\n")
# Ordinal association: Goodman and Kruskal's gamma and Kendall's tau-b from the
# concordant and discordant pairs, each with two standard errors (Agresti 2002,
# Brown and Benedetti 1977): one for a test that the measure is 0, one under
# independence
o <- unclass(tab)
R <- nrow(o)
C <- ncol(o)
conc <- disc <- o * 0
for (i in 1:R) for (j in 1:C) {
conc[i, j] <- sum(o[(1:R) > i, (1:C) > j]) + sum(o[(1:R) < i, (1:C) < j])
disc[i, j] <- sum(o[(1:R) > i, (1:C) < j]) + sum(o[(1:R) < i, (1:C) > j])
}
cc <- sum(o * conc) # concordant pairs, counted twice
dc <- sum(o * disc) # discordant pairs, counted twice
gamma <- (cc - dc) / (cc + dc)
se_gamma <- 4 / (cc + dc)^2 * sqrt(sum(o * (dc * conc - cc * disc)^2))
v_ind <- sum(o * (conc - disc)^2) - (cc - dc)^2 / n
se_gamma_ind <- 2 / (cc + dc) * sqrt(v_ind)
rows <- rowSums(o)
cols <- colSums(o)
dr <- n^2 - sum(rows^2)
dcl <- n^2 - sum(cols^2)
taub <- (cc - dc) / sqrt(dr * dcl)
vij <- outer(rows, cols, function(a, b) a * dcl + b * dr)
vt <- sum(o * (2 * sqrt(dr * dcl) * (conc - disc) + taub * vij)^2) -
n^3 * taub^2 * (dr + dcl)^2
se_taub <- sqrt(vt) / (dr * dcl)
se_taub_ind <- 2 * sqrt(v_ind / (dr * dcl))
z <- qnorm(0.975)
show <- function(label, est, se) {
cat(label, " SE =", six(se), " ", pv(2 * pnorm(-abs(est / se))),
" 95% CI =", six(est - z * se), "to", six(est + z * se), "\n")
}
cat("Goodman-Kruskal gamma =", six(gamma), "\n")
show("Approximate test of gamma = 0:", gamma, se_gamma)
show("Approximate test of independence:", gamma, se_gamma_ind)
cat("Kendall tau-b =", six(taub), "\n")
show("Approximate test of tau-b = 0:", taub, se_taub)
show("Approximate test of independence:", taub, se_taub_ind)