Normality Tests
Menu location: Analysis_Parametric_Normality.
This function enables you to explore the distribution of a sample and test for certain patterns of non-normality.
A normal probability plot is provided, after some basic descriptive statistics and five hypothesis tests. The common hypothesis is that the sample has been drawn at random from a normal distribution. The test of skewness focuses on the symmetry of the distribution; the test of kurtosis examines its peakedness; and the omnibus chi-square (or k-square) test combines these characteristics (D'Agostino et al., 1990). These tests are calculated using moments about the mean, therefore they are quite sensitive to outliers. You might wish to clean your dataset carefully before exploring its distribution in this way. The two other tests are semi-parametric analyses of variance: Shapiro-Wilk W (Conover, 1999; Shapiro and Wilk, 1965; Royston, 1982a, 1982b, 1991a, 1995) and Shapiro-Francia W' (Shapiro and Francia, 1972; Royston 1983).
StatsDirect requires a random sample of between 3 and 2,000 for the Shapiro-Wilk test, or between 5 and 5,000 for the Shapiro-Francia test. The omnibus chi-square test can be used with larger samples but requires a minimum of 8 observations.
In addition to P values the semi-parametric tests provide V or V' which have a value approximately equal to 1 for samples from normal distributions. Large values of V or V' indicate non-normality. 95% critical values for V lie between 1.2 and 2.4 depending on sample size, and 2.0 and 2.8 for V' (Royston 1991a).
The omnibus chi-square test is updated from the original method to adjust for the fact that the test statistic is not quite distributed as chi-square with 2 degrees of freedom (Royston 1991b).
Significant P values for any of the tests indicate non-normality. There is no perfect test for all forms of non-normality therefore we have provided several here. In addition to the tests you should always examine the normal probability plot. Deviations from the straight line on the plot indicate non-normality. Also look at histograms and spread/dot-plots to examine the distribution of your data. Investigate outlying observations and consider excluding them. You might also wish to transform your data, e.g. into natural logarithms, then examine the normality of the transformed data.
Example
Test workbook (Parametric worksheet: Penicillin).
Consider the following 30 penicillin yields:
| 0.0987 | 0.0000 |
| 0.0533 | -0.0026 |
| 0.0293 | -0.0036 |
| 0.0246 | -0.0042 |
| 0.0200 | -0.0114 |
| 0.0194 | -0.0139 |
| 0.0191 | -0.0222 |
| 0.0180 | -0.0333 |
| 0.0172 | -0.0348 |
| 0.0132 | -0.0363 |
| 0.0102 | -0.0363 |
| 0.0084 | -0.0402 |
| 0.0077 | -0.0583 |
| 0.0058 | -0.1184 |
| 0.0016 | -0.1420 |
To test these data for non-normality using StatsDirect you must first prepare them in a workbook column. Alternatively, open the test workbook using the file open function of the file menu. Then select the normality test from the parametric methods section of the analysis menu. Select the column marked "Penicillin" when prompted.
For this example:
| Normality | |
| Sample name | Penicillin |
| Sample size | 30 |
| Mean | -0.007033 |
| Standard deviation | 0.0454 |
| Skewness | -0.942148, P = 0.025 |
| Kurtosis | 5.373413, P = 0.0156 |
| Royston chi-sq | 9.081201, P = 0.0107 |
| Shapiro-Wilk W | 0.892184, V = 3.426927, P = 0.0054 |
| Shapiro-Francia W' | 0.873427, V' = 4.442896, P = 0.0034 |
| Sample unlikely to be from a normal distribution | |
Here the test statistics are clearly significant at P = 0.05 which rejects the null hypothesis that these data are from a normal distribution. The departures from the straight line in the normal plot demonstrate this pattern in more detail. In fact, these data were from a 2 by 5 factor grouping experiment.
R code
This R code reproduces the example above. It needs no packages and was checked with R 4.6.1. Paste it into R, or save it as a script and run it.
# Normality tests: the StatsDirect help example (30 penicillin yields) in R
yield <- c(0.0987, 0.0000, 0.0533, -0.0026, 0.0293, -0.0036, 0.0246, -0.0042, 0.0200,
-0.0114, 0.0194, -0.0139, 0.0191, -0.0222, 0.0180, -0.0333, 0.0172, -0.0348,
0.0132, -0.0363, 0.0102, -0.0363, 0.0084, -0.0402, 0.0077, -0.0583, 0.0058,
-0.1184, 0.0016, -0.1420)
n <- length(yield)
# R's standard test gives the Shapiro-Wilk W and its P (Royston 1995), as
# StatsDirect reports them.
sw <- shapiro.test(yield)
print(sw)
six <- function(x) formatC(x, digits = 6, format = "f", drop0trailing = TRUE)
pv <- function(p) {
paste("P =", formatC(p, digits = 4, format = "f", drop0trailing = TRUE))
}
cat("Sample size", n, " Mean", six(mean(yield)), " Standard deviation",
six(sd(yield)), "\n")
# Moments about the mean give the skewness (sqrt b1) and kurtosis (b2)
m2 <- mean((yield - mean(yield))^2)
m3 <- mean((yield - mean(yield))^3)
m4 <- mean((yield - mean(yield))^4)
skew <- m3 / m2^1.5
kurt <- m4 / m2^2
# Test of skewness (D'Agostino 1970), needing at least 8 observations
stopifnot(n >= 8)
y <- skew * sqrt((n + 1) * (n + 3) / (6 * (n - 2)))
beta2 <- 3 * (n^2 + 27 * n - 70) * (n + 1) * (n + 3) /
((n - 2) * (n + 5) * (n + 7) * (n + 9))
w2 <- -1 + sqrt(2 * (beta2 - 1))
z1 <- asinh(y / sqrt(2 / (w2 - 1))) / sqrt(log(sqrt(w2)))
cat("Skewness", paste0(six(skew), ","), pv(2 * pnorm(-abs(z1))), "\n")
# Test of kurtosis (Anscombe and Glynn 1983)
xx <- (kurt - 3 * (n - 1) / (n + 1)) /
sqrt(24 * n * (n - 2) * (n - 3) / ((n + 1)^2 * (n + 3) * (n + 5)))
mo <- 6 * (n^2 - 5 * n + 2) / ((n + 7) * (n + 9)) *
sqrt(6 * (n + 3) * (n + 5) / (n * (n - 2) * (n - 3)))
a <- 6 + 8 / mo * (2 / mo + sqrt(1 + 4 / mo^2))
z2 <- (1 - 2 / (9 * a) - ((1 - 2 / a) / (1 + xx * sqrt(2 / (a - 4))))^(1 / 3)) /
sqrt(2 / (9 * a))
cat("Kurtosis", paste0(six(kurt), ","), pv(2 * pnorm(-abs(z2))), "\n")
# Omnibus test: K2 = Z1^2 + Z2^2 (D'Agostino, Belanger and D'Agostino 1990) with
# Royston's (1991) correction, which makes it closer to chi-square with 2 df.
# The corrected K2 is -2 log P.
zc <- -qnorm(exp(-(z1^2 + z2^2) / 2)) # K2's P as a normal deviate
ln <- log(n)
cut <- 0.55 * n^0.2 - 0.21
a1 <- (-5 + 3.46 * ln) * exp(-1.37 * ln)
b1 <- 1 + (0.854 - 0.148 * ln) * exp(-0.55 * ln)
b2 <- 2.13 / (1 - 2.37 * ln)
zr <- if (zc < -1) zc else if (zc < cut) a1 + b1 * zc else
a1 - b2 * cut + (b2 + b1) * zc
pk <- pnorm(-zr)
cat("Royston chi-sq", paste0(six(-2 * log(pk)), ","), pv(pk), "\n")
# StatsDirect adds V = (1 - W) / the median of 1 - W under normality, from the
# approximation for 12 or more observations (it uses another for 4 to 11)
stopifnot(n >= 12)
mu <- -1.5861 + ln * (-0.31082 + ln * (-0.083751 + ln * 0.0038915))
cat("Shapiro-Wilk W", paste0(six(sw$statistic), ","),
"V =", paste0(six((1 - sw$statistic) / exp(mu)), ","), pv(sw$p.value), "\n")
# Shapiro-Francia W': the squared correlation between the ordered values and
# Blom's normal scores for their positions, with P and V' from Royston's (1983)
# normalising transformation
x <- sort(yield)
m <- qnorm((1:n - 0.375) / (n + 0.25))
Wf <- cor(x, m)^2
h <- ln - 5
l <- -0.0480157 + h * (0.01971964 - 0.0119065 * h^2)
mf <- -exp(1.6930674 + h * (0.1441647 + h * (-0.01849276 + h * (0.031074485 +
h * 0.0055717663))))
f <- exp(-0.510725 + h * (-0.1160364 + h * (-0.006702098 + h * (0.054465944 +
h * 0.0087397329))))
psf <- pnorm(-(((1 - Wf)^l - 1) / l - mf) / f)
cat("Shapiro-Francia W'", paste0(six(Wf), ","),
"V' =", paste0(six((1 - Wf) / (l * mf + 1)^(1 / l)), ","), pv(psf), "\n")
# StatsDirect's conclusion, from the smaller of the last two P values
p <- min(sw$p.value, psf)
cat(if (p < 0.05) "Sample unlikely to be from a normal distribution"
else if (p < 0.1) "Tests not quite significant but do not assume normality"
else "No non-normality detected by tests: examine plot", "\n")
# The normal plot at the end of the report: each value against its Blom normal
# score put on the scale of the data (mean + z sd, with sd taken over n), and the
# line of equality
z <- qnorm((rank(yield) - 0.375) / (n + 0.25))
expected <- mean(yield) + z * sqrt(m2)
lim <- range(expected, yield)
plot(expected, yield, xlim = lim, ylim = lim, xlab = "Normal (Penicillin)",
ylab = "Observed (Penicillin)")
abline(0, 1)