Mann-Whitney U Test
Menu location: Analysis_Nonparametric_Mann-Whitney.
This is a method for the comparison of two independent random samples (x and y):
The Mann Whitney U statistic is defined as:
- where samples of size n1 and n2 are pooled and Ri are the ranks.
U can be resolved as the number of times observations in one sample precede observations in the other sample in the ranking.
Wilcoxon rank sum, Kendall's S and the Mann-Whitney U test are exactly equivalent tests. In the presence of ties the Mann-Whitney test is also equivalent to a chi-square test for trend.
In most circumstances a two sided test is required; here the alternative hypothesis is that x values tend to be distributed differently to y values. For a lower side test the alternative hypothesis is that x values tend to be smaller than y values. For an upper side test the alternative hypothesis is that x values tend to be larger than y values.
Assumptions of the Mann-Whitney test:
- random samples from populations
- independence within samples and mutual independence between samples
- measurement scale is at least ordinal
A confidence interval for the difference between two measures of location is provided with the sample medians. The assumptions of this method are slightly different from the assumptions of the Mann-Whitney test:
- random samples from populations
- independence within samples and mutual independence between samples
- two population distribution functions are identical apart from a possible difference in location parameters
The theta statistic [U'/(n1*n2), where U' = n1*n2 - U] is provided as an additional effect size reflecting the distance between the two underlying frequency distributions from which the samples are drawn (Newcombe, 2006a). It is equivalent to the area under the receiver operating characteristic (ROC) curve.
Technical Validation
StatsDirect uses the sampling distribution of U to give exact probabilities. Each one sided probability includes the observed value of U, so the two add up to more than one; the two sided probability is twice the smaller of them. These calculations may take an appreciable time to complete when many data are tied.
Confidence intervals are constructed for the difference between the means or medians (any measure of location in fact). The level of confidence used will be as close as is theoretically possible to the one you specify. StatsDirect approaches the selected confidence level from the conservative side via the Hodges-Lehman estimator (Monahan, 1984).
When samples are large a normal approximation is used. For the hypothesis test this happens when both samples exceed 100 observations, or when the exact calculation would need more than a million workspace elements (with ties this can happen below 100 observations per sample, e.g. 91 vs 91, and depends mainly on the size of the first sample). For the confidence interval the exact distribution of U is used unless the product of the two sample sizes is about 200,000 or more (e.g. 447 vs 447), in which case an approximate K is used and is labelled as such. Note that StatsDirect uses more accurate P value calculations than some other statistical software, therefore, you may notice a difference in results (Conover, 1999; Dinneen and Blakesley, 1973; Harding, 1984; Neumann, 1988).
A confidence interval for theta is constructed using Newcombe's fifth method (Newcombe, 2006b). Note that the confidence interval for theta is directly comparable with the P value for the Mann-Whitney test whereas the confidence interval for the difference between the two measures of location (medians) is not.
Example
From Conover (1999, p. 218).
Test workbook (Nonparametric worksheet: Farm Boys, Town Boys).
The following data represent fitness scores from two groups of boys of the same age, those from homes in the town and those from farm homes.
| Farm Boys | Town Boys |
| 14.8 | 12.7 |
| 7.3 | 14.2 |
| 5.6 | 12.6 |
| 6.3 | 2.1 |
| 9.0 | 17.7 |
| 4.2 | 11.8 |
| 10.6 | 16.9 |
| 12.5 | 7.9 |
| 12.9 | 16.0 |
| 16.1 | 10.6 |
| 11.4 | 5.6 |
| 2.7 | 5.6 |
| 7.6 | |
| 11.3 | |
| 8.3 | |
| 6.7 | |
| 3.6 | |
| 1.0 | |
| 2.4 | |
| 6.4 | |
| 9.1 | |
| 6.7 | |
| 18.6 | |
| 3.2 | |
| 6.2 | |
| 6.1 | |
| 15.3 | |
| 10.6 | |
| 1.8 | |
| 5.9 | |
| 9.9 | |
| 10.6 | |
| 14.8 | |
| 5.0 | |
| 2.6 | |
| 4.0 |
To analyse these data in StatsDirect you must first enter them in two separate workbook columns. Alternatively, open the test workbook using the file open function of the file menu. Then select the Mann-Whitney from the Nonparametric section of the analysis menu. Select the columns marked "Farm Boys" and "Town Boys" when prompted for data.
For this example:
Mann-Whitney U test
Observations (x) in Farm Boys = 12 median = 9.8 rank sum = 321
Observations (y) in Town Boys = 36 median = 7.75
U = 243 U' = 189
Exact probability (adjusted for ties):
Lower side P = 0.7394 (H1: x tends to be less than y)
Upper side P = 0.2645 (H1: x tends to be greater than y)
Two sided P = 0.529 (H1: x tends to be distributed differently to y)
Theta (U'/mn) = 0.4375 (95% CI: 0.271772 to 0.621223)
95.14% confidence interval for difference between medians or means:
K = 134 Median difference = 0.8 (CI: -2.3 to 4.4)
Here we have assumed that these groups are independent and that they represent at least hypothetical random samples of the sub-populations they represent. In this analysis, we are clearly unable to reject the null hypothesis that one group does NOT tend to yield different fitness scores to the other. This lack of statistical evidence of a difference is reflected in the confidence interval for the difference between population means, in that the interval spans zero. Note that the quoted 95.14% confidence interval is as close as you can get to 95% because of the very nature of the mathematics involved in nonparametric methods like this.
R code
This R code reproduces the example above. It needs no packages and was checked with R 4.6.1. Paste it into R, or save it as a script and run it.
# Mann-Whitney U test: the StatsDirect help example (Conover 1999, p. 218) in R
farm <- c(14.8, 7.3, 5.6, 6.3, 9.0, 4.2, 10.6, 12.5, 12.9, 16.1, 11.4, 2.7)
town <- c(12.7, 14.2, 12.6, 2.1, 17.7, 11.8, 16.9, 7.9, 16.0, 10.6, 5.6, 5.6,
7.6, 11.3, 8.3, 6.7, 3.6, 1.0, 2.4, 6.4, 9.1, 6.7, 18.6, 3.2,
6.2, 6.1, 15.3, 10.6, 1.8, 5.9, 9.9, 10.6, 14.8, 5.0, 2.6, 4.0)
n1 <- length(farm)
n2 <- length(town)
# R's standard test. Its W is the U for x (243). These data have ties. R 4.6.0
# or later gives exact results that allow for the ties. Older versions warn and
# use a normal approximation with continuity correction, as exact = FALSE does
# (two sided P = 0.5279).
# Use alternative = "less" for the lower side P and "greater" for the upper.
print(wilcox.test(farm, town, conf.int = TRUE))
# Medians, rank sum, U, U' and theta
r <- rank(c(farm, town)) # tied values share a mid-rank
ranksum <- sum(r[1:n1])
U <- ranksum - n1 * (n1 + 1) / 2
Uprime <- n1 * n2 - U
theta <- Uprime / (n1 * n2)
cat("x: n =", n1, " median =", median(farm), " rank sum =", ranksum, "\n")
cat("y: n =", n2, " median =", median(town), "\n")
cat("U =", U, " U' =", Uprime, "\n")
# Theta: 95% confidence interval by method 5 of Newcombe (2006b). Each search
# stops just short of the end that would solve the equation if theta is 0 or 1.
z <- qnorm(0.975)
h <- (n1 + n2) / 2 - 1
e <- 1e-9
f <- function(y, s) y + s * z * sqrt(y * (1 - y) *
(1 + h * ((1 - y) / (2 - y) + y / (1 + y))) / (n1 * n2)) - theta
ci <- c(uniroot(f, c(0, 1 - e), s = 1, tol = 1e-10)$root, # lower limit
uniroot(f, c(e, 1), s = -1, tol = 1e-10)$root) # upper limit
cat(sprintf("Theta = %.6f (95%% CI: %.6f to %.6f)\n", theta, ci[1], ci[2]))
# Exact P given the ties, in any version of R: count the ways of choosing n1
# of the ranks, by their sum. Ranks are doubled to make mid-ranks whole numbers.
r2 <- 2 * r
S <- sum(r2)
ways <- matrix(0, n1 + 1, S + 1) # rows: how many ranks, columns: sum
ways[1, 1] <- 1
for (v in r2) ways[-1, (v + 1):(S + 1)] <-
ways[-1, (v + 1):(S + 1)] + ways[-(n1 + 1), 1:(S + 1 - v)]
p <- ways[n1 + 1, ] / choose(n1 + n2, n1) # P(doubled rank sum = 0, 1, ... S)
w <- 2 * ranksum + 1 # position of the observed value
lower <- sum(p[1:w]) # each side includes the observed U
upper <- sum(p[w:(S + 1)])
cat("Lower side P =", round(lower, 4), " Upper side P =", round(upper, 4),
" Two sided P =", round(min(1, 2 * min(lower, upper)), 4), "\n")
# Confidence interval for the difference: K from the exact distribution of U
# without ties, then the Kth smallest and Kth largest of the n1 * n2 differences.
# For these data R 4.6.0 or later gives the same limits but labels them 95.2
# percent, as it allows for the ties. For other tied data its limits can differ.
K <- qwilcox(0.025, n1, n2)
d <- sort(outer(farm, town, "-"))
cat(sprintf("%.2f%% confidence interval for the difference: K = %d\n",
100 * (1 - 2 * pwilcox(K - 1, n1, n2)), K))
cat(sprintf("median difference = %g CI = %g to %g\n",
median(d), d[K], rev(d)[K]))