Wilcoxon's Signed Ranks Test
Menu location: Analysis_Nonparametric_Wilcoxon Signed Ranks.
This is a method for the comparison of a pair of samples.
The Wilcoxon signed ranks test statistic T+ is the sum of the ranks of the positive, non-zero differences (Di) between a pair of samples. Zero differences are set aside before the other differences are ranked, as in Conover (1999). Some other software ranks the zero differences with the rest before setting them aside (Pratt, 1959), which can give a different P value when some differences are zero.
Assumptions of tests on T+:
- distribution of each Di is symmetrical
- all Di are mutually independent
- all Di have the same mean
- measurement scale of Di is at least interval
In most situations you should use a two sided test. A two sided test is based upon the null hypothesis that the common median of the differences is zero. The approximate alternative hypothesis in this case is that the differences tend not to be zero. For a lower side test the approximate alternative hypothesis is that differences tend to be less than zero. For an upper side test the approximate alternative hypothesis is that differences tend to be greater than zero.
A confidence interval is constructed for the difference between the population medians. In sample terms this is called the confidence interval for the median or mean difference. It is also known as the Hodges-Lehmann estimate of shift. The assumptions of this method are:
- distribution of each Di is symmetrical
- all Di are mutually independent
- all Di have the same median
- measurement scale of Di is at least interval
Technical Validation
Exact permutational probability associated with the test statistic is calculated for up to 200 non-zero differences (pairs that do not differ are set aside before ranking). A normal approximation is used with more than 200 non-zero differences. Note that StatsDirect uses more accurate methods than some other statistical software for calculating probabilities associated with this statistic, therefore, you may notice a difference in results, especially where there are tied observations - you should find that StatsDirect and StatXact give the same answers. Confidence limits are calculated using the exact critical value k from the signed rank distribution for samples of 4 to 199 observations, or by calculating the normal approximation K* for samples of 200 or more observations; the exact confidence level attained is evaluated for samples of up to 1000 observations (Conover, 1999; Neumann, 1988).
Example
From Conover (1999).
Test workbook (Nonparametric worksheet: First Born, Second Born).
The following data represent aggressivity scores for 12 pairs of monozygotic twins.
| First Born | Second Born |
| 86 | 88 |
| 71 | 77 |
| 77 | 76 |
| 68 | 64 |
| 91 | 96 |
| 72 | 72 |
| 77 | 65 |
| 91 | 90 |
| 70 | 65 |
| 71 | 80 |
| 88 | 81 |
| 87 | 72 |
To analyse these data in StatsDirect you must first enter them into two columns in the workbook. Alternatively, open the test workbook using the file open function of the file menu. Then select the Wilcoxon Signed Ranks from the Nonparametric methods section of the analysis menu. Select the columns marked "First Born" and "Second Born" when prompted for data.
For this example:
Wilcoxon's signed ranks test
First Born vs. Second Born
Number of non-zero differences ranked = 11
Sum of ranks for positive differences = 41.5
Exact probability (adjusted for ties):
Lower side P = 0.7739 (H1: differences tend to be less than zero)
Upper side P = 0.2378 (H1: differences tend to be greater than zero)
Two sided P = 0.4756 (H1: differences tend not to be zero)
95.8% confidence interval for difference between population medians:
K = 14 Median difference = 1.5 (-2.5 to 6.5)
Assuming that the paired differences come from a symmetrical distribution then these results show that one group did not tend to yield different results to the other group which was paired with it, i.e. there was no statistically significant difference between the aggressivity scores of the first born as compared with the second twin. The extent of this lack of difference is shown well by the confidence interval which clearly encompasses zero. Note that the quoted 95.8% confidence interval is as close as you can get to 95% because of the mathematics involved in nonparametric methods like this.
R code
This R code reproduces the example above. It needs no packages and was checked with R 4.6.1. Paste it into R, or save it as a script and run it.
# Wilcoxon signed ranks test: the StatsDirect help example (Conover 1999) in R
first <- c(86, 71, 77, 68, 91, 72, 77, 91, 70, 71, 88, 87)
second <- c(88, 77, 76, 64, 96, 72, 65, 90, 65, 80, 81, 72)
d <- first - second # one difference is 0 and some are tied
n <- length(d)
# R's standard test. Its V is the sum of ranks for positive differences. R 4.6.0
# or later gives exact results that allow for the zero and the ties, but it ranks
# the zero with the other differences before setting it aside (the method of
# Pratt, 1959): V = 48.5 and two sided P = 0.458. StatsDirect, like Conover, sets
# zeros aside before ranking: the sum is 41.5 and two sided P = 0.4756. Each P is
# exact for its own method. Older versions of R warn and use a normal
# approximation with continuity correction, as exact = FALSE does (V = 41.5,
# two sided P = 0.4765).
print(wilcox.test(first, second, paired = TRUE, conf.int = TRUE))
# With zeros set aside first, R 4.6.0 or later gives StatsDirect's two sided P.
# Use alternative = "less" for the lower side P and "greater" for the upper.
# Keep the confidence interval from the first call: it needs every difference.
nz <- d[d != 0]
print(wilcox.test(nz))
# Number of non-zero differences and the sum of ranks for positive differences
r <- rank(abs(nz)) # tied values share a mid-rank
Tplus <- sum(r[nz > 0])
cat("Number of non-zero differences ranked =", length(nz), "\n")
cat("Sum of ranks for positive differences =", Tplus, "\n")
# Exact P given the ties, in any version of R: each rank is as likely to belong
# to a positive as to a negative difference, so build the distribution of the
# sum one rank at a time. Ranks are doubled to make mid-ranks whole numbers.
p <- 1 # P(doubled sum = 0), before any rank
for (v in 2 * r) p <- (c(p, numeric(v)) + c(numeric(v), p)) / 2
w <- 2 * Tplus + 1 # position of the observed value
lower <- sum(p[1:w]) # each side includes the observed sum
upper <- sum(p[w:length(p)])
cat("Lower side P =", round(lower, 4), " Upper side P =", round(upper, 4),
" Two sided P =", round(min(1, 2 * min(lower, upper)), 4), "\n")
# Confidence interval, from all n differences: K from the exact distribution of
# the statistic without ties, then the Kth smallest and Kth largest of the
# n(n + 1)/2 averages of two differences (each is also averaged with itself).
# With fewer than 6 differences K is 0 and there is no 95% interval.
# StatsDirect's median difference is the median of these averages, which R calls
# the (pseudo)median; it is not median(d), which is 1 here.
# R 4.6.0 or later gives the same limits for these data but works out the level
# given the ties (96.2 percent in R 4.6.1). For other tied data its limits can
# differ.
K <- qsignrank(0.025, n)
a <- outer(d, d, "+") / 2
a <- sort(a[upper.tri(a, diag = TRUE)])
cat(sprintf("%.1f%% confidence interval for difference between population medians:\n",
100 * (1 - 2 * psignrank(K - 1, n))))
cat(sprintf("K = %g Median difference = %g (%g to %g)\n",
K, median(a), a[K], rev(a)[K]))