Spearman's Rank Correlation
Menu location: Analysis_Nonparametric_Spearman Rank Correlation.
Spearman's rank correlation provides a distribution free test of independence between two variables. It is, however, insensitive to some types of dependence. Kendall's rank correlation gives a better measure of correlation and is also a better two sided test for independence.
Spearman's rank correlation coefficient (r or rho) is calculated as:
- where R(x) and R(y) are the ranks of a pair of variables (x and y) each containing n observations.
Technical Validation
Rho is calculated as Pearson's r based on ranks and average ranks using the above formula. The probability associated with r, when there are no tied ranks present, is evaluated using an exact permutational method when n ≤ 10 and an Edgeworth series approximation when n > 10 (Best and Roberts, 1975). The exact probability calculation employs a corrected version of the Best and Roberts (1975) algorithm, which was further improved in version 3 of StatsDirect. When ties are present the standard asymptotic approximation [2_tailed_t(df=n-2,t=|rho|sqrt(n-2)/sqrt(1-rho^2)]) is used. A confidence interval for rho is constructed using Fisher's z transformation (Conover, 1999; Gardner and Altman, 1989; Hollander and Wolfe, 1973). Note that StatsDirect uses more accurate definitions of rho and the probabilities associated with it than some other statistical software, therefore, there may be differences in results: for example, using the data below in version 3.2.2 of R will give different P values from the corr.test function compared with the spearman.test function of the pspearman package because the mistake in AS89 has been carried over from the original FORTRAN.
Example
From Armitage and Berry (1994, p. 466).
Test workbook (Nonparametric worksheet: Career, Psychology).
The following data represent a tutor's ranking of ten clinical psychology students as to their suitability for their career and their knowledge of psychology:
| Career | Psychology |
| 4 | 5 |
| 10 | 8 |
| 3 | 6 |
| 1 | 2 |
| 9 | 10 |
| 2 | 3 |
| 6 | 9 |
| 7 | 4 |
| 8 | 7 |
| 5 | 1 |
To analyse these data in StatsDirect you must first enter them into two columns in the workbook. Alternatively, open the test workbook using the file open function of the file menu. Then select Spearman Rank Correlation from the Nonparametric section of the analysis menu. Select the columns marked "Career" and "Psychology" when prompted for data.
For this example:
Spearman's rank correlation
Career vs. Psychology
Observations per sample = 10
Spearman's score = 52
Spearman's rank correlation coefficient (Rho) = 0.684848
95% CI for rho (Fisher's Z transformed) = 0.097085 to 0.918443
Upper side P = 0.0173 (H1: positive correlation)
Lower side P = 0.9847 (H1: negative correlation)
Two sided P = 0.0347 (H1: any correlation)
From these results we reject the null hypothesis of mutual independence between the tutor's ranking of students suitability for their career and their knowledge of psychology. With a two sided test we are considering the possibility of a positive or a negative correlation, i.e. we can't be sure of this direction at the outset. A one sided test would have been restricted to correlation in one direction only i.e. large values of one group associated with big values of the other (positive correlation) or large values of one group associated with small values of the other (negative correlation). In our example we can conclude that there is a statistically significant lack of independence between career suitability and psychology knowledge rankings of the students by the tutor. The tutor tended to rank students with apparently greater knowledge as more suitable to their career than those with apparently less knowledge and vice versa.
R code
This R code reproduces the example above. It needs no packages and was checked with R 4.6.1. Paste it into R, or save it as a script and run it.
# Spearman's rank correlation: the StatsDirect help example (Armitage and Berry
# 1994, p. 466) in R
career <- c(4, 10, 3, 1, 9, 2, 6, 7, 8, 5)
psychology <- c(5, 8, 6, 2, 10, 3, 9, 4, 7, 1)
n <- length(career)
# R's standard test. Its S is the sum of squared differences between the ranks,
# which StatsDirect calls Spearman's score (52). With 10 or more pairs R takes
# P from a series approximation (two sided P = 0.0351 here); with fewer it is
# exact. StatsDirect's P is exact for 10 or fewer pairs (0.0347).
# Use alternative = "greater" for the upper side P and "less" for the lower.
print(cor.test(career, psychology, method = "spearman"))
# Rho, the score and a 95% confidence interval by Fisher's z transformation
rx <- rank(career)
ry <- rank(psychology)
rho <- cor(rx, ry)
score <- sum((rx - ry)^2)
ci <- tanh(atanh(rho) + c(-1, 1) * qnorm(0.975) / sqrt(n - 3))
cat("Spearman's score =", score, "\n")
cat(sprintf("Rho = %.6f 95%% CI = %.6f to %.6f\n", rho, ci[1], ci[2]))
# Exact P, for data without ties and up to 12 pairs: count, for every way of
# pairing the ranks, the sum of squared differences. Pairs are added one at a
# time; a row of 'ways' is the set of ranks of y already used and a column is
# the sum so far.
stopifnot(!anyDuplicated(career), !anyDuplicated(psychology), n <= 12)
maxsum <- n * (n^2 - 1) / 3
ways <- matrix(0, 2^n, maxsum + 1)
ways[1, 1] <- 1
for (set in 0:(2^n - 2)) {
used <- bitwAnd(set, 2^(0:(n - 1))) > 0
i <- sum(used) + 1 # the next rank of x to be paired
for (j in which(!used)) {
s <- (i - j)^2
to <- set + 2^(j - 1) + 1
ways[to, (s + 1):(maxsum + 1)] <- ways[to, (s + 1):(maxsum + 1)] +
ways[set + 1, 1:(maxsum + 1 - s)]
}
}
p <- ways[2^n, ] / factorial(n) # P(score = 0, 1, ... maxsum)
upper <- sum(p[1:(score + 1)]) # a small score is a large rho
lower <- sum(p[(score + 1):(maxsum + 1)]) # each side includes the observed score
cat("Upper side P =", round(upper, 4), " Lower side P =", round(lower, 4),
" Two sided P =", round(min(1, 2 * min(upper, lower)), 4), "\n")