Nonparametric Linear Regression
Menu location: Analysis_Nonparametric_Nonparametric Linear Regression.
This is a distribution free method for investigating a linear relationship between two variables Y (dependent, outcome) and X (predictor, independent).
The slope b of the regression (Y=bX+a) is calculated as the median of the gradients from all possible pairwise contrasts of your data. A confidence interval based upon Kendall's t is constructed for the slope.
Nonparametric linear regression is much less sensitive to extreme observations (outliers) than is simple linear regression based upon the least squares method. If your data contain extreme observations which may be erroneous but you do not have sufficient reason to exclude them from the analysis then nonparametric linear regression may be appropriate.
Assumptions:
- The sample is random (X can be non-random provided that Ys are independent with identical conditional distributions).
- The regression of Y on X is linear (this implies an interval measurement scale for both X and Y).
This function also provides you with an approximate two sided Kendall's rank correlation test for independence between the variables.
Technical Validation
Note that the two sided confidence interval for the slope is the inversion of the two sided Kendall's test. The approximate two sided P value for Kendall's t or tb is given but the exact quantile from Kendall's distribution is used to construct the confidence interval, therefore, there may be slight disagreement between the P value and confidence interval. If there are many ties then this situation is compounded (Conover, 1999).
Example
From Conover (1999, p. 338).
Test workbook (Nonparametric worksheet: GPA, GMAT).
The following data represent test scores for 12 graduates respectively:
| GPA | GMAT |
| 4.0 | 710 |
| 4.0 | 610 |
| 3.9 | 640 |
| 3.8 | 580 |
| 3.7 | 545 |
| 3.6 | 560 |
| 3.5 | 610 |
| 3.5 | 530 |
| 3.5 | 560 |
| 3.3 | 540 |
| 3.2 | 570 |
| 3.2 | 560 |
To analyse these data in StatsDirect you must first enter them into two columns in the workbook. Alternatively, open the test workbook using the file open function of the file menu. Then select Nonparametric Linear Regression from the Nonparametric section of the analysis menu. Select the columns marked "GPA" and "GMAT" when prompted for Y and X variables respectively.
For this example:
GPA vs. GMAT
Observations per sample = 12
Median slope (95% CI) = 0.003485 (0 to 0.008)
Y-intercept = 1.581061
Kendall's rank correlation coefficient tau b = 0.439039
Two sided (on continuity corrected z) P = 0.0678 (H1: any correlation)
If you plot GPA against GMAT scores using the scatter plot function in the graphics menu, you will see that there is a reasonably straight line relationship between GPA and GMAT. Here we can infer with 95% confidence that the true population value of the slope of a linear regression line for these two variables lies between 0 and 0.008. The regression equation is estimated at Y = 1.5811 + 0.0035X.
From the two sided Kendall's rank correlation test, we can not reject the null hypothesis of mutual independence between the pairs of results for the twelve graduates. Note that the zero lower confidence interval is a marginal result and we may have rejected the null hypothesis had we used a different method for testing independence.
R code
This R code reproduces the example above. It needs no packages and was checked with R 4.6.1. Paste it into R, or save it as a script and run it.
# Nonparametric linear regression: the StatsDirect help example (Conover 1999,
# p. 338) in R
gpa <- c(4.0, 4.0, 3.9, 3.8, 3.7, 3.6, 3.5, 3.5, 3.5, 3.3, 3.2, 3.2) # outcome
gmat <- c(710, 610, 640, 580, 545, 560, 610, 530, 560, 540, 570, 560) # predictor
n <- length(gpa)
# R has no standard function for this method (Theil's regression). The slope is
# the median of the slopes between all pairs of points that differ in x, and the
# intercept makes the line pass through the two medians.
s <- outer(gpa, gpa, "-") / outer(gmat, gmat, "-")
s <- sort(s[upper.tri(s) & outer(gmat, gmat, "!=")])
N <- length(s) # 62 slopes: 4 of the 66 pairs tie in x
slope <- median(s)
cat("Observations per sample =", n, "\n")
cat(sprintf("Median slope = %.6f Y-intercept = %.6f\n", slope,
median(gpa) - slope * median(gmat)))
# 95% confidence interval for the slope (Conover 1999): with w the 0.975 quantile
# of Kendall's statistic (concordant minus discordant pairs) for n pairs, the
# limits are the rth smallest and rth largest slope, r = (N - w) / 2 rounded down.
# The null distribution of the statistic comes from the number of inversions of
# a random ordering: multiply the polynomials 1 + q + ... + q^i for i below n.
ways <- 1
for (i in 1:(n - 1)) ways <- convolve(ways, rep(1, i + 1), type = "open")
S <- n * (n - 1) / 2 - 2 * (seq_along(ways) - 1) # the statistic, largest first
upper <- cumsum(round(ways)) / factorial(n) # P(statistic >= S)
w <- max(S[upper > 0.025]) # 28: P(statistic > 28) is 0.0224
r <- floor((N - w) / 2)
cat(sprintf("95%% CI for the slope = %g to %g\n", s[r], s[N + 1 - r]))
# Kendall's rank correlation between outcome and predictor. With ties R gives
# tau b and a normal approximation; continuity = TRUE matches StatsDirect's P.
print(cor.test(gmat, gpa, method = "kendall", exact = FALSE, continuity = TRUE))