Preference Group Allocation
Menu location: Analysis_Randomization_Preference Allocation.
This function allocates subjects to groups according to their preferences. A uniform random allocation procedure is used to select subjects for inclusion in groups which are over-subscribed. The procedure is best explained by example:
Suppose ten students were asked to apply for a choice of five courses, which have two, one, two, three and five places respectively. The students are asked to list their top three course preferences in order.
StatsDirect can allocate the students to a course based on their preferences and on a uniform random selection procedure for over-subscribed courses. There is no weighting procedure for any round of selections as this would encourage tactical preference choice, i.e. the probability that a student is allocated his/her first preference is not influenced by the subscription rates for his/her other preferences.
Say our ten students mark the following preferences for courses 1 to 5:
| Student | Preference 1 | Preference 2 | Preference 3 |
| 1 | 3 | 5 | 1 |
| 2 | 3 | 4 | 2 |
| 3 | 5 | 1 | 3 |
| 4 | 3 | 1 | 4 |
| 5 | 5 | 3 | 1 |
| 6 | 5 | 4 | 3 |
| 7 | 2 | 3 | 5 |
| 8 | 4 | 1 | 5 |
| 9 | 3 | 5 | 1 |
| 10 | 1 | 3 | 5 |
The maximum capacity of each group is as follows:
| Group | Capacity/Places |
| 1 | 2 |
| 2 | 1 |
| 3 | 2 |
| 4 | 3 |
| 5 | 5 |
To use StatsDirect to allocate the students to groups you must first enter the above columns of data into a workbook (they are in the Other worksheet of the test workbook as Group capacity, 1st choice, 2nd choice and 3rd choice). Then select Preference Allocation from the Randomization section of the Analysis menu. Enter 10 as the seed for the random number generator. When asked for the group capacity column, select it: in this column the rows represent the allocation groups to which the preference data refer (i.e. if the entry in row 3 of the capacity column was 5 this would mean that group 3 can hold a maximum of 5 subjects). When asked for columns of preferences you must select the columns in the correct order, i.e. preference 1, 2, 3. For this example the random allocation procedure with seed 10 gives the results below. The seed offered by default is taken from the computer's timer, so if you accept it you are likely to get different results each time; enter a seed of your own to make an allocation repeatable.
The output is:
Random allocation to groups by preference
Groups = 5
Total group capacity = 13
Subjects = 10
Randomized with seed: 10
| Subject | Group |
| 1 | 5 |
| 2 | 3 |
| 3 | 5 |
| 4 | 3 |
| 5 | 5 |
| 6 | 5 |
| 7 | 2 |
| 8 | 4 |
| 9 | 5 |
| 10 | 1 |
R code
This R code reproduces the procedure of the example above, using R's own random number generator, which differs from StatsDirect's, so the same seed gives a different allocation. It needs no packages and was checked with R 4.6.1. Paste it into R, or save it as a script and run it.
# Preference group allocation: the StatsDirect help example (ten students ranking
# five courses; the test workbook's Other worksheet columns Group capacity and 1st,
# 2nd and 3rd choice) in R
capacity <- c(2, 1, 2, 3, 5) # places in groups 1 to 5
choice <- cbind(c(3, 3, 5, 3, 5, 5, 2, 4, 3, 1), # each student's 1st choice
c(5, 4, 1, 1, 3, 4, 3, 1, 5, 3), # 2nd choice
c(1, 2, 3, 4, 1, 3, 5, 5, 1, 5)) # 3rd choice
groups <- length(capacity)
subjects <- nrow(choice)
stopifnot(sum(capacity) >= subjects, choice >= 1, choice <= groups)
# R has no function for this, so the procedure is written out. Round by round (1st
# choices, then 2nd, then 3rd) each group with places left takes the unallocated
# subjects who chose it in that round; when more want it than there are places, a
# random sample of them gets in, so each has the same chance. Subjects still without
# a group after the last round are placed at random in the places left over.
# set.seed() makes an allocation repeatable. R's random number generator is not
# StatsDirect's (both are Mersenne twisters, but seeded and used differently), so the
# same seed gives a different allocation in R from the one in the help.
set.seed(10)
group <- rep(NA_integer_, subjects)
for (round in seq_len(ncol(choice))) {
for (g in seq_len(groups)) {
space <- capacity[g] - sum(group == g, na.rm = TRUE)
wanting <- which(is.na(group) & choice[, round] == g)
if (space > 0 && length(wanting) > 0) {
taken <- if (length(wanting) > space) sample(wanting, space) else wanting
group[taken] <- g
}
}
}
left <- which(is.na(group))
if (length(left) > 0) {
spare <- rep(seq_len(groups), capacity - tabulate(group, groups)) # one entry a place
group[left] <- spare[sample.int(length(spare), length(left))] # even if one is left
}
# The report: the allocation, and (not in the report) which choice each student got
cat("Random allocation to groups by preference\n")
cat("Groups =", groups, "\n")
cat("Total group capacity =", sum(capacity), "\n")
cat("Subjects =", subjects, "\n")
cat("Randomized with seed: 10\n")
got <- apply(choice == group, 1, function(hit) if (any(hit)) which(hit)[1] else NA)
print(data.frame(Subject = seq_len(subjects), Group = group, Choice = got),
row.names = FALSE)
cat("Places used:", paste(tabulate(group, groups), "of", capacity, collapse = ", "),
"\n")