Histogram (Text-based)
Menu location: Graphics_Histogram (text-based).
The text-based frequency distribution histogram is plotted horizontally across the screen with the count for each value-range (bin) displayed at the left hand side.
The variable that you select is divided into m ranges (bins, bars). The variable is then sorted and the first k values less than or equal to the upper limit of the first bin are counted as the frequency of the first bin. This counting is then repeated for each bin, and bars are plotted to represent the counts. The numbers displayed on the x axis are the middle values (mid-points) for the value range of each bin.
StatsDirect will attempt to select a "neat" number of bins and "neat" mid-point values for your data. You may over-ride this selection and set your own bin specifications when prompted.
See also the graphical histogram.
Example
From Aziz et al. (1996).
Test workbook (Graphics worksheet: SDI not conceived).
The data are sperm deformity index (SDI) values from the semen samples of 116 men in an infertility study whose partners did not conceive. The same values, with those of the men whose partners did conceive, are used in the ROC curve example.
To draw this histogram in StatsDirect open the test workbook using the file open function of the file menu. Then select Histogram (text-based) from the Graphics menu and select the column marked "SDI not conceived" when prompted for data. Keep the bins that StatsDirect suggests.
For this example:
Distribution of SDI not conceived
Counts Mid-points
2 | 190 ===
0 | 180
1 | 170 ==
14 | 160 ========================
29 | 150 ==================================================
33 | 140 =========================================================
20 | 130 ==================================
12 | 120 =====================
5 | 110 =========
/+--------+-------+--------+-------+--------+-------+--------+
0 5 10 15 20 25 30 35
Mid-points for SDI not conceived
StatsDirect chose nine bins of width 10 with mid-points 110 to 190, so the first bin holds the values above 105 up to 115 and the last those above 185 up to 195; a value equal to the upper limit of a bin is counted in that bin. The bars are drawn to the scale of the count ruler beneath them, so the end of a bar reads its count off the ruler. The distribution is roughly symmetrical about the 140 bin, with 33 of the 116 values between 135 and 145, and has a longer tail to the right: 17 values are 125 or below and only three are above 165.
R code
This R code reproduces the example above. It needs no packages and was checked with R 4.6.1. It reads the data from sdi_not_conceived.csv, the test workbook's column saved with its heading as a csv file (see the first comment in the code).
# Histogram (text-based): the StatsDirect help example (Aziz et al. 1996, the
# sperm deformity index of 116 men whose partners did not conceive; the test
# workbook's Graphics worksheet column "SDI not conceived") in R
# The data are read from sdi_not_conceived.csv: that column of the Graphics
# worksheet of the StatsDirect test workbook, saved with its heading as a csv file
# in R's working directory.
sdi <- read.csv("sdi_not_conceived.csv")[[1]]
cat("n =", length(sdi), " range", min(sdi), "to", max(sdi), "\n")
# StatsDirect chose nine bins of width 10 with "neat" mid-points 110 to 190, so
# the bin edges are 105, 115, ..., 195. A value equal to a bin's upper limit is
# counted in that bin, which is hist()'s default (right = TRUE), and a value on the
# lowest edge would go into the first bin (include.lowest = TRUE, also the default).
edges <- seq(105, 195, by = 10)
h <- hist(sdi, breaks = edges, plot = FALSE)
print(data.frame(mid_point = h$mids, count = h$counts))
# The text histogram as StatsDirect prints it: the highest bin first, its count at
# the left, then the mid-point and a bar of "=" signs. The bars are on the scale of
# the count ruler drawn beneath them, which ends at a neat value above the largest
# count (35 here, which pretty() also gives; for other data the two may differ),
# and the full width of 60 characters stands for that value
ruler <- pretty(c(0, max(h$counts)))
bar <- strrep("=", round(h$counts / max(ruler) * 60))
# a bin with a count too small for one sign is marked with a colon
bar[h$counts > 0 & bar == ""] <- ":"
cat("Distribution of SDI not conceived\n")
cat("Counts Mid-points\n")
for (i in rev(seq_along(h$mids)))
cat(formatC(h$counts[i], width = -6), "|", formatC(h$mids[i], width = 4), bar[i],
"\n")
cat("Count ruler:", ruler, "\n")
cat("Mid-points for SDI not conceived\n")
# The same bins drawn as an ordinary histogram, labelled at the bin mid-points
plot(h, main = "Distribution of SDI not conceived", ylab = "Counts",
xlab = "Mid-points for SDI not conceived", xaxt = "n")
axis(1, at = h$mids)