3  Summary Statistics


We often summarize distributions with statistics: functions of data. We will go through some specific examples mathematically below, but can intuitively understand that they are generally computed as statistic <- function(x){ .... }. This chapter covers measures of center (mean, median), spread (variance, IQR, MAD), and shape (skewness, kurtosis) for cardinal data. It then covers the statistics that remain meaningful for factor data: the mode, the Herfindahl index, and quantiles of ordered levels.

Note that functions can take functions as arguments, meaning we can also program statistics generally as

Code
statistic <- function(x_hat, f){
    stat_hat <- f(x_hat)
    return(stat_hat)
}

x0_hat <- c(0, 1, 3, 10, 6) # Data
statistic(x0_hat, sum)
## [1] 20

The most basic way to compute statistics is with summary, which reports multiple values that can all be calculated individually.

Code
# A random sample (real data)
x_hat <- USArrests[, 'Murder'] # lab data: murder arrests per 100,000
summary(x_hat)
##    Min. 1st Qu.  Median    Mean 3rd Qu.    Max. 
##   0.800   4.075   7.250   7.788  11.250  17.400

# A random sample (computer simulation)
x1_sim <- runif(1000)
summary(x1_sim)
##     Min.  1st Qu.   Median     Mean  3rd Qu.     Max. 
## 0.001003 0.248650 0.504267 0.506181 0.764271 0.999970

# Another random sample (computer simulation)
x2_sim <- rnorm(1000)
summary(x2_sim)
##     Min.  1st Qu.   Median     Mean  3rd Qu.     Max. 
## -3.08461 -0.68394 -0.05505 -0.05304  0.60773  3.11937

Together, the sample mean and variance statistics summarize the central tendency and dispersion of a distribution. In some special cases, such as with the normal distribution, they completely describe the distribution. Other distributions are better described with other statistics, either as an alternative or in addition to the mean and variance. After discussing those other statistics, we will return to the two most basic statistics in theoretical detail.

3.1 Mean and Variance

The mean and variance are the two most basic statistics that summarize the center and how spread apart the values are for data in your sample. As before, we represent data as vector \(\hat{x}=(\hat{x}_{1}, \hat{x}_{2}, ....\hat{x}_{n})\), where there are \(n\) observations and \(\hat{x}_{i}\) is the value of the \(i\)th one.

Mean

The first step in describing a sample is to reduce its center to a single number.

ImportantKey Definition

The sample mean is the sum of the observations divided by the number of observations, \[\hat{m} = \frac{1}{n} \sum_{i=1}^{n} \hat{x}_{i}.\]

The mean (also called the sample mean or empirical mean) is the simplest and most commonly reported single-number summary of a sample, and the natural measure of central tendency when the data are roughly symmetric. The mean is useful when totals matter: total income across a population equals the mean income times the number of people, a relation that does not hold for the median or other rank-based summaries. Because every observation contributes equally to the sum, a single extreme value can pull the mean substantially in its direction.

For example, a dataset of \(\{1,4,10\}\) has a mean of \([1+4+10]/3=5\).

Code
x0_hat <- c(1, 4, 10)
sum(x0_hat)/length(x0_hat)
## [1] 5
mean(x0_hat)
## [1] 5
Code
# compute the mean of a random sample
m_hat <- mean(x_hat)
m_hat
## [1] 7.788

# visualize on a histogram
hist(x_hat, border=NA, main=NA, freq=FALSE)
abline(v=m_hat, col=rgb(1, 0, 0, .8), lwd=2)
title(paste0('mean= ', round(m_hat, 2)), font.main=1)

Variance

Once we know where the data are centered, the next question is how widely they vary around that center.

ImportantKey Definition

The sample variance is the average squared deviation from the mean, \[\hat{v} = \frac{1}{n} \sum_{i=1}^{n} (\hat{x}_{i} - \hat{m})^2.\] The sample standard deviation is its square root, \(\hat{s} = \sqrt{\hat{v}}\), and is in the same units as the data.

The variance and standard deviation are the second-most-common single-number summaries of a sample. They tell us how concentrated or spread out the data are around the mean: a small variance means the observations cluster tightly, a large variance means they fan out. The standard deviation in particular is useful because it sits in the original units of the data, making sentences like “monthly returns averaged \(0.6\%\) with a standard deviation of \(4\%\)” interpretable at a glance. Squaring the deviations also means a few extreme values can dominate the result, which is one reason the median absolute deviation (below) is often preferred for heavy-tailed data.

For example, a dataset of \(\{1,4,10\}\) has a mean of \([1+4+10]/3=5\). The variance is \([(1-5)^2+(4-5)^2+(10-5)^2]/3=[16+1+25]/3=42/3=14\) and the standard deviation is \(\sqrt{14}\).

Code
x0_hat <- c(1, 4, 10)
m_hat <- mean(x0_hat)
v_hat <- mean( (x0_hat - m_hat)^2 )
sqrt(v_hat)
## [1] 3.741657
Code
s1_hat <- sd(x_hat) # sqrt(var(x_hat))
hist(x_hat, border=NA, main=NA, freq=FALSE)
m1_sd_bounds <- c(m_hat - s1_hat,  m_hat + s1_hat)
abline(v=m_hat, col=rgb(1, 0, 0, .8), lwd=2) #mean
abline(v=m1_sd_bounds, col=rgb(0, 0, 1, .8)) # +- sd
text(m1_sd_bounds, -.02,
    c( expression(bar(X)-s[X]), expression(bar(X)+s[X])),
    col=rgb(0, 0, 1, .8), adj=0)
title(paste0('mean ± 1 sd= ', round(s1_hat, 2)), font.main=1)

Note that a “unbiased version” of the empirical variance is used by R and many statisticians: \(\hat{v}' =\frac{\sum_{i=1}^{n} [\hat{x}_{i} - \hat{m}]^2}{n-1}\) and \(\hat{s}' = \sqrt{\hat{v}'}\). In this class, we use the version defined previously because we do not yet know about “bias/unbiased” statistics. Do not be concerned, as there is hardly any difference when \(n\) is large: e.g., \(\frac{1}{n}\approx \frac{1}{n-1}\) for \(n=100,000\).

Code
x0_hat <- c(1, 4, 10)

m_hat <- mean(x0_hat)
v_hat <- mean( (x0_hat - m_hat)^2 )
v_hat
## [1] 14

# Corrected Version
n <- length(x0_hat)
v_corrected_hat <- sum( (x0_hat - m_hat)^2 )/(n-1)
v_corrected_hat
## [1] 21

var(x0_hat) # R-Version
## [1] 21

3.2 Other Center/Spread Statistics

A general rule of applied statistics is that there are multiple ways to measure something. Mean and Variance are measurements of Center and Spread, but there are others that have different theoretical properties and may be better suited for your dataset.

Medians and Absolute Deviations

When the data are skewed or contain outliers, statistics built from quantiles often give a more representative summary of center and spread than the mean and variance.

ImportantKey Definition

The sample median \(\tilde{m}\) is the value where one half of the data fall below and one half above; it is the \(0.5\) quantile.

The interquartile range (\(IQR\)) is the upper quartile minus the lower quartile, the spread of the middle half of the data.

The median absolute deviation (\(\text{MAD}\)) is the median of the absolute deviations from the median, \(\text{Med}(|\hat{x}_{i} - \tilde{m}|)\).

These three statistics share a common rationale: by using ranks rather than arithmetic averages, they resist the influence of extreme values. A single outlier can pull the mean by an arbitrary amount, but moves the median by at most one position; the \(IQR\) and \(\text{MAD}\) inherit the same protection because they are built from medians and quantiles (recall the \(p\)-th quantile from the previous chapter). Reach for these robust summaries when the distribution is skewed or contains values you cannot trust to be typical: income, house prices, response times, and similar heavy-tailed quantities where a small fraction of large values would swamp the mean.

Code
median(x_hat)
## [1] 7.25
quantile(x_hat, prob=0.5)
##  50% 
## 7.25

Examine robustness to an extreme value

Code
x1_extreme_hat <- c(x_hat, 1000) # add one extreme value
#par(mfrow=c(1, 2)) # visualize side-by-side
#hist(x_hat)
#hist(x1_extreme_hat)

# Which measures of central tendency are robust
# to a single extreme value?
mean( x_hat)
## [1] 7.788
mean( x1_extreme_hat )
## [1] 27.24314

quantile(x_hat, prob=0.5)
##  50% 
## 7.25
quantile(x1_extreme_hat, prob=0.5)
## 50% 
## 7.3

The \(IQR\) corresponds to the box in the boxplot, and the \(\text{MAD}\) uses the same robust idea but measures distances from the median directly rather than reading off quartile cut-offs.

Code
IQR(x_hat)
## [1] 7.175
mad(x_hat)
## [1] 5.41149

Compute the \(IQR\) statistic for the dataset \(\{-100,1,4,10,10\}\).

\(IQR =\) Upper quartile \(-\) Lower quartile \(= 10 - 1 = 9\).

Code
x0_hat <- c(-100, 1, 4, 10, 10)
# An intuitive alternative to sd(x0_hat), used in the boxplot
quants <- quantile(x0_hat, probs=c(.25, .75))
quants[2]-quants[1]
## 75% 
##   9
IQR(x0_hat)
## [1] 9

Compute the \(\text{MAD}\) statistic for the dataset \(\{1,4,10\}\). First compute the median, \(\text{Med}(-100,1,4,10,10)=4\). Then compute \(\text{Med}(,~ |99-4|, |1-4|,~ |4-4|,~ |10-4|,~ |10-4| )= \text{Med}(95, 3, 0, 6, 6 ) = 6\).

Code
#Another alternative to sd(x0_hat)
mad(x0_hat, constant=1)
## [1] 6

# Computationally equivalent
# med_hat <- quantile(x0_hat, probs=.5)
# absdev_hat <- abs(x0_hat -med_hat)
# quantile(absdev_hat, probs=.5)

Compare the robustness of various “spread” metrics to an extreme value

Code
sd(x_hat)
## [1] 4.35551
sd(x1_extreme_hat)
## [1] 139.0044

IQR(x_hat)
## [1] 7.175
IQR(x1_extreme_hat)
## [1] 7.2

mad(x_hat, constant=1)
## [1] 3.65
mad(x1_extreme_hat, constant=1)
## [1] 3.8

Note that there other “absolute deviation” statistics

Code
# sometimes seen elsewhere
mean( abs(x_hat - mean(x_hat)) )
mean( abs(x_hat - median(x_hat)) )

Weighted Statistics

The mean generalizes to a weighted mean: an average where different values contribute to the final result with varying levels of importance. For each outcome \(x\) we have a weight \(W_{x}\) and compute \[\begin{aligned} \hat{m} &= \frac{\sum_{x} x W_{x}}{\sum_{x} W_{x}} = \sum_{x} x w_{x}, \end{aligned}\] where \(w_{x}=\frac{W_{x}}{\sum_{x'} W_{x'}}\) is normalized version of \(W_{x}\) that implies \(\sum_{x}w_{x}=1\).

For another example, suppose a student has these scores

Code
Homework1 <- c(score=88, weight=0.25)
Homework2 <- c(score=92, weight=0.25)
Exam1 <- c(score=67, weight=0.2)
Exam2 <- c(score=90, weight=0.3)

Grades <- rbind(Homework1, Homework2, Exam1, Exam2)
Grades
##           score weight
## Homework1    88   0.25
## Homework2    92   0.25
## Exam1        67   0.20
## Exam2        90   0.30

We can compute the final grade as a weighted mean

Code
# Manual Way
88*0.25 + 92*0.25 + 67*0.2 + 90*0.3
## [1] 85.4

# Computerized Way
Values <- Grades[, 'score'] * Grades[, 'weight']
FinalGrade <- sum(Values)
FinalGrade
## [1] 85.4

When data are discrete, we can also compute the mean using “probability weights” \(w_{x} = \hat{p}(x)=\sum_{i=1}^{n}\mathbf{1}\left(\hat{x}_{i}=x\right)/n\).

E.g., the dataset \(\{1,2,1,3,1\}\) has \(\hat{p}(1)=\frac{3}{5}\), \(\hat{p}(2)=\frac{1}{5}\), \(\hat{p}(3)=\frac{1}{5}\). Then, sorting the dataset as {1,1,1,2,3}, we can see that \(\hat{m} = [1+1+1+2+3]/5 = [1+1+1]/5 + 2/5 + 3/5 = 1 \frac{3}{5} + 2 \frac{1}{5} + 3 \frac{1}{5} = 1 \hat{p}(1) + 2 \hat{p}(2) + 3 \hat{p}(3)\). In either case, we end up with \(\frac{3+2+3}{5}=8/5=1.6\)

Code
x0_hat <- c(1, 2, 1, 3, 1)
mean(x0_hat)
## [1] 1.6

proportions <- table(x0_hat)/length(x0_hat)
proportions
## x0_hat
##   1   2   3 
## 0.6 0.2 0.2
vals <- sort(unique(x0_hat))
sum(vals*proportions)
## [1] 1.6

In principle, we can also compute other weighted statistics such as a weighted median.

See that we can also compute weighted quantiles

Code
weighted.quantile <- function(x_hat, w, probs){
    #See spatstat.univar::weighted.quantile
    oo <- order(x_hat)
    x_hat <- x_hat[oo]
    w <- w[oo]
    F_hat <- cumsum(w)/sum(w)
    quantile_id <- max(which(F_hat <= probs))+1
    q_hat <- x_hat[quantile_id]
    return(q_hat)
}

## Unweighted
quantile(x0_hat, probs=.5)
## 50% 
##   1
weights <- rep(1, length(x0_hat))
weighted.quantile(x_hat=x0_hat, w=weights, probs=.5)
## [1] 1

## Weighted
weights <- seq(x0_hat)
weights <- weights/sum(weights) # normalize
weighted.quantile(x_hat=x0_hat, w=weights, probs=.5)
## [1] 1

3.3 Shape Statistics

Central tendency and dispersion are often insufficient to describe a distribution. To further describe shape, we can compute sample skew and kurtosis to measure asymmetry and extreme values.

There are many other statistics we could compute on an ad-hoc basis. However, shape is often best understood with graphical descriptions: histogram, ECDF, Boxplot. These should be made before numerical descriptions: skewness and kurtosis statistics.

Skewness

Center and spread alone leave open whether the distribution is symmetric; the first shape question asks whether one tail extends further than the other.

ImportantKey Definition

The sample skewness measures asymmetry: the average cubed deviation from the mean, scaled by the cube of the standard deviation, \[\text{skew} = \frac{\sum_{i=1}^{n}(\hat{x}_{i}-\hat{m})^3 / n}{\hat{s}^3}.\] Positive skew means a long right tail, negative skew a long left tail, and zero skew (e.g., the Normal) is symmetric.

Cubing the deviations preserves their sign: positive deviations stay positive and negative deviations stay negative. The numerator is therefore close to zero when the data are roughly balanced around the mean, and large in absolute value when one tail is much longer than the other. Dividing by the cubed standard deviation makes skewness dimensionless, so the values are comparable across datasets measured in different units. Skewness is useful for flagging whether the mean and median will agree closely (low skew) or diverge sharply (high skew), and for deciding whether a log or square-root transformation might tame the distribution before further analysis.

Consider the right-skewed dataset \(\{0,1,2,3,9\}\) with mean \(\hat{m}=15/5=3\). The cubed deviations from the mean are \((-3)^3, (-2)^3, (-1)^3, 0^3, 6^3\), which equal \(-27, -8, -1, 0, 216\). A cube keeps the sign of its deviation, so the single large positive deviation outweighs the small negative ones. Their average is the numerator: \(\sum_{i=1}^{n}[\hat{x}_{i}-\hat{m}]^3/n = 180/5 = 36\). Dividing by \(\hat{s}^3 \approx 44.2\) gives a skewness of \(36/44.2 \approx 0.81\). The positive value reflects the long right tail created by the value \(9\).

Code
x_demo_hat <- c(0, 1, 2, 3, 9)
m3_hat <- mean( (x_demo_hat - mean(x_demo_hat))^3 ) # numerator
m3_hat
## [1] 36
m3_hat / sd(x_demo_hat)^3 # skewness
## [1] 0.814587
Code
hist( x_hat^2, border=NA,
    xlab='[Murder Rate]^2',
    main=NA, freq=FALSE, breaks=20)

Code

skewness <-  function(x_hat){
 m_hat <- mean(x_hat)
 m3_hat <- mean((x_hat - m_hat)^3)
 s3_hat <- sd(x_hat)^3
 skew <- m3_hat/s3_hat
 return(skew)
}

skewness( x_hat^2 )
## [1] 1.078254
skewness( x_hat^3 )
## [1] 1.658012

We can automatically compare against the normal distribution, which has a skew of 0.

Kurtosis

A second shape question, independent of asymmetry, asks how much of the variability comes from a few far-from-center observations rather than from the typical bulk.

ImportantKey Definition

The sample kurtosis measures how heavy the tails are: the average fourth power of deviations from the mean, scaled by \(\hat{s}^4\), \[\text{kurt} = \frac{\sum_{i=1}^{n}(\hat{x}_{i}-\hat{m})^4 / n}{\hat{s}^4}.\] The Normal distribution has kurtosis \(3\), so subtracting \(3\) gives a comparison-to-Normal version (excess kurtosis).

Where skewness uses the cube (and so can be negative), kurtosis uses the fourth power, which forces every term to be positive and disproportionately amplifies the largest deviations. A handful of far-from-the-mean values can drive kurtosis up substantially even when the bulk of the distribution looks unremarkable, which is why it is often described as a “tail-heaviness” measure. Kurtosis is useful for distinguishing distributions that look similar in their bulk but differ in how often extreme values occur: daily stock returns, for example, have moderate variance but kurtosis well above \(3\), reflecting occasional crashes that a Normal model would treat as vanishingly rare.

Using the same dataset \(\{0,1,2,3,9\}\) with mean \(\hat{m}=3\), the quartic deviations are \((-3)^4, (-2)^4, (-1)^4, 0^4, 6^4\), which equal \(81, 16, 1, 0, 1296\). Because the power is even, every term is positive, and the large deviation from the value \(9\) dominates the sum. Their average is the numerator: \(\sum_{i=1}^{n}[\hat{x}_{i}-\hat{m}]^4/n = 1394/5 = 278.8\). Dividing by \(\hat{s}^4 = 156.25\) gives a kurtosis of \(278.8/156.25 \approx 1.78\).

Code
x_demo_hat <- c(0, 1, 2, 3, 9)
m4_hat <- mean( (x_demo_hat - mean(x_demo_hat))^4 ) # numerator
m4_hat
## [1] 278.8
m4_hat / sd(x_demo_hat)^4 # kurtosis
## [1] 1.78432

Some authors further subtract \(3\) to explicitly compare against the normal distribution (the normal distribution has a kurtosis of \(3\)).

Boxplot whiskers are a great way to examine kurtosis, with more circle “outliers” indicating more kurtosis. You can also see skew in the boxplot when one quartile is much further from the median than the other.

Code
x1_hat <- x_hat^1/mean(x_hat^1)
x2_hat <- x_hat^2/mean(x_hat^2)
x3_hat <- x_hat^3/mean(x_hat^3)
x4_hat <- x_hat^4/mean(x_hat^4)
boxplot(x1_hat, x2_hat, x3_hat, x4_hat,
    names=c(1, 2, 3, 4), xlab='Data Transformations',
    main=NA)

Code


kurtosis <- function(x_hat){
 m_hat <- mean(x_hat)
 m4_hat <- mean((x_hat - m_hat)^4)
 s4_hat <- sd(x_hat)^4
 kurt <- m4_hat/s4_hat
 return(kurt)
 # use instead to compare against normal
 # excess_kurt <- kurt - 3 
}

kurtosis( x1_hat )
## [1] 2.05077
kurtosis( x2_hat )
## [1] 3.247877
kurtosis( x3_hat )
## [1] 5.166225
kurtosis( x4_hat )
## [1] 7.511867

Clusters/Gaps

You can also describe distributions in terms of how clustered the values are, including the number of modes (peaks), bunching, and many other statistics. A gap is a range with few or no observations, and a cluster is a group of values bunched together. A distribution with one peak is unimodal, and one with two peaks is bimodal. The two histograms below are each bimodal: two clusters of values separated by a gap. The mean and variance alone would miss this structure, since two very different distributions can share the same mean and variance. So remember that “a picture is worth a thousand words”: plot your data before summarizing it numerically.

3.4 Factor Data

The statistics above all add, subtract, or rank the observations, so they require cardinal data. For factor data, we instead start from the proportions \(\hat{p}(x)\) from the empirical mass function and summarize those. Which summaries make sense depends on whether the factor is ordered.

Ordered Factors

Ordered factors, such as letter grades or survey responses from strongly disagree to strongly agree, have a ranking but no meaningful distance between levels. The mode and Herfindahl index apply as before, since they only use the proportions. Because the levels can be sorted, the median and other quantiles also apply: the median is the middle level, and the quartiles bracket the middle half of the observations. Quantiles are useful for ordered factors because they only use the ordering: “half the class earned a \(B\) or better” is a statement the data can support. The mean and variance do not apply, because there is nothing to add. A grade point average gets around this by coding the levels as numbers, treating the gap between \(C\) and \(B\) as equal to the gap between \(B\) and \(A\), but that is an assumption imposed by the grader rather than a property of the data. Surveys often record a cardinal variable only in bands, such as income brackets, which turns it into an ordered factor.

For example, the letter grades \(\{B, A, C, B, F, B, A\}\) of \(n=7\) students sort from lowest to highest as \(F, C, B, B, B, A, A\). The median is the fourth sorted value, \(B\), the lower quartile is the second, \(C\), and the upper quartile is the sixth, \(A\). The mode is also \(B\).

Code
grade <- factor(c('B', 'A', 'C', 'B', 'F', 'B', 'A'),
    levels=c('F', 'D', 'C', 'B', 'A'), ordered=TRUE)
sort(grade)
## [1] F C B B B A A
## Levels: F < D < C < B < A

# median and quartiles (type=1 is required for ordered factors)
quantile(grade, probs=c(.25, .5, .75), type=1)
## 25% 50% 75% 
##   C   B   A 
## Levels: F < D < C < B < A

Compute the median of the original USArrests[, 'UrbanPop'] and check which band it falls in. Then move the inner breaks from \(50\) and \(75\) to \(40\) and \(60\) and recompute the median band. Does the median band always contain the cardinal median?

Below, the urban share of each state’s population is recorded in three bands.

Code
# urban population in bands, an ordered factor
urban_band <- cut(USArrests[, 'UrbanPop'], breaks=c(0, 50, 75, 100),
    labels=c('mostly rural', 'mixed', 'mostly urban'), ordered_result=TRUE)
p_hat <- table(urban_band)/length(urban_band)
plot(p_hat, col=grey(0, .5),
    xlab='Urban Population', ylab='Proportion of States')

Code

# mode and Herfindahl index
names(p_hat)[p_hat==max(p_hat)]
## [1] "mixed"
sum(p_hat^2)
## [1] 0.4024

# median and quartiles
quantile(urban_band, probs=c(.25, .5, .75), type=1)
##          25%          50%          75% 
##        mixed        mixed mostly urban 
## Levels: mostly rural < mixed < mostly urban

Unordered Factors

For categories with no natural order, such as regions or brands, the center of a distribution is simply the most common category. Categories have no distance between them, so for factor data spread instead means how evenly the observations divide across the categories.

ImportantKey Definition

The sample mode is the value that occurs most often: the \(x\) with the largest proportion \(\hat{p}(x)\).

The Herfindahl index is the sum of the squared proportions across the distinct values, \(\hat{H} = \sum_{x} \hat{p}(x)^2.\)

The mode is useful when the question is “which category is typical?”: the most common region, the best-selling brand, the most frequent answer on a survey. It is the only measure of center that applies to unordered factors, since the categories can be neither averaged nor ranked. If several categories tie for the highest proportion, the data have more than one mode.

The Herfindahl index ranges from \(1/K\), when every category has the same proportion, up to \(1\), when every observation falls in a single category. It is useful for asking how dominated a distribution is by its largest categories. Economists use the same statistic, with market shares in place of proportions, to measure how concentrated an industry is, and antitrust agencies report it with shares in percentage points so that the scale runs to \(10{,}000\). The complement \(1-\hat{H}\) is sometimes reported instead so that larger values mean a more even spread, and after rescaling by \(K/(K-1)\) it is called the index of qualitative variation.

For example, the dataset \(\{A, B, A, C, C, A\}\) has \(n=6\) observations in \(K=3\) categories. \(A\) appears three times, \(C\) twice, and \(B\) once, so \(\hat{p}(A)=1/2\), \(\hat{p}(B)=1/6\), \(\hat{p}(C)=1/3\), and the mode is \(A\). Notice the mode is not the alphabetically middle letter, \(B\). The Herfindahl index is \((1/2)^2+(1/6)^2+(1/3)^2=[9+1+4]/36=14/36\approx 0.39\), above the minimum of \(1/3\) that would occur if all three letters were equally common.

Code
x0_hat <- c('A', 'B', 'A', 'C', 'C', 'A')
p_hat <- table(x0_hat)/length(x0_hat)
p_hat
## x0_hat
##         A         B         C 
## 0.5000000 0.1666667 0.3333333

# mode(s)
names(p_hat)[p_hat==max(p_hat)]
## [1] "A"

# Herfindahl index
sum(p_hat^2)
## [1] 0.3888889
Code
# census region of each state, an unordered factor
region <- state.region
p_hat <- table(region)/length(region)
plot(p_hat, col=grey(0, .5),
    xlab='Census Region', ylab='Proportion of States')

Code

# mode
names(p_hat)[p_hat==max(p_hat)]
## [1] "South"

# Herfindahl index, compared to its minimum 1/K
K <- length(p_hat)
c(H_hat=sum(p_hat^2), H_min=1/K)
## H_hat H_min 
##  0.26  0.25

Suppose \(100\) customers each buy from one of four firms: \(60\) from firm \(A\), \(20\) from \(B\), \(10\) from \(C\), and \(10\) from \(D\). Compute each firm’s market share, the mode, and the Herfindahl index. Then suppose firms \(C\) and \(D\) merge into a single firm and recompute. Does the merger raise or lower concentration?

Code
x0_hat <- rep(c('A', 'B', 'C', 'D'), times=c(60, 20, 10, 10))
p_hat <- table(x0_hat)/length(x0_hat)
sum(p_hat^2)
## [1] 0.42

# after the merger, C and D are one firm
x_merged_hat <- x0_hat
x_merged_hat[x_merged_hat=='D'] <- 'C'

Each statistic in this chapter suits some data types and not others. Here is a table to help you recall which apply. Note that the mode and Herfindahl index need repeated values, so for continuous cardinal data we can describe proportions with a histogram instead (seeing the highest peak as the mode, for example).

Statistic Factor Cardinal
Mean, variance, skewness, kurtosis No Yes
Median, quantiles, \(IQR\) Ordered only (not Unordered) Yes
Mode, Herfindahl index Yes Discrete only (not Continuous)*

3.5 Exercises

  1. Comment the script you wrote for this chapter, then restart R and check that it runs from a clean session, then check the script with AI as explained in Working with AI. Write three sentences from memory on the main statistical idea of this chapter, and ask the assistant what is wrong, vague, or missing. Finish with your own questions about whatever you found hardest.

  2. Explain in your own words why the median is more robust to extreme values than the mean. Give a concrete example with a small dataset where adding one outlier changes the mean substantially but barely moves the median.

  3. Using the dataset \(\{2, 5, 5, 8, 12\}\), compute by hand: (a) the sample mean \(\hat{m}\), (b) the sample variance \(\hat{v}\), (c) the sample standard deviation \(\hat{s}\), and (d) the \(IQR\). Show your work for each step.

  4. Load USArrests in R. Compute the mean, standard deviation, and MAD (with constant=1) of the Assault variable. Then apply a log transformation with log(USArrests[,'Assault']) and compute the skewness of both the original and transformed data using the skewness function defined in the chapter. Which version is less skewed?

  5. A survey asks \(12\) customers to rate a product on a five-level scale from poor to excellent, with responses \(\{good, fair, excellent, good, poor, good, fair, excellent, good, very~good, fair, good\}\). Store the responses as an ordered factor in R, then report the proportions, the mode, the median, and the Herfindahl index. Explain why the median is a meaningful summary here but the mean of the numeric codes \(1\) to \(5\) requires an extra assumption.

Further Reading

Recall

This chapter compressed a distribution into single numbers: the mean and variance (and standard deviation) for center and spread, the median, IQR, and MAD as robust alternatives, and the skewness and kurtosis for shape. For factor data, where nothing can be added, the mode and the Herfindahl index \(\hat{H}\) took over: the four census regions in state.region have mode South and \(\hat{H}=0.26\), barely above the \(0.25\) floor for four equally common categories, and ordered factors such as letter grades also keep the median and quartiles. The running five-point dataset \(\{0, 1, 2, 3, 9\}\) tied the shape statistics together: its mean is \(3\), its cubed deviations average to \(36\) giving a skewness of about \(0.81\) (long right tail), and its quartic deviations give a kurtosis of about \(1.78\) (both driven by the single outlier at \(9\)). On real data such as USArrests[, 'Murder'], a histogram plus the mean and standard deviation already reveal most of the shape; the median, MAD, skew, and kurtosis fill in what symmetry assumptions hide.