Appendix E — Packages and Datasets


The book works from a small number of datasets on purpose. Reusing USArrests across twenty chapters means that by the time you reach multiple regression you already know what the variables are, so the new material is the only new thing on the page. This appendix lists each dataset, its variables, where it comes from, and which chapters use it. Use it when a chapter refers to a variable you have forgotten, or when you want a dataset you already understand for an exercise. The last section lists every package the book uses, so you can install them all before you start.

E.1 Summary

Dataset Source Rows Variables Chapters
USArrests datasets (built in) 50 Murder, Assault, UrbanPop, Rape 2–3, 5–8, 10–11, 13–18, 20–25, 30
state.region datasets (built in) 50 a 4-level factor 3, 13, 22, 24, 30
anscombe datasets (built in) 11 x1–x4, y1–y4 15
UCBAdmissions datasets (built in) 24 cells Admit, Gender, Dept 18, 22
LifeCycleSavings datasets (built in) 50 sr, pop15, pop75, dpi, ddpi 22
Wages1 Ecdat 3294 exper, sex, school, wage 2, 5–8, 11–18, 20, 22, 26
finance-charts-apple.csv web, via read.csv 506 Date, AAPL.Open, AAPL.High, AAPL.Low, AAPL.Close, AAPL.Volume, AAPL.Adjusted, and four derived columns 27
tylervigen.csv web, via read.csv varies many unrelated time series 29

E.2 Built-in Datasets

These come with R, so they need no package and no download. Type the name to see the data, and ?USArrests for the help page.

USArrests

Violent crime rates by US state in 1973, and the book’s main working dataset.

Variable Meaning
Murder Murder arrests per 100,000 residents
Assault Assault arrests per 100,000 residents
UrbanPop Percent of the state population living in urban areas
Rape Rape arrests per 100,000 residents

Two features make it useful for teaching. The row names are state names rather than a column, which is why the book indexes with USArrests[,'Murder'] rather than by position. And the variables are rates per 100,000 rather than counts, so they are comparable across states of very different size.

Code
dim(USArrests)
## [1] 50  4
head(USArrests, 3)
##         Murder Assault UrbanPop Rape
## Alabama   13.2     236       58 21.2
## Alaska    10.0     263       48 44.5
## Arizona    8.1     294       80 31.0
summary(USArrests[,'Murder'])
##    Min. 1st Qu.  Median    Mean 3rd Qu.    Max. 
##   0.800   4.075   7.250   7.788  11.250  17.400

Note that these are arrest rates, not offense rates. Differences across states reflect policing and reporting practice as well as underlying crime, which is worth remembering whenever a regression in a later chapter treats one of these variables as an outcome.

state.region

A factor of length 50 giving the census region of each state, in the same state order as USArrests. The four levels are Northeast, South, North Central, and West. The book uses it as the example of an unordered factor and as the grouping variable for multiple-group comparisons.

Code
table(state.region)
## state.region
##     Northeast         South North Central          West 
##             9            16            12            13

Because the ordering matches, you can attach it to USArrests directly without a merge.

anscombe

Anscombe’s quartet: four \(x\) and \(y\) pairs with nearly identical means, variances, correlations, and regression lines, but very different scatterplots. Eleven rows each. The book uses it to make the point that summary statistics do not determine the shape of a relationship.

UCBAdmissions

A three-dimensional table of graduate admissions at the University of California, Berkeley, by admission outcome, gender, and department. The table has two admission outcomes, two gender categories, and six departments, for \(2\times2\times6=24\) cells. The book uses it for three-way tables in Chapter 22 and for Simpson’s paradox in Chapter 18.

Code
ftable(UCBAdmissions)
##                 Dept   A   B   C   D   E   F
## Admit    Gender                             
## Admitted Male        512 353 120 138  53  22
##          Female       89  17 202 131  94  24
## Rejected Male        313 207 205 279 138 351
##          Female       19   8 391 244 299 317

LifeCycleSavings

Savings rates and demographics for fifty countries, averaged over 1960 to 1970.

Variable Meaning
sr Aggregate personal savings as a percent of disposable income
pop15 Percent of the population under age 15
pop75 Percent of the population over age 75
dpi Real per-capita disposable income, in US dollars
ddpi Percent growth rate of dpi

The book uses it in Chapter 22 as a second dataset with several cardinal variables, since its scatterplot matrix has both tightly related pairs and a clear outlier.

E.3 Package Datasets

Wages1

A cross-section of 3294 workers, from the Ecdat package. This is the book’s main dataset once samples need to be larger than 50. Sampling & Resampling, Population Statistics, Confidence Intervals, and Hypothesis Testing pretend its 3294 workers are the whole population, so sample statistics can be compared with a known truth. Simple Regression, Inference, Association is not Causation, and Testing Theory do the same with school and wage, stored in population_xy.

Variable Meaning
exper Years of work experience
sex male or female
school Years of schooling
wage Hourly wage

Install the package once, then load the dataset when you need it.

install.packages('Ecdat')   # once ever
data(Wages1, package='Ecdat')

With \(n = 3294\) the scatterplots need transparency to be readable, which is why Figures ties the alpha channel to sample size.

E.4 Datasets Read from the Web

Two chapters read a CSV directly from a URL rather than from a package. This needs a working internet connection, and the file could change or move, which is itself a reproducibility lesson worth noticing.

finance-charts-apple.csv

Daily Apple share prices over 506 trading days, read into an object called stock in Chapter 27. Alongside Date it carries AAPL.Open, AAPL.High, AAPL.Low, AAPL.Close, AAPL.Volume, and AAPL.Adjusted, plus four columns (dn, mavg, up, direction) that someone else derived from the prices before publishing the file. The book uses it for time-series plots, where observations are not independent across rows.

stock <- read.csv('https://raw.githubusercontent.com/plotly/datasets/master/finance-charts-apple.csv')

tylervigen.csv

A collection of unrelated time series compiled to illustrate spurious correlation, read into vigen_csv in Chapter 29. Any two of its columns will tend to correlate strongly, which is exactly the point being made.

vigen_csv <- read.csv(
    'https://raw.githubusercontent.com/the-mad-statter/whysospurious/master/data-raw/tylervigen.csv'
)

E.5 Packages

R starts with a handful of packages already loaded. These include stats for lm() and loess(), graphics for plot() and hist(), grDevices for colors such as rgb(), utils for head() and read.csv(), and datasets for USArrests. Every other package must be installed once and then loaded in each session, as explained in Data & Visualization. Code Conventions explains when to load a package with library() and when to call one function with ::.

Package What the book uses it for Chapters
Ecdat The Wages1 dataset 2, 5–8, 11–18, 20, 22, 26
wooldridge The wage1 and countymurders datasets 12, 30
mvtnorm Draws from a multivariate normal distribution, with rmvnorm() 14, 18–19
plotly Interactive figures, with plot_ly() 15, 21, 23, 27
reactable Interactive tables 21
stargazer Formatted regression and summary tables 21, 28–29
psych Scatterplot matrices, with pairs.panels() 22
fixest Fixed-effects and instrumental-variable regressions, with feols() 24, 28–29
sf Vector spatial data, such as county borders 27
terra Raster spatial data, such as an elevation grid 27
units Attaching units such as kilometers to distances 27
knitr Placing images and tables on the page, in code the book hides 1, 17, 21

You can install all of them in one line.

# Run once, in the console, not in your script
install.packages(c('Ecdat', 'wooldridge', 'mvtnorm', 'plotly', 'reactable',
    'stargazer', 'psych', 'fixest', 'sf', 'terra', 'units', 'knitr'))

On Windows and macOS these install like any other package. On Linux, sf and terra also need the system libraries GDAL, GEOS, and PROJ.

Some chapters also name a package in a comment or in the text, as a pointer to a ready-made version of something the chapter builds by hand. You do not need these to run the book’s code.

Package Named for Chapter
spatstat.univar Weighted quantiles 3
VGAM The Dagum distribution 10
DescTools Cramer’s V 14
strucchange The Chow test for a structural break 26
segmented Regression with an estimated break point 26
ivreg Instrumental-variable regression 28
car Variance inflation factors, as an example of :: Appendix B

Chapter 22 draws a scatterplot matrix with psych::pairs.panels(). Restart R and run pairs.panels(USArrests) without the psych:: prefix. What happens, and what are two ways to fix it?

R reports that it cannot find the function pairs.panels, because psych is installed but not loaded. Either write psych::pairs.panels(USArrests), or run library(psych) once near the top of your script. If R instead says there is no package called psych, then it is not installed, so run install.packages('psych') in the console first.

Further Reading