The book works from a small number of datasets on purpose. Reusing USArrests across twenty chapters means that by the time you reach multiple regression you already know what the variables are, so the new material is the only new thing on the page. This appendix lists each dataset, its variables, where it comes from, and which chapters use it. Use it when a chapter refers to a variable you have forgotten, or when you want a dataset you already understand for an exercise. The last section lists every package the book uses, so you can install them all before you start.
Summary
USArrests |
datasets (built in) |
50 |
Murder, Assault, UrbanPop, Rape |
2–3, 5–8, 10–11, 13–18, 20–25, 30 |
state.region |
datasets (built in) |
50 |
a 4-level factor |
3, 13, 22, 24, 30 |
anscombe |
datasets (built in) |
11 |
x1–x4, y1–y4 |
15 |
UCBAdmissions |
datasets (built in) |
24 cells |
Admit, Gender, Dept |
18, 22 |
LifeCycleSavings |
datasets (built in) |
50 |
sr, pop15, pop75, dpi, ddpi |
22 |
Wages1 |
Ecdat |
3294 |
exper, sex, school, wage |
2, 5–8, 11–18, 20, 22, 26 |
finance-charts-apple.csv |
web, via read.csv |
506 |
Date, AAPL.Open, AAPL.High, AAPL.Low, AAPL.Close, AAPL.Volume, AAPL.Adjusted, and four derived columns |
27 |
tylervigen.csv |
web, via read.csv |
varies |
many unrelated time series |
29 |
Built-in Datasets
These come with R, so they need no package and no download. Type the name to see the data, and ?USArrests for the help page.
USArrests
Violent crime rates by US state in 1973, and the book’s main working dataset.
Murder |
Murder arrests per 100,000 residents |
Assault |
Assault arrests per 100,000 residents |
UrbanPop |
Percent of the state population living in urban areas |
Rape |
Rape arrests per 100,000 residents |
Two features make it useful for teaching. The row names are state names rather than a column, which is why the book indexes with USArrests[,'Murder'] rather than by position. And the variables are rates per 100,000 rather than counts, so they are comparable across states of very different size.
Code
dim(USArrests)
## [1] 50 4
head(USArrests, 3)
## Murder Assault UrbanPop Rape
## Alabama 13.2 236 58 21.2
## Alaska 10.0 263 48 44.5
## Arizona 8.1 294 80 31.0
summary(USArrests[,'Murder'])
## Min. 1st Qu. Median Mean 3rd Qu. Max.
## 0.800 4.075 7.250 7.788 11.250 17.400
Note that these are arrest rates, not offense rates. Differences across states reflect policing and reporting practice as well as underlying crime, which is worth remembering whenever a regression in a later chapter treats one of these variables as an outcome.
state.region
A factor of length 50 giving the census region of each state, in the same state order as USArrests. The four levels are Northeast, South, North Central, and West. The book uses it as the example of an unordered factor and as the grouping variable for multiple-group comparisons.
Code
table(state.region)
## state.region
## Northeast South North Central West
## 9 16 12 13
Because the ordering matches, you can attach it to USArrests directly without a merge.
anscombe
Anscombe’s quartet: four \(x\) and \(y\) pairs with nearly identical means, variances, correlations, and regression lines, but very different scatterplots. Eleven rows each. The book uses it to make the point that summary statistics do not determine the shape of a relationship.
UCBAdmissions
A three-dimensional table of graduate admissions at the University of California, Berkeley, by admission outcome, gender, and department. The table has two admission outcomes, two gender categories, and six departments, for \(2\times2\times6=24\) cells. The book uses it for three-way tables in Chapter 22 and for Simpson’s paradox in Chapter 18.
Code
ftable(UCBAdmissions)
## Dept A B C D E F
## Admit Gender
## Admitted Male 512 353 120 138 53 22
## Female 89 17 202 131 94 24
## Rejected Male 313 207 205 279 138 351
## Female 19 8 391 244 299 317
LifeCycleSavings
Savings rates and demographics for fifty countries, averaged over 1960 to 1970.
sr |
Aggregate personal savings as a percent of disposable income |
pop15 |
Percent of the population under age 15 |
pop75 |
Percent of the population over age 75 |
dpi |
Real per-capita disposable income, in US dollars |
ddpi |
Percent growth rate of dpi |
The book uses it in Chapter 22 as a second dataset with several cardinal variables, since its scatterplot matrix has both tightly related pairs and a clear outlier.
Package Datasets
Wages1
A cross-section of 3294 workers, from the Ecdat package. This is the book’s main dataset once samples need to be larger than 50. Sampling & Resampling, Population Statistics, Confidence Intervals, and Hypothesis Testing pretend its 3294 workers are the whole population, so sample statistics can be compared with a known truth. Simple Regression, Inference, Association is not Causation, and Testing Theory do the same with school and wage, stored in population_xy.
exper |
Years of work experience |
sex |
male or female |
school |
Years of schooling |
wage |
Hourly wage |
Install the package once, then load the dataset when you need it.
install.packages('Ecdat') # once ever
data(Wages1, package='Ecdat')
With \(n = 3294\) the scatterplots need transparency to be readable, which is why Figures ties the alpha channel to sample size.
Datasets Read from the Web
Two chapters read a CSV directly from a URL rather than from a package. This needs a working internet connection, and the file could change or move, which is itself a reproducibility lesson worth noticing.
finance-charts-apple.csv
Daily Apple share prices over 506 trading days, read into an object called stock in Chapter 27. Alongside Date it carries AAPL.Open, AAPL.High, AAPL.Low, AAPL.Close, AAPL.Volume, and AAPL.Adjusted, plus four columns (dn, mavg, up, direction) that someone else derived from the prices before publishing the file. The book uses it for time-series plots, where observations are not independent across rows.
stock <- read.csv('https://raw.githubusercontent.com/plotly/datasets/master/finance-charts-apple.csv')
tylervigen.csv
A collection of unrelated time series compiled to illustrate spurious correlation, read into vigen_csv in Chapter 29. Any two of its columns will tend to correlate strongly, which is exactly the point being made.
vigen_csv <- read.csv(
'https://raw.githubusercontent.com/the-mad-statter/whysospurious/master/data-raw/tylervigen.csv'
)
Packages
R starts with a handful of packages already loaded. These include stats for lm() and loess(), graphics for plot() and hist(), grDevices for colors such as rgb(), utils for head() and read.csv(), and datasets for USArrests. Every other package must be installed once and then loaded in each session, as explained in Data & Visualization. Code Conventions explains when to load a package with library() and when to call one function with ::.
Ecdat |
The Wages1 dataset |
2, 5–8, 11–18, 20, 22, 26 |
wooldridge |
The wage1 and countymurders datasets |
12, 30 |
mvtnorm |
Draws from a multivariate normal distribution, with rmvnorm() |
14, 18–19 |
plotly |
Interactive figures, with plot_ly() |
15, 21, 23, 27 |
reactable |
Interactive tables |
21 |
stargazer |
Formatted regression and summary tables |
21, 28–29 |
psych |
Scatterplot matrices, with pairs.panels() |
22 |
fixest |
Fixed-effects and instrumental-variable regressions, with feols() |
24, 28–29 |
sf |
Vector spatial data, such as county borders |
27 |
terra |
Raster spatial data, such as an elevation grid |
27 |
units |
Attaching units such as kilometers to distances |
27 |
knitr |
Placing images and tables on the page, in code the book hides |
1, 17, 21 |
You can install all of them in one line.
# Run once, in the console, not in your script
install.packages(c('Ecdat', 'wooldridge', 'mvtnorm', 'plotly', 'reactable',
'stargazer', 'psych', 'fixest', 'sf', 'terra', 'units', 'knitr'))
On Windows and macOS these install like any other package. On Linux, sf and terra also need the system libraries GDAL, GEOS, and PROJ.
Some chapters also name a package in a comment or in the text, as a pointer to a ready-made version of something the chapter builds by hand. You do not need these to run the book’s code.
spatstat.univar |
Weighted quantiles |
3 |
VGAM |
The Dagum distribution |
10 |
DescTools |
Cramer’s V |
14 |
strucchange |
The Chow test for a structural break |
26 |
segmented |
Regression with an estimated break point |
26 |
ivreg |
Instrumental-variable regression |
28 |
car |
Variance inflation factors, as an example of :: |
Appendix B |
Chapter 22 draws a scatterplot matrix with psych::pairs.panels(). Restart R and run pairs.panels(USArrests) without the psych:: prefix. What happens, and what are two ways to fix it?
R reports that it cannot find the function pairs.panels, because psych is installed but not loaded. Either write psych::pairs.panels(USArrests), or run library(psych) once near the top of your script. If R instead says there is no package called psych, then it is not installed, so run install.packages('psych') in the console first.