Appendix B — Code Conventions
This is the house style used by every piece of R code in this book, not a standard that R enforces or that other R programmers follow. R will happily run code written any other way, and you will meet plenty of it: other books use = for assignment, load the tidyverse, and prefer double quotes. The conventions here are collected for two reasons. The first is practical: when your code looks like the book’s code, you can compare the two line by line and see where you diverged. The second is that most of these choices exist to prevent a specific mistake, and the notes below say which one. Where a convention is a matter of taste rather than safety, it says so. The code blocks on this page are shown for reference and are not run.
B.1 Assignment and Naming
Always use <- for assignment, never =. The two are not interchangeable, because = also passes arguments to functions, and mixing the two makes it hard to see which is happening.
Use snake_case for multi-word names and _hat for generic names that hold observed values or sample summaries. Use y_fit for fitted or predicted values and e_hat for residuals. Keep capitals where the mathematical notation uses them, such as the design matrix X and the empirical CDF F_hat.
| Kind of name | Convention | Examples |
|---|---|---|
| Multi-word object | snake_case | sample_means, boot_se, col_high |
| Observed sample values | lowercase with _hat |
x_hat, y_hat, xy_hat |
| Small numerical example | label 0 before _hat |
x0_hat, xy0_hat |
| Simulated sample values | _sim |
x_sim, xy_sim |
| Sample mean, variance, standard deviation | lowercase with _hat |
m_hat, v_hat, s_hat |
| Sample median | med_hat |
med_hat <- median(x_hat) |
| Empirical CDF | F_hat |
F_hat <- ecdf(x_hat) |
| Kernel density or histogram | f_hat |
f_hat <- density(x_hat), f_hat <- hist(x_hat, plot=FALSE) |
| Fitted values and residuals | distinguish the fit from the observations | y_fit, e_hat |
| Bootstrap sample | _boot |
x_boot |
| Bootstrap means and their average | distinguish the collection from its average | boot_means, m_boot |
| Loop index or count | lowercase | i, n, b, n_boot |
| Mathematical matrix | uppercase | X (design matrix) |
| Coefficient | lowercase, with context where needed | b0, b1, b_k, b_hat |
x_hat <- USArrests[, 'Murder']
m_hat <- mean(x_hat)
v_hat <- var(x_hat)
s_hat <- sd(x_hat)
med_hat <- median(x_hat)Names that match the notation in Notation let you read a formula and the code that implements it side by side. For example, x_hat[i] corresponds to \(\hat{x}_i\), m_hat to \(\hat{m}\), and y_fit[i] to \(\hat{y}_i^{\mathrm{fit}}\). The sample median is med_hat in code and \(\tilde{m}\) in formulas. Use F_hat for an empirical cumulative distribution function and f_hat for a kernel density estimate or a stored histogram. The capital F distinguishes the cumulative distribution from the density. Add a group or variable label when needed, as in x1_hat, x2_hat, m_x_hat, and m_y_hat. For two distributions, use names such as F1_hat and F2_hat, or f_x_hat and f_y_hat. Descriptive names such as wage and region_means can stay descriptive, and existing dataset column names keep their spelling. Possible outcomes and evaluation points can use x or x_values, while known population quantities use names such as population_mean. These are different roles from the sample values in x_hat.
In the Part 1 labs, x_hat always holds the murder arrest rates USArrests[, 'Murder']. Each chapter defines it once, where the lab data first appears. Small hand-typed examples, such as c(3, 3.1, 0.02), use x0_hat, and simulated draws use x_sim, so neither overwrites the lab data. If you copy a lab chunk into a new session, run the definition of x_hat first. In Sampling & Resampling and Population Statistics, the Wages1 hourly wages play the role of a known population. They are stored in population_wages, with population_wage_mean for the population mean, and samples drawn from them use x_sim.
In the Part 2 labs, xy_hat plays the same role for bivariate data. It holds USArrests[, c('UrbanPop', 'Murder')] with columns renamed x_hat and y_hat, so a regression reads lm(y_hat ~ x_hat, data=xy_hat). Small hand-typed examples use xy0_hat, and simulated data use xy_sim with columns x_sim and y_sim. The Wages1 lab data use xy2_hat, with years of schooling in x_hat and hourly wage in y_hat.
In the Part 3 labs, xy_hat holds the multivariate data, usually USArrests with a Region column added from state.region. Here the columns keep their names, so a regression reads lm(Murder ~ Assault + UrbanPop, data=xy_hat) and its output names each variable. Bootstrap resamples of the data use xy_boot, and simulated data use xy_sim with columns such as x1_sim, x2_sim, and y_sim.
B.2 Base R
This book uses base R throughout, with no tidyverse verbs (mutate, filter, select, arrange) and no pipes. Use aggregate(), subset(), merge(), and bracket indexing instead. The one exception is %>% where a package requires chaining, as plotly does.
Access a package with :: when you use it once or twice, and with library() only when a chapter uses it repeatedly.
car::vif(reg) # used once
library('wooldridge') # used throughout a chapterWriting car::vif() rather than loading the package makes it obvious where a function came from, which matters when an assistant suggests a function and you cannot find it.
Data Access
Use bracket notation with the column name for subsetting, the dollar sign for quick single-column access inline, and double brackets for pulling an element out of a list.
xy_hat <- USArrests[, c('UrbanPop', 'Murder')]
colnames(xy_hat) <- c('x_hat', 'y_hat')
x_hat <- xy_hat[, 'x_hat']
y_hat <- xy_hat[, 'y_hat']
# also acceptable for single columns
assault_high <- USArrests$Assault > median(USArrests$Assault)Naming the column, rather than its position, means your code keeps working when the column order changes.
B.3 Functions and Comments
Put the opening brace on the same line as function(), indent the body, end with an explicit return(), and close the brace on its own line.
skewness <- function(x_hat) {
m_hat <- mean(x_hat)
m3_hat <- mean((x_hat - m_hat)^3)
s3_hat <- sd(x_hat)^3
skew <- m3_hat / s3_hat
return(skew)
}R returns the last expression evaluated even without return(), so an explicit return() is for the reader rather than the interpreter. It makes the output of a function unambiguous at a glance.
Use # with a space after it, and let comments explain purpose rather than mechanics.
x_sim <- rnorm(1000) # simulated sample values
m_hat <- mean(x_sim) # sample meanA comment that restates the code (# take the mean) adds nothing, while one that says why (# sample mean, to compare against the population mean) does. This is the style Working with AI asks you to comment in, and Step 3 there checks these comments against what the code actually does.
Spacing and Strings
Put spaces around binary operators and after commas, and no spaces inside parentheses. The one exception is = when it names a function argument, which takes no spaces. This is deliberate: x <- y + 1 is arithmetic and gets room to breathe, while col='red' is one argument and reads as a single unit.
x <- y + 1
c(1, 2, 3)
X[1, 2]
grey(0, .5)
mean(x_hat) # not mean( x_hat )
hist(x_hat, breaks=25, freq=FALSE) # not breaks = 25Use single quotes for all strings, including column names, package names, and axis labels. R treats 'a' and "a" as identical, so this one is purely for consistency.
USArrests[,'Murder']
library('wooldridge')
xlab='Murder arrests (per 100k)'Logical Values
Write logical values in full, as TRUE and FALSE.
hist(x_hat, freq=FALSE, border=NA)
x_boot <- sample(x_hat, replace=TRUE)R also accepts the abbreviations T and F, and you will see them constantly in other people’s code, but they are not safe. TRUE and FALSE are reserved words that cannot be reassigned, while T and F are ordinary variables that merely start out holding those values. Someone who writes T <- 0 earlier in a script silently breaks every freq=T after it, and nothing warns you. The full words cost three or four extra characters and remove the possibility.
B.4 Simulation and Loops
Call set.seed() once at the top of the block that generates random numbers, and never reset it partway through. Resetting midway makes the results depend on where the reset happened, which is very hard to debug later.
Use replicate() for repeated simulations that return a vector or matrix, sapply() and lapply() for apply-style iteration, and an explicit for loop when the iteration itself is the point being taught. Pre-allocate with rep(NA, n).
set.seed(1)
n_boot <- 999
boot_means <- rep(NA, n_boot)
for (b in seq(n_boot)) {
x_boot <- sample(x_hat, replace=TRUE)
boot_means[b] <- mean(x_boot)
}
m_boot <- mean(boot_means)Here boot_means[b] is one replicate mean, while m_boot averages all the replicate means. If a loop stores its current replicate mean separately, use m_b_hat. The corresponding jackknife collection and its average are jack_means and m_jack.
Pre-allocating with rep(NA, n) rather than numeric(n) matters because numeric(n) fills the vector with zeros. If a bug leaves some entries unfilled, zeros look like real results and NA does not.
seq and seq_along
These two are not interchangeable, and the book uses each for a different job.
for (b in seq(n_boot)) { ... } # n_boot is a count, so seq(n_boot) gives 1, 2, ..., n_boot
for (i in seq_along(x_hat)) { ... } # x_hat is a vector, so this gives 1, ..., length(x_hat)Writing seq_along(n_boot) when n_boot is a single number is a silent bug: it returns just 1, so the loop runs once and the code appears to work. Writing seq(x_hat) when x_hat is a vector happens to work, but it reads as though x_hat were a count.
One edge case is worth knowing. seq(n_boot) returns c(1, 0) when n_boot is zero, so a loop that should not run at all runs twice. seq_len(n_boot) returns nothing in that case and is the safer form. The book uses seq() because its counts are never zero.
Formulas and Output
Use standard R formula syntax with spaces around ~.
reg <- lm(y_hat ~ x_hat)
y_fit <- fitted(reg)
e_hat <- y_hat - y_fit
plot(y_hat ~ x_hat, pch=16, col=grey(0, .5))Round for display with round(value, 2), build label strings with paste0(), and rely on auto-printing rather than calling print() except inside a loop or conditional.
title(paste0('mean= ', round(m_hat, 2)), font.main=1)B.5 Reference Card
<-for assignment, never=.- Base R only, no tidyverse and no pipes.
library()for repeated use,pkg::fun()for one-off calls.- Bracket notation
[,'Name']for column subsetting. - Explicit
return()in every function you write. set.seed()once at the top of a random block, never reset.- Single quotes for all strings.
- Spaces around operators, none around an argument’s
=. TRUEandFALSEin full, neverTandF.- Pre-allocate with
rep(NA, n), notnumeric(n). seq()for counts,seq_along()for vectors.- Comments explain purpose, not mechanics.
x_hat,y_hat,m_hat,v_hat,s_hat, andmed_hatfor generic sample quantities;xy_hatfor the Part 2 and Part 3 lab data;x0_hatandxy0_hatfor small examples;x_simandxy_simfor simulated draws;y_fitfor fitted values ande_hatfor residuals.F_hatfor ECDFs;f_hatfor kernel density estimates and stored histograms.
Further Reading
- https://style.tidyverse.org/ – a widely used R style guide. It assumes the tidyverse rather than base R, so read it for the general principles rather than the specific verbs.
- Figures – the matching conventions for plots.