Appendix B — Code Conventions


This is the house style used by every piece of R code in this book, not a standard that R enforces or that other R programmers follow. R will happily run code written any other way, and you will meet plenty of it: other books use = for assignment, load the tidyverse, and prefer double quotes. The conventions here are collected for two reasons. The first is practical: when your code looks like the book’s code, you can compare the two line by line and see where you diverged. The second is that most of these choices exist to prevent a specific mistake, and the notes below say which one. Where a convention is a matter of taste rather than safety, it says so. The code blocks on this page are shown for reference and are not run.

B.1 Assignment and Naming

Always use <- for assignment, never =. The two are not interchangeable, because = also passes arguments to functions, and mixing the two makes it hard to see which is happening.

Use snake_case for multi-word names and _hat for generic names that hold observed values or sample summaries. Use y_fit for fitted or predicted values and e_hat for residuals. Keep capitals where the mathematical notation uses them, such as the design matrix X and the empirical CDF F_hat.

Kind of name Convention Examples
Multi-word object snake_case sample_means, boot_se, col_high
Observed sample values lowercase with _hat x_hat, y_hat, xy_hat
Small numerical example label 0 before _hat x0_hat, xy0_hat
Simulated sample values _sim x_sim, xy_sim
Sample mean, variance, standard deviation lowercase with _hat m_hat, v_hat, s_hat
Sample median med_hat med_hat <- median(x_hat)
Empirical CDF F_hat F_hat <- ecdf(x_hat)
Kernel density or histogram f_hat f_hat <- density(x_hat), f_hat <- hist(x_hat, plot=FALSE)
Fitted values and residuals distinguish the fit from the observations y_fit, e_hat
Bootstrap sample _boot x_boot
Bootstrap means and their average distinguish the collection from its average boot_means, m_boot
Loop index or count lowercase i, n, b, n_boot
Mathematical matrix uppercase X (design matrix)
Coefficient lowercase, with context where needed b0, b1, b_k, b_hat
x_hat <- USArrests[, 'Murder']
m_hat <- mean(x_hat)
v_hat <- var(x_hat)
s_hat <- sd(x_hat)
med_hat <- median(x_hat)

Names that match the notation in Notation let you read a formula and the code that implements it side by side. For example, x_hat[i] corresponds to \(\hat{x}_i\), m_hat to \(\hat{m}\), and y_fit[i] to \(\hat{y}_i^{\mathrm{fit}}\). The sample median is med_hat in code and \(\tilde{m}\) in formulas. Use F_hat for an empirical cumulative distribution function and f_hat for a kernel density estimate or a stored histogram. The capital F distinguishes the cumulative distribution from the density. Add a group or variable label when needed, as in x1_hat, x2_hat, m_x_hat, and m_y_hat. For two distributions, use names such as F1_hat and F2_hat, or f_x_hat and f_y_hat. Descriptive names such as wage and region_means can stay descriptive, and existing dataset column names keep their spelling. Possible outcomes and evaluation points can use x or x_values, while known population quantities use names such as population_mean. These are different roles from the sample values in x_hat.

In the Part 1 labs, x_hat always holds the murder arrest rates USArrests[, 'Murder']. Each chapter defines it once, where the lab data first appears. Small hand-typed examples, such as c(3, 3.1, 0.02), use x0_hat, and simulated draws use x_sim, so neither overwrites the lab data. If you copy a lab chunk into a new session, run the definition of x_hat first. In Sampling & Resampling and Population Statistics, the Wages1 hourly wages play the role of a known population. They are stored in population_wages, with population_wage_mean for the population mean, and samples drawn from them use x_sim.

In the Part 2 labs, xy_hat plays the same role for bivariate data. It holds USArrests[, c('UrbanPop', 'Murder')] with columns renamed x_hat and y_hat, so a regression reads lm(y_hat ~ x_hat, data=xy_hat). Small hand-typed examples use xy0_hat, and simulated data use xy_sim with columns x_sim and y_sim. The Wages1 lab data use xy2_hat, with years of schooling in x_hat and hourly wage in y_hat.

In the Part 3 labs, xy_hat holds the multivariate data, usually USArrests with a Region column added from state.region. Here the columns keep their names, so a regression reads lm(Murder ~ Assault + UrbanPop, data=xy_hat) and its output names each variable. Bootstrap resamples of the data use xy_boot, and simulated data use xy_sim with columns such as x1_sim, x2_sim, and y_sim.

B.2 Base R

This book uses base R throughout, with no tidyverse verbs (mutate, filter, select, arrange) and no pipes. Use aggregate(), subset(), merge(), and bracket indexing instead. The one exception is %>% where a package requires chaining, as plotly does.

Access a package with :: when you use it once or twice, and with library() only when a chapter uses it repeatedly.

car::vif(reg)      # used once
library('wooldridge')  # used throughout a chapter

Writing car::vif() rather than loading the package makes it obvious where a function came from, which matters when an assistant suggests a function and you cannot find it.

Data Access

Use bracket notation with the column name for subsetting, the dollar sign for quick single-column access inline, and double brackets for pulling an element out of a list.

xy_hat <- USArrests[, c('UrbanPop', 'Murder')]
colnames(xy_hat) <- c('x_hat', 'y_hat')
x_hat <- xy_hat[, 'x_hat']
y_hat <- xy_hat[, 'y_hat']

# also acceptable for single columns
assault_high <- USArrests$Assault > median(USArrests$Assault)

Naming the column, rather than its position, means your code keeps working when the column order changes.

B.3 Functions and Comments

Put the opening brace on the same line as function(), indent the body, end with an explicit return(), and close the brace on its own line.

skewness <- function(x_hat) {
    m_hat <- mean(x_hat)
    m3_hat <- mean((x_hat - m_hat)^3)
    s3_hat <- sd(x_hat)^3
    skew <- m3_hat / s3_hat
    return(skew)
}

R returns the last expression evaluated even without return(), so an explicit return() is for the reader rather than the interpreter. It makes the output of a function unambiguous at a glance.

Use # with a space after it, and let comments explain purpose rather than mechanics.

x_sim <- rnorm(1000)        # simulated sample values
m_hat <- mean(x_sim)        # sample mean

A comment that restates the code (# take the mean) adds nothing, while one that says why (# sample mean, to compare against the population mean) does. This is the style Working with AI asks you to comment in, and Step 3 there checks these comments against what the code actually does.

Spacing and Strings

Put spaces around binary operators and after commas, and no spaces inside parentheses. The one exception is = when it names a function argument, which takes no spaces. This is deliberate: x <- y + 1 is arithmetic and gets room to breathe, while col='red' is one argument and reads as a single unit.

x <- y + 1
c(1, 2, 3)
X[1, 2]
grey(0, .5)
mean(x_hat)                        # not mean( x_hat )
hist(x_hat, breaks=25, freq=FALSE)     # not breaks = 25

Use single quotes for all strings, including column names, package names, and axis labels. R treats 'a' and "a" as identical, so this one is purely for consistency.

USArrests[,'Murder']
library('wooldridge')
xlab='Murder arrests (per 100k)'

Logical Values

Write logical values in full, as TRUE and FALSE.

hist(x_hat, freq=FALSE, border=NA)
x_boot <- sample(x_hat, replace=TRUE)

R also accepts the abbreviations T and F, and you will see them constantly in other people’s code, but they are not safe. TRUE and FALSE are reserved words that cannot be reassigned, while T and F are ordinary variables that merely start out holding those values. Someone who writes T <- 0 earlier in a script silently breaks every freq=T after it, and nothing warns you. The full words cost three or four extra characters and remove the possibility.

B.4 Simulation and Loops

Call set.seed() once at the top of the block that generates random numbers, and never reset it partway through. Resetting midway makes the results depend on where the reset happened, which is very hard to debug later.

Use replicate() for repeated simulations that return a vector or matrix, sapply() and lapply() for apply-style iteration, and an explicit for loop when the iteration itself is the point being taught. Pre-allocate with rep(NA, n).

set.seed(1)
n_boot <- 999
boot_means <- rep(NA, n_boot)
for (b in seq(n_boot)) {
    x_boot <- sample(x_hat, replace=TRUE)
    boot_means[b] <- mean(x_boot)
}
m_boot <- mean(boot_means)

Here boot_means[b] is one replicate mean, while m_boot averages all the replicate means. If a loop stores its current replicate mean separately, use m_b_hat. The corresponding jackknife collection and its average are jack_means and m_jack.

Pre-allocating with rep(NA, n) rather than numeric(n) matters because numeric(n) fills the vector with zeros. If a bug leaves some entries unfilled, zeros look like real results and NA does not.

seq and seq_along

These two are not interchangeable, and the book uses each for a different job.

for (b in seq(n_boot))        { ... }   # n_boot is a count, so seq(n_boot) gives 1, 2, ..., n_boot
for (i in seq_along(x_hat))   { ... }   # x_hat is a vector, so this gives 1, ..., length(x_hat)

Writing seq_along(n_boot) when n_boot is a single number is a silent bug: it returns just 1, so the loop runs once and the code appears to work. Writing seq(x_hat) when x_hat is a vector happens to work, but it reads as though x_hat were a count.

One edge case is worth knowing. seq(n_boot) returns c(1, 0) when n_boot is zero, so a loop that should not run at all runs twice. seq_len(n_boot) returns nothing in that case and is the safer form. The book uses seq() because its counts are never zero.

Formulas and Output

Use standard R formula syntax with spaces around ~.

reg <- lm(y_hat ~ x_hat)
y_fit <- fitted(reg)
e_hat <- y_hat - y_fit
plot(y_hat ~ x_hat, pch=16, col=grey(0, .5))

Round for display with round(value, 2), build label strings with paste0(), and rely on auto-printing rather than calling print() except inside a loop or conditional.

title(paste0('mean= ', round(m_hat, 2)), font.main=1)

B.5 Reference Card

  1. <- for assignment, never =.
  2. Base R only, no tidyverse and no pipes.
  3. library() for repeated use, pkg::fun() for one-off calls.
  4. Bracket notation [,'Name'] for column subsetting.
  5. Explicit return() in every function you write.
  6. set.seed() once at the top of a random block, never reset.
  7. Single quotes for all strings.
  8. Spaces around operators, none around an argument’s =.
  9. TRUE and FALSE in full, never T and F.
  10. Pre-allocate with rep(NA, n), not numeric(n).
  11. seq() for counts, seq_along() for vectors.
  12. Comments explain purpose, not mechanics.
  13. x_hat, y_hat, m_hat, v_hat, s_hat, and med_hat for generic sample quantities; xy_hat for the Part 2 and Part 3 lab data; x0_hat and xy0_hat for small examples; x_sim and xy_sim for simulated draws; y_fit for fitted values and e_hat for residuals.
  14. F_hat for ECDFs; f_hat for kernel density estimates and stored histograms.

Further Reading

  • https://style.tidyverse.org/ – a widely used R style guide. It assumes the tidyverse rather than base R, so read it for the general principles rather than the specific verbs.
  • Figures – the matching conventions for plots.