A hypothesis test turns a question about the population into a decision: given the data, do we have enough evidence to reject a specific value for the parameter? This chapter conducts those tests two ways (by inverting a confidence interval and by imposing the null), then summarizes the evidence with a \(p\)-value and a \(t\)-value. We also cover one-sided versions and common pitfalls so that “statistically significant” stays linked to its statistical meaning.
Testing Methods
In this chapter, we test hypotheses using data-driven methods that assume much less about the data generating process. There are two main ways to conduct a hypothesis test: inverting a confidence interval and imposing the null. The first treats the distribution of estimates directly; the second explicitly enforces the null hypothesis to evaluate how unusual the observed statistic is. Both approaches rely on the bootstrap to approximate sampling variability.
Before either test can run, we have to name the claim being tested and the alternative we would accept if the data spoke against it.
The null hypothesis \(H_{0}\) is the specific claim about the parameter that we will test.
The alternative hypothesis \(H_{A}\) is the competing claim.
A two-sided hypothesis test covers deviations from the null in either direction (alternative values are higher or lower than the null).
This setup is useful because it forces a single direction of inference: the data either contradict \(H_{0}\) enough to reject it, or they do not, but in no case do they “prove” \(H_{A}\). The most common test concerns the mean, \(H_{0}: \mathbb{E}[X_i] = m_0\), where \(m_0\) is a hypothesized value. The two-sided alternative is \(H_A: \mathbb{E}[X_i] \neq m_0\). E.g., you hypothesize your stock returns have earnings \(m_0=0\) and consider that the returns could truly be either positive or negative.
Invert a CI
One way to test a hypothesis is to look at whether a confidence interval contains the hypothesized value: if the value lies outside the interval, sampling variability alone cannot easily explain the gap.
The rejection region is the set of values for the test statistic (or hypothesized parameter) that lead us to reject \(H_{0}\). For a two-sided test at level \(\alpha=0.05\), the rejection region is the two tails outside the central \(95\%\) of the null distribution.
The rejection region is useful as a decision rule that mirrors how we read a confidence interval:
- reject \(H_{0}\) if the hypothesized value falls outside the interval (it is in the rejection region)
- fail to reject \(H_{0}\) if it falls inside
A \(95\%\) confidence interval corresponds to a \(5\%\) rejection region split between the two tails, which links the test level directly to the coverage of the interval (see the previous Confidence Intervals chapter for how the interval is computed).
For example, suppose you hypothesize the mean murder arrest rate is \(9\) per \(100{,}000\). You then construct a bootstrap distribution with \(95\%\) confidence interval, and find your hypothesized value falls outside of the confidence interval. Then, after accounting for sampling variability (which you estimate), it still seems extremely unlikely that the theoretical mean actually equals \(9\), so you reject that hypothesis. (If the theoretical value landed in the interval, you would “fail to reject” the theoretical mean equals \(9\).)
Code
x_hat <- USArrests[, 'Murder'] # lab data: murder arrests per 100,000
m_hat <- mean(x_hat)
# Bootstrap Distribution
n <- length(x_hat)
set.seed(1) # to be replicable
boot_means <- rep(NA, 999)
for(b in seq_along(boot_means)){
x_boot <- sample(x_hat, replace=TRUE) # c.f. jackknife
m_b_hat <- mean(x_boot)
boot_means[b] <- m_b_hat
}
hist(boot_means, breaks=25,
border=NA,
freq=FALSE,
main=NA,
xlab='Bootstrap Samples')
# CI
ci_95 <- quantile(boot_means, probs=c(.025, .975))
abline(v=ci_95, lwd=2)
# H0: mean=9
abline(v=9, col=rgb(1, 0, 0, .8), lwd=2)
The above procedure also generalizes to many other statistics. Perhaps the most informative are those for spread or shape. E.g., you can conduct hypothesis tests for sd and IQR, or skew and kurtosis.
Code
# Bootstrap Distribution for SD
s_hat <- sd(x_hat)
boot_sds <- rep(NA, 999)
for(b in seq_along(boot_sds)){
x_boot <- sample(x_hat, replace=TRUE)
s_b_hat <- sd(x_boot)
boot_sds[b] <- s_b_hat
}
# Test for SD Differences (Invert CI)
sd_null <- 3.6
hist(boot_sds, freq=FALSE,
border=NA, xlab='Bootstrap',
main=NA)
title('Standard Deviations (Invert CI)', font.main=1)
sd_ci <- quantile(boot_sds, probs=c(0.025, .975) )
abline(v=sd_ci, lwd=2)
abline(v=sd_null, lwd=2, col=rgb(1, 0, 0, .8))
To better your understanding, try redoing the above for any function (such as IQR(x_boot)/median(x_boot))
Suppose a student scores \(77\%\) on a \(44\)-question multiple-choice exam but believes they are really an \(85\%\) student who had a bad day. The professor is skeptical: “where’s your evidence?” The student bootstraps a confidence interval for their “true” proportion correct, resampling from the \(44\) right/wrong answers on that one exam.
This creates a conundrum. If the confidence interval comes out narrow, it may exclude \(85\%\), and the \(85\%\) hypothesis is rejected. But an interval wide enough to contain \(85\%\) is also wide enough to contain \(70\%\), so the same evidence fails to rule out a worse student who simply got lucky. A single exam cannot easily distinguish the two stories.
Code
# One exam: 34 correct out of 44 questions
n_questions <- 44
n_correct <- 34
outcomes <- c(rep(1, n_correct), rep(0, n_questions - n_correct)) # 1=correct, 0=incorrect
# Bootstrap Distribution for the proportion correct
set.seed(1)
boot_props <- rep(NA, 999)
for(b in seq_along(boot_props)){
x_boot <- sample(outcomes, replace=TRUE)
boot_props[b] <- mean(x_boot)
}
hist(boot_props, breaks=25,
border=NA, freq=FALSE,
main=NA, xlab='Bootstrap Proportion Correct')
ci_95 <- quantile(boot_props, probs=c(.025, .975))
abline(v=ci_95, lwd=2)
abline(v=c(.70, .85), col=rgb(1, 0, 0, .8), lwd=2, lty=c(2, 1))
The interval comfortably contains both \(70\%\) and \(85\%\), so the student can reject neither story.
The deeper issue is what the bootstrap resample actually represents. Resampling only reshuffles the same \(44\) questions, so it approximates how the score would vary if the student retook that exact exam. It says nothing about how the student would do on a genuinely different exam with different questions, which is the sampling distribution the student actually needs to make their case. Estimating that sampling distribution properly needs new, independent samples: more exams, not more resamples of one exam, and that is a cost the student is unlikely to accept for the rest of the semester.
Impose the Null
We can also compute a null distribution: the sampling distribution of the statistic under the null hypothesis (assuming your null hypothesis was true). We use the bootstrap to loop through a large number of “resamples”. In each iteration of the loop, we impose the null hypothesis and re-estimate the statistic of interest. We then calculate the range of the statistic across all resamples and compare how extreme the original value we observed is.
For example, suppose you again hypothesize the mean murder arrest rate is \(9\). You then construct a \(95\%\) confidence interval around the null bootstrap distribution (resamples centered around \(9\)). If your sample mean falls outside of that interval, then even after accounting for sampling variability (which you estimate), it seems extremely unlikely that the theoretical mean actually equals \(9\), so you reject that hypothesis. (If the sample mean landed in the interval, you would “fail to reject” the theoretical mean equals \(9\).)
Code
m_hat <- mean(x_hat)
# Bootstrap NULL: mean=9
# Bootstrap shift: center each bootstrap resample so that the distribution satisfies the null hypothesis on average.
set.seed(1)
null_mean <- 9
boot_means_null <- rep(NA, 999)
for(b in seq_along(boot_means_null)){
x_boot <- sample(x_hat, replace=TRUE)
m_b_hat <- mean(x_boot) + (null_mean - m_hat) # impose the null via Bootstrap shift
boot_means_null[b] <- m_b_hat
}
hist(boot_means_null, breaks=25, border=NA, freq=FALSE,
main=NA,
xlab='Null Bootstrap Samples')
ci_95 <- quantile(boot_means_null, probs=c(.025, .975)) # critical region
abline(v=ci_95, lwd=2)
abline(v=m_hat, lwd=2, col=rgb(0, 0, 1, .8))
Why does adding \((m_0 - \hat{m})\) impose the null? A bootstrap resample of the data has a mean that varies around the sample mean \(\hat{m}\), not around the hypothesized \(m_0\). Adding the constant \((m_0 - \hat{m})\) slides the whole resampling distribution over so it is centered at \(m_0\) instead. A shift moves the center but leaves the spread and shape unchanged, so the result is a sampling distribution with the null mean and the same variability as the data.
For example, take the three observations \(\{2, 4, 9\}\), which have mean \(\hat{m}=5\), and the null \(m_0=6\). Adding \(m_0-\hat{m}=1\) to every observation gives \(\{3, 5, 10\}\), which has mean \(6\) and the same spread.
Code
x0_hat <- c(2, 4, 9)
m0_hat <- mean(x0_hat)
x0_null <- x0_hat + (6 - m0_hat) # shift so the mean equals 6
x0_null
## [1] 3 5 10
c(mean(x0_null), sd(x0_hat), sd(x0_null))
## [1] 6.000000 3.605551 3.605551
The Normal approximation from the previous chapter can also impose the null: center the interval at the hypothesized value \(m_0\) and use \(m_0 \pm q(\alpha/2) \cdot SE(M)\), where \(SE(M)\) is estimated via bootstrap or classical formulas. (While we could also use a Null Jackknife distribution, that is rarely done in practice.) Altogether, there are two different types of confidence intervals that “impose the null”. Until you know more, a conservative rule-of-thumb is to take the larger estimate.
Types of Confidence Interval Estimates that “impose the null”
| Bootstrap (Percentile) |
randomly resample \(n\) observations with replacement and shift |
| Normal |
assume sampling distribution is Normal and use formula \(\pm q(\alpha/2) \cdot SE\) |
\(p\)-values
Reading a hypothesis test off a confidence interval is binary (reject or not); often we also want a number that describes how extreme the observed statistic is relative to the null.
The \(p\)-value is the probability, under the null hypothesis, of seeing a statistic at least as extreme as the one observed. Small \(p\)-values give evidence against the null; the threshold \(p \leq 0.05\) is a convention, not a law of nature. The \(p\)-value is not the probability that the null is true.
The \(p\)-value is useful as a continuous summary of how strongly the data oppose the null, whereas a confidence interval gives only a yes/no answer at a fixed coverage level. For the running mean example, we want the probability that the random variable \(M\) is at least as extreme (far from the null mean of \(9\)) as our observed sample mean \(\hat{m}\).
Recall that we used the bootstrap to estimate the distribution of the sample statistic like the mean, and the null-bootstrap shifted the bootstrap to be centered at a hypothesized value. The bootstrap idea here is to approximate \(M-\mathbb{E}[X_i]\), the difference between the sample mean \(M\) and the unknown theoretical mean \(\mathbb{E}[X_i]\), with \(M^{\text{boot}}-\hat{m}\). To test \(H_0: \mathbb{E}[X_i]=m_0\), we recenter the bootstrap mean as \(M^{\text{boot}}_{0}=M^{\text{boot}}+(m_0-\hat{m})\), so \(M^{\text{boot}}_{0}-m_0=M^{\text{boot}}-\hat{m}\). For the running example, \(m_0=9\). \[\begin{aligned}
& Prob( |M - m_0| \geq |\hat{m} - m_0| \mid \mathbb{E}[X_i] = m_0 ) \\
& \approx Prob( |M^{\text{boot}}_{0}- m_0| \geq |\hat{m}- m_0| ) \\
& = 1-\hat{F}^{|\text{boot}|}_{0}(|\hat{m}-9|),
\end{aligned}\] where \(\hat{F}^{|\text{boot}|}_{0}\) is the ECDF of \(|M^{\text{boot}}_{0}- m_0|\).
Using the null bootstrap distribution boot_means_null from above, we compute the two-sided \(p\)-value for the murder data.
Code
# Two-Sided Test, ALTERNATIVE: mean < 9 or mean >9
# Visualize Two Sided Prob. & reject region boundary
par(mfrow=c(1, 2))
hist(boot_means_null-null_mean,
freq=FALSE, breaks=20,
border=NA,
main=NA,
xlab=expression('Null Bootstrap for M - '~m[0]))
abline(v=m_hat-null_mean, col=rgb(0, 0, 1, .8))
ci_95 <- quantile(boot_means_null-null_mean, probs=c(0.025, .975))
abline(v=ci_95, lwd=2)
# Equivalent Visualization
boot_absval <- abs(boot_means_null-null_mean)
F_hat <- ecdf(boot_absval)
plot(F_hat,
main=NA,
xlab=expression('Null Bootstrap for |M - '~m[0]~'|'))
abline(v=abs(m_hat-null_mean), col=rgb(0, 0, 1, .8))
# with Two Sided Probability
p2 <- 1 - F_hat( abs(m_hat-null_mean) )
title( paste0('p=', round(p2, 3)), font.main=1)
Continuing the exam example from above: the student scored \(\hat{p}=34/44 \approx 77\%\) and wants to test \(H_{0}: p=0.85\). Imposing the null shifts the bootstrap distribution of the proportion correct so that it is centered at \(0.85\) instead of at \(\hat{p}\), exactly as was done for the mean.
Code
n_questions <- 44
n_correct <- 34
outcomes <- c(rep(1, n_correct), rep(0, n_questions - n_correct))
p_hat <- mean(outcomes)
set.seed(1)
# Bootstrap NULL: p=0.85
# Bootstrap shift: center each bootstrap resample so that the distribution satisfies the null hypothesis on average.
p0 <- 0.85
boot_props_null <- rep(NA, 999)
for(b in seq_along(boot_props_null)){
x_boot <- sample(outcomes, replace=TRUE)
boot_props_null[b] <- mean(x_boot) + (p0 - p_hat) # impose the null via Bootstrap shift
}
# Two-Sided p-value
boot_absval <- abs(boot_props_null - p0)
F_hat <- ecdf(boot_absval)
p2_exam <- 1 - F_hat( abs(p_hat - p0) )
hist(boot_props_null, breaks=25, border=NA, freq=FALSE,
main=NA, xlab='Null Bootstrap Proportion (H0: p=0.85)')
ci_95 <- quantile(boot_props_null, probs=c(.025, .975))
abline(v=ci_95, lwd=2)
abline(v=p_hat, lwd=2, col=rgb(0, 0, 1, .8))
title( paste0('p=', round(p2_exam, 3)), font.main=1)
The two-sided \(p\)-value comes out around \(0.2\), far above the conventional \(5\%\) threshold, so we fail to reject \(H_{0}: p=0.85\). This matches the earlier conclusion from the confidence interval, but the \(p\)-value adds a precise number for how weak the evidence against \(85\%\) actually is, rather than just a reject/fail-to-reject label. Testing \(H_{0}: p=0.70\) the same way would also fail to reject: the \(p\)-value is large in both directions because a single \(44\)-question exam simply cannot pin down the student’s true ability closely enough to rule out either story.
You can conduct hypothesis test using \(p\)-values instead of confidence intervals. It is common to use this decision rule:
- reject the null at the \(5\%\) level if \(p \leq 0.05\)
- fail to reject the null at the \(5\%\) level if \(p > 0.05\)
Caveats
Beware that a common misreading of the \(p\)-value as “the probability the null is true”. That is false. A \(p\)-value is the frequency you see something at least as extreme as your statistic under the null hypothesis.
Often, one may also see or hear “\(p<.05\): statistically significant” and “\(p>.05\): not statistically significant”. That is decision making on purely statistical grounds, and it may or may not be suitable for your context. You simply need to know that whoever says those things is using \(5\%\) as a critical value to reject a null hypothesis.
Running many hypothesis tests inflates the chance of a false rejection. At the \(5\%\) level, you expect roughly \(1\) in \(20\) true nulls to be rejected by chance alone, so if you test \(20\) independent nulls and find \(1\) “significant” result, that result is what you would expect even when every null is true. The fix is to correct the level (e.g., Bonferroni: use \(0.05/k\) when running \(k\) tests) or to pre-specify a small set of hypotheses before looking at the data.
Code
# Purely-Statistical Decision Making
# via Two Sided Test
if(p2 >.05){
print('fail to reject the null that mean=9, at the 5% level')
} else {
print('reject the null that mean=9 in favor of either <9 or >9, at the 5% level')
}
## [1] "reject the null that mean=9 in favor of either <9 or >9, at the 5% level"
Also note that the \(p\)-value is itself a function of data, and hence a random variable that changes from sample to sample. To see this, return to the wage data from Sampling and pretend its \(3294\) workers are the whole population. Then we know the population mean, and we can draw many samples and test the null that is actually true.
Code
# Known population: hourly wages of 3294 workers
data(Wages1, package='Ecdat')
population_wages <- Wages1[, 'wage']
population_wage_mean <- mean(population_wages)
# 300 samples, each tested against the true mean
p_values <- rep(NA, 300)
for(s in seq_along(p_values)){
x_sim <- sample(population_wages, 50, replace=FALSE)
m_sim <- mean(x_sim)
boot_means_null_sim <- rep(NA, 999)
for(b in seq_along(boot_means_null_sim)){
x_boot <- sample(x_sim, replace=TRUE)
boot_means_null_sim[b] <- mean(x_boot) + (population_wage_mean - m_sim) # impose the (true) null
}
F_hat <- ecdf( abs(boot_means_null_sim-population_wage_mean) )
p_values[s] <- 1 - F_hat( abs(m_sim-population_wage_mean) )
}
hist(p_values, breaks=seq(0, 1, by=.05),
freq=FALSE, border=NA, main=NA,
xlab='p-values when the null is true')
abline(v=.05, col=rgb(1, 0, 0, .8), lwd=2)
Code
# Share of samples that falsely reject at the 5% level
mean(p_values <= .05)
## [1] 0.06666667
The null is true in every one of these samples, yet the \(p\)-values spread across the whole range from \(0\) to \(1\). A few samples still reject at the \(5\%\) level, which is exactly the false rejection described above. Given that the \(5\%\) level is somewhat arbitrary, and that the \(p\)-value both varies from sample to sample and is often misunderstood, it makes sense to give specific \(p\)-values a limited role in decision making.
The \(p\)-value also varies a little each time you rerun the bootstrap on the same data, because each run draws different resamples. This variation comes from the computer, not from the sample, and it shrinks as you use more resamples. Here we repeat the murder test \(300\) times on the one sample we have.
Code
p_values_boot <- rep(NA, 300)
for(b2 in seq(p_values_boot)){
boot_means_null_p <- rep(NA, 999)
for(b in seq_along(boot_means_null_p)){
x_boot <- sample(x_hat, replace=TRUE)
m_b_hat <- mean(x_boot) + (null_mean - m_hat) # impose the null
boot_means_null_p[b] <- m_b_hat
}
F_hat <- ecdf( abs(boot_means_null_p-null_mean) )
p_values_boot[b2] <- 1- F_hat( abs(m_hat-null_mean) )
}
hist(p_values_boot, freq=FALSE,
border=NA, main=NA)
abline(v=.05, col=rgb(1, 0, 0, .8), lwd=2)
The \(p\)-values here land on both sides of \(0.05\), so rerunning the bootstrap alone could flip the decision. This is one more reason not to lean on the exact \(5\%\) cutoff.
Other Statistics
\(t\)-values
Different studies have different \(SE(M)\), so the raw difference \(\hat{m} - m_0\) is hard to compare across studies; we need to express the deviation in standard-error units.
The \(t\)-value (or \(t\)-statistic) standardizes the deviation of \(\hat{m}\) from the hypothesized \(m_0\) by the standard error, \[t = \frac{M - m_0}{SE(M)}.\] Using a standardized statistic makes results comparable across studies and connects to a well-studied reference distribution (Normal in large samples, Student’s \(t\) in small).
The \(t\)-value is useful as a portable test statistic: a \(t\) of \(2.05\) has roughly the same evidentiary weight whether the underlying data are measured in dollars, hours, or grams, because the units cancel between numerator and denominator. For any specific sample we must approximate the standard error. Using the theory-driven approach, we compute \(\hat{t}=(\hat{m}-m_0)/(\hat{s}/\sqrt{n})\). Using the data-driven approach, we compute \(\hat{t}=(\hat{m}-m_0)/(\hat{SE}^{\text{jack}})\) or \(\hat{t}=(\hat{m}-m_0)/(\hat{SE}^{\text{boot}})\). In any case, we can use bootstrapping to estimate the variability of the \(t\) statistic, just like we did with the mean.
Take the four observations \(\{3, 5, 7, 9\}\) and the null \(m_0=4\). The sample mean is \(\hat{m}=6\) and the sample standard deviation is \(\hat{s}\approx 2.58\). The theory-driven standard error is \(\hat{s}/\sqrt{4} \approx 1.29\), so \(\hat{t} \approx (6-4)/1.29 \approx 1.55\). The sample mean is about one and a half standard errors above the null.
Code
x0_hat <- c(3, 5, 7, 9)
se0_hat <- sd(x0_hat)/sqrt(length(x0_hat))
t0_hat <- (mean(x0_hat) - 4)/se0_hat
c(se0_hat, t0_hat)
## [1] 1.290994 1.549193
For the murder data, we compute \(\hat{t}\) with jackknife standard errors and then bootstrap its null distribution.
Code
#null hypothesis
null_mean <- 9
# t statistic
jack_means <- rep(NA, length(x_hat))
for(i in seq_along(jack_means)){
x_jack <- x_hat[-i]
jack_means[i] <- mean(x_jack)
}
jackknife_se <- sd(jack_means)*sqrt(length(x_hat))
t_hat <- (m_hat - null_mean)/jackknife_se
# Boostrap Null Distribution
boot_t_null <- rep(NA, 999)
for(b in seq_along(boot_t_null)){
x_boot <- sample(x_hat, replace=TRUE)
m_b_hat <- mean(x_boot) + (null_mean - m_hat) # impose the null by recentering
# Compute t stat using jackknife ses (same as above)
jack_means_b <- rep(NA, length(x_boot))
for(i in seq_along(jack_means_b)){
jack_means_b[i] <- mean(x_boot[-i])
}
jackknife_se_b <- sd(jack_means_b)*sqrt(length(x_boot))
t_b_hat <- (m_b_hat - null_mean)/jackknife_se_b
boot_t_null[b] <- t_b_hat
}
# Plot the null distribution and CI
par(mfrow=c(1, 2))
hist(boot_t_null, border=NA, breaks=50,
freq=FALSE, main=NA, xlab='Null Bootstrap for t')
abline(v=t_hat, col=rgb(0, 0, 1, .8))
ci_95 <- quantile(boot_t_null, probs=c(0.025, 0.975) )
abline(v=ci_95, lwd=2)
# Compute the p-value for two-sided test
F_hat <- ecdf(abs(boot_t_null))
plot(F_hat,
xlim=range(boot_t_null, t_hat),
xlab='Null Bootstrap for |t|',
main=NA)
abline(v=abs(t_hat), col=rgb(0, 0, 1, .8))
p <- 1 - F_hat( abs(t_hat) )
title( paste0('p=', round(p, 3)), font.main=1)
Code
if(p >.05){
print('fail to reject the null that mean=9, at the 5% level')
} else {
print('reject the null that mean=9 in favor of either <9 or >9, at the 5% level')
}
## [1] "fail to reject the null that mean=9, at the 5% level"
There are several benefits to this statistic:
- uses the same statistic for different hypothesis tests
- makes the statistic comparable across different studies
- removes dependence on unknown parameters by normalizing with a standard error
- makes the null distribution theoretically known asymptotically (approximately)
The last point implies we are typically dealing with a normal distribution that is well-studied, or another well-studied distribution derived from it.
Quantiles and Shape Statistics
Bootstrap allows hypothesis tests for any statistic, not just the mean, without relying on parametric theory. For example, the above procedures generalize from means to quantile statistics like medians.
Code
# Test for Median Differences (Impose the Null)
# Bootstrap Null Distribution for the median
# Each Bootstrap shifts medians so that median = null_median
med_hat <- quantile(x_hat, probs=.5)
null_median <- 7.8
boot_medians_null <- rep(NA, 999)
for(b in seq_along(boot_medians_null)){
x_boot <- sample(x_hat, replace=TRUE) #bootstrap sample
med_b_hat <- quantile(x_boot, probs=.5) # median
med_null_b_hat <- med_b_hat + (null_median - med_hat) # impose the null
boot_medians_null[b] <- med_null_b_hat
}
# 2-Sided Test for Medians
hist(boot_medians_null-null_median,
border=NA, freq=FALSE, xlab='Null Bootstrap',
main=NA)
title('Medians (Impose Null)', font.main=1)
median_ci <- quantile(boot_medians_null-null_median, probs=c(.025, .975))
abline(v=median_ci, lwd=2)
abline(v=med_hat-null_median, lwd=2, col=rgb(0, 0, 1, .8))
Code
# 2-Sided Test for Median Difference
## Null: No Median Difference
1 - ecdf( abs(boot_medians_null-null_median))( abs(med_hat-null_median) )
## [1] 0.5435435
Conduct a hypothesis test for whether the upper quartile is statistically different from \(12\).
Code
q_hat <- quantile(x_hat, probs=.75)
One-Sided Tests
Above, we tested whether the observed statistic is either extremely high or low. This is known as a two-sided test. There are also two one-sided tests (left tail: observed statistic is extremely low, right tail: observed statistic is extremely high). For a concrete example, consider whether the mean statistic, \(M\), is centered on a theoretical value of \(\mathbb{E}[X_i]=9\) for the population. If your null hypothesis is that the theoretical mean is nine, \(H_{0}: \mathbb{E}[X_i] =9\), and you calculated the mean for your sample as \(\hat{m}\), then you can consider any one of these three alternative hypotheses:
- \(H_{A}: \mathbb{E}[X_i] \neq 9\), a two-tail test
- \(H_{A}: \mathbb{E}[X_i] < 9\), a left-tail test
- \(H_{A}: \mathbb{E}[X_i] > 9\), a right-tail test
A fitness tracker manufacturer claims that users take more than \(8{,}000\) steps per day on average. A sample of \(25\) users has \(\hat{m}=8{,}492\) steps and \(\hat{s}=1{,}200\) steps. Test at the \(5\%\) level using theory-based intervals.
(a) Hypotheses. The claim is a right-tail alternative: \(H_{0}: \mathbb{E}[X_i] = 8000\) versus \(H_{A}: \mathbb{E}[X_i] > 8000\).
(b) Test statistic. The classical standard error is \(\hat{s}/\sqrt{n} = 1200/\sqrt{25} = 240\), so \(t = (8492 - 8000)/240 \approx 2.05\).
(c) Decision. The right-tail critical value at \(5\%\) is \(q_{0.95} \approx 1.645\) from the Normal. Since \(t = 2.05 > 1.645\), we reject \(H_{0}\).
(d) Conclusion. At the \(5\%\) level, the data are consistent with the manufacturer’s claim: there is evidence the population mean exceeds \(8{,}000\) steps per day.
Code
m_steps_hat <- 8492
s_steps_hat <- 1200
n <- 25
steps_null_mean <- 8000
SE <- s_steps_hat / sqrt(n)
t_stat <- (m_steps_hat - steps_null_mean) / SE
t_stat
## [1] 2.05
# Right-tail p-value from the standard Normal approximation
1 - pnorm(t_stat)
## [1] 0.02018222
One-sided hypothesis tests can be conducted by inverting a one-sided confidence interval (covered in the previous chapter) or by computing a one-sided \(p\)-value.
One-Sided \(p\)-values
The \(p\)-value for a one-sided test is more straightforward to implement via a bootstrap null distribution.
For a left-tail test, we examine \[\begin{aligned}
p = Prob( M < \hat{m} \mid \mathbb{E}[X_i] = 9 )
&\approx Prob( M^{\text{boot}}_{0} < \hat{m} ) = \hat{F}^{\text{boot}}_{0}(\hat{m}),
\end{aligned}\] where \(\hat{F}^{\text{boot}}_{0}\) is the ECDF of the bootstrap null distribution. We reject the null if \(p < 0.05\) at the \(5\%\) level, and otherwise fail to reject.
For a right-tail test, we examine \(p=Prob( M > \hat{m} \mid \mathbb{E}[X_i] = 9 ) \approx 1-\hat{F}^{\text{boot}}_{0}(\hat{m})\).
Code
# Right-tail Test, ALTERNATIVE: mean > 9
# Equivalent Visualization with p-value
F_hat <- ecdf(boot_means_null) # Look at right tail
plot(F_hat,
main=NA,
xlab='Null Bootstrap')
abline(v=m_hat, col=rgb(0, 0, 1, .8))
p1 <- 1- F_hat(m_hat) #Compute right tail prob: 0.987
title( paste0('p=', round(p1, 3)), font.main=1)
Code
if(p1 >.05){
print('fail to reject the null that mean=9, at the 5% level')
} else {
print('reject the null that mean=9 in favor of >9, at the 5% level')
}
## [1] "fail to reject the null that mean=9, at the 5% level"
The left-tail test uses the other side of the same null distribution.
Code
# Left-tail Test, ALTERNATIVE: mean < 9
# Compute left-tail prob: p = F(m_hat)
p_left <- F_hat(m_hat)
p_left
## [1] 0.01301301
if(p_left >.05){
print('fail to reject the null that mean=9, at the 5% level')
} else {
print('reject the null that mean=9 in favor of <9, at the 5% level')
}
## [1] "reject the null that mean=9 in favor of <9, at the 5% level"
For two-sided tests, measure distance from \(m_0\) by subtracting it before taking absolute values. For one-sided tests, subtracting \(m_0\) from both the null-bootstrap statistic and the observed statistic leaves their ordering unchanged. Specifically, \(p = Prob( M < \hat{m} \mid \mathbb{E}[X_i] = 9 ) = Prob( M - m_0 < \hat{m} - m_0 \mid \mathbb{E}[X_i] = 9 )\). That is intuitively also why the \(t\)-value can be used for both one and two-sided hypothesis tests.
Code
# See that the 'recentering' matters for two-sided tests
ecdf( abs(boot_means_null-null_mean) )( abs(m_hat-null_mean) )
## [1] 0.966967
ecdf( abs(boot_means_null) )( abs(m_hat) )
## [1] 0.01301301
# See that the 'recentering' doesn't matter for one-sided ones
ecdf( boot_means_null-null_mean)( m_hat-null_mean)
## [1] 0.01301301
ecdf( boot_means_null )( m_hat)
## [1] 0.01301301
Exercises
Comment the script you wrote for this chapter, then restart R and check that it runs from a clean session, then check the script with AI as explained in Working with AI. Write three sentences from memory on the main statistical idea of this chapter, and ask the assistant what is wrong, vague, or missing. Finish with your own questions about whatever you found hardest.
Explain the difference between “rejecting the null hypothesis” and “proving the alternative hypothesis is true.” Why is the phrase “fail to reject” used instead of “accept the null”?
Suppose you collect a sample of \(n = 50\) observations with \(\hat{m} = 22.4\) and \(\hat{s} = 6.0\). You want to test \(H_{0}: \mathbb{E}[X_i] = 20\) against \(H_{A}: \mathbb{E}[X_i] \neq 20\) at the \(5\%\) level. Compute the \(t\)-value using theory-driven standard errors and determine whether you reject or fail to reject.
Using the USArrests dataset in R, test the hypothesis that the population mean of Assault equals \(150\) by constructing a bootstrap null distribution. Compute the two-sided \(p\)-value and state whether you reject or fail to reject at the \(5\%\) level.
Recall
This chapter built hypothesis tests two ways (invert a CI and impose the null), introduced the \(p\)-value and the \(t\)-value, and flagged the multiple-comparisons pitfall. The fitness-tracker example tied these ideas together: with \(\hat{m}=8492\), \(\hat{s}=1200\), and \(n=25\), the standard error is \(SE = 1200/\sqrt{25} = 240\) and the \(t\)-value for \(H_{0}: \mathbb{E}[X_i] = 8000\) is \((8492 - 8000)/240 \approx 2.05\), which exceeds the right-tail critical value \(1.645\) at the \(5\%\) level and so rejects the null.