Hypothesis testing for a population mean
Master hypothesis testing for a population mean in Year 12 VCE Specialist Mathematics. A hypothesis test uses a single sample to decide between a null hypothesis \(H_0:\mu=\mu_0\) and an alternative \(H_1\), standardising the sample mean to the test statistic \(z=\dfrac{\bar x-\mu_0}{\sigma/\sqrt n}\). It sits in the Data analysis, probability and statistics area of study of the VCE Mathematics Study Design (VCAA), within the Statistical inference topic of Unit 4.
You will learn to state the hypotheses, choose a one- or two-tailed test, compute the test statistic and a p-value or critical value, then decide at a chosen level of significance \(\alpha\) and interpret the result as evidence in context — never as proof.
Theory
A hypothesis test for a population mean uses a single sample to decide between a null hypothesis \(H_0:\mu=\mu_0\) and an alternative \(H_1\). In Year 12 Specialist Mathematics you standardise the sample mean to the test statistic \(z=\dfrac{\bar x-\mu_0}{\sigma/\sqrt n}\), find a p-value or a critical value, then decide at a chosen level of significance \(\alpha\) and interpret the result in context.
A hypothesis test weighs the evidence a sample gives against a claim about the population mean \(\mu\). The null hypothesis \(H_0:\mu=\mu_0\) states that the mean equals a specified value; the alternative hypothesis \(H_1\) is what you test for.
The alternative sets the number of tails. A two-tailed test \(H_1:\mu\ne\mu_0\) asks whether the mean has changed in either direction; a one-tailed test \(H_1:\mu>\mu_0\) or \(H_1:\mu<\mu_0\) asks whether it has changed in one specific direction.
The test statistic standardises the sample mean under \(H_0\): \(z=\dfrac{\bar x-\mu_0}{\sigma/\sqrt n}\), where \(\sigma\) is the known population standard deviation and \(n\) the sample size. It measures how many standard errors \(\bar x\) sits from \(\mu_0\). (For a large sample, \(n\ge 30\), the sample standard deviation \(s\) may replace \(\sigma\).)
The level of significance \(\alpha\) (commonly \(0.05\) or \(0.01\)) is the risk you accept of rejecting a true \(H_0\). The p-value is the probability, if \(H_0\) were true, of a test statistic at least as extreme as the one observed. You reject \(H_0\) when \(p<\alpha\) (equivalently, when \(|z|\) exceeds the critical value); otherwise you do not reject \(H_0\).
Interpretation is evidence, not proof. Rejecting \(H_0\) means the sample gives sufficient evidence to reject it at the \(\alpha\) level — it never proves \(H_1\). Failing to reject \(H_0\) means there is insufficient evidence against it, not that \(H_0\) is true.
The standardised test statistic for a sample of size \(n\) with mean \(\bar x\), drawn from a population with mean \(\mu_0\) under \(H_0\) and known standard deviation \(\sigma\):
The p-value is the tail area beyond the observed \(z\), using the standard normal cdf \(\Phi\). For a one-tailed test take one tail; for a two-tailed test double it:
Decision. Reject \(H_0\) when \(p<\alpha\), or equivalently when \(|z|\) exceeds the critical value for the test:
| Test | \(5\%\) level | \(1\%\) level |
|---|---|---|
| One-tailed critical value | \(1.645\) | \(2.326\) |
| Two-tailed critical value | \(1.96\) | \(2.576\) |
How to run a hypothesis test for the mean
- State the hypotheses. Write \(H_0:\mu=\mu_0\) and choose \(H_1\) from the context: \(\ne\) for a change either way (two-tailed), \(>\) or \(<\) for a specific direction (one-tailed).
- Standardise. Compute the test statistic \(z=\dfrac{\bar x-\mu_0}{\sigma/\sqrt n}\) using the sample mean \(\bar x\), the value \(\mu_0\), the known \(\sigma\) and the size \(n\).
- Measure the evidence. Find the p-value — \(1-\Phi(|z|)\) for one tail, \(2\bigl(1-\Phi(|z|)\bigr)\) for two — or read off the critical value for the level \(\alpha\).
- Decide. Reject \(H_0\) if \(p<\alpha\) (equivalently \(|z|\) beyond the critical value); otherwise do not reject \(H_0\).
- Interpret in context. State the conclusion as evidence: e.g. "sufficient evidence at the \(5\%\) level that the mean differs from \(\mu_0\)" — never "this proves \(H_1\)".
State the hypotheses (a two-tailed test):
| \(H_0\) | \(:\) | \(\mu=250\) |
| \(H_1\) | \(:\) | \(\mu\ne250\) |
Standardise \(z=\dfrac{\bar x-\mu_0}{\sigma/\sqrt n}\):
| \(z\) | \(=\) | \(\dfrac{246-250}{8/\sqrt{16}}\) |
| \(=\) | \(\dfrac{-4}{2}\) | |
| \(=\) | \(-2\) |
Two-tailed p-value \(2\bigl(1-\Phi(|z|)\bigr)\):
| \(p\) | \(=\) | \(2\bigl(1-\Phi(2)\bigr)\) |
| \(=\) | \(2(1-0.9772)\) | |
| \(=\) | \(0.0455\) |
Compare and decide at \(5\%\):
| \(|z|\) | \(=\) | \(2>1.96\) |
| \(\Rightarrow\) | \(\text{reject } H_0\) |
Reject \(H_0\): there is sufficient evidence at the \(5\%\) level that the mean volume differs from \(250\) mL.
State the hypotheses (a one-tailed, lower test):
| \(H_0\) | \(:\) | \(\mu=500\) |
| \(H_1\) | \(:\) | \(\mu<500\) |
Standardise:
| \(z\) | \(=\) | \(\dfrac{490-500}{30/\sqrt{36}}\) |
| \(=\) | \(\dfrac{-10}{5}\) | |
| \(=\) | \(-2\) |
One-tailed p-value \(1-\Phi(|z|)\):
| \(p\) | \(=\) | \(1-\Phi(2)\) |
| \(=\) | \(1-0.9772\) | |
| \(=\) | \(0.0228\) |
Compare with the \(5\%\) critical value \(1.645\):
| \(|z|\) | \(=\) | \(2>1.645\) |
| \(\Rightarrow\) | \(\text{reject } H_0\) |
Reject \(H_0\): there is sufficient evidence at the \(5\%\) level that the mean breaking strength is less than \(500\) N.
State the hypotheses (a one-tailed, upper test):
| \(H_0\) | \(:\) | \(\mu=70\) |
| \(H_1\) | \(:\) | \(\mu>70\) |
Standardise:
| \(z\) | \(=\) | \(\dfrac{75-70}{12/\sqrt{36}}\) |
| \(=\) | \(\dfrac{5}{2}\) | |
| \(=\) | \(2.5\) |
One-tailed p-value \(1-\Phi(|z|)\):
| \(p\) | \(=\) | \(1-\Phi(2.5)\) |
| \(=\) | \(1-0.9938\) | |
| \(=\) | \(0.0062\) |
Compare with the \(5\%\) critical value \(1.645\):
| \(z\) | \(=\) | \(2.5>1.645\) |
| \(\Rightarrow\) | \(\text{reject } H_0\) |
Reject \(H_0\): there is sufficient evidence at the \(5\%\) level that the mean score has increased above \(70\).
State the hypotheses (a two-tailed test):
| \(H_0\) | \(:\) | \(\mu=20\) |
| \(H_1\) | \(:\) | \(\mu\ne20\) |
Standardise:
| \(z\) | \(=\) | \(\dfrac{20.9-20}{5/\sqrt{100}}\) |
| \(=\) | \(\dfrac{0.9}{0.5}\) | |
| \(=\) | \(1.8\) |
Two-tailed p-value \(2\bigl(1-\Phi(|z|)\bigr)\):
| \(p\) | \(=\) | \(2\bigl(1-\Phi(1.8)\bigr)\) |
| \(=\) | \(2(1-0.9641)\) | |
| \(=\) | \(0.0719\) |
Compare with the \(1\%\) critical value \(2.576\):
| \(|z|\) | \(=\) | \(1.8<2.576\) |
| \(\Rightarrow\) | \(\text{do not reject } H_0\) |
Do not reject \(H_0\): there is insufficient evidence at the \(1\%\) level that the mean sugar content differs from \(20\) g.
Common pitfalls
Frequently asked questions
What is the test statistic for a hypothesis test on a mean?
It is \(z=\dfrac{\bar x-\mu_0}{\sigma/\sqrt n}\), the standardised sample mean under \(H_0\). It counts how many standard errors the sample mean \(\bar x\) sits from the hypothesised mean \(\mu_0\).
How do I know whether to use a one-tailed or two-tailed test?
Read the alternative hypothesis. "Differs from" or "changed" means \(H_1:\mu\ne\mu_0\) (two-tailed); "increased", "greater than", "decreased" or "less than" means \(H_1:\mu>\mu_0\) or \(H_1:\mu<\mu_0\) (one-tailed).
How is the p-value calculated?
It is the tail area beyond the observed \(z\). For a one-tailed test \(p=1-\Phi(|z|)\); for a two-tailed test \(p=2\bigl(1-\Phi(|z|)\bigr)\), where \(\Phi\) is the standard normal cdf.
What are the critical values I need to remember?
For a \(5\%\) level: \(1.645\) one-tailed, \(1.96\) two-tailed. For a \(1\%\) level: \(2.326\) one-tailed, \(2.576\) two-tailed. Reject \(H_0\) when \(|z|\) exceeds the value.
When do I reject the null hypothesis?
Reject \(H_0\) when the p-value is less than the level of significance \(\alpha\), or equivalently when \(|z|\) exceeds the critical value. Both routes always give the same decision.
Does rejecting H0 prove the alternative is true?
No. Rejecting \(H_0\) means there is sufficient evidence, at the chosen level, to support \(H_1\). It is evidence, not proof; and not rejecting \(H_0\) does not prove \(H_0\) is true.