Resources For Teachers For Tutors For Students & Parents Pricing
Year 12 Maths - Specialist (Unit 3 & Unit 4) Statistical inference

Hypothesis testing for a population mean

20 practice questions 0 video lessons Theory + worked examples

Master hypothesis testing for a population mean in Year 12 VCE Specialist Mathematics. A hypothesis test uses a single sample to decide between a null hypothesis \(H_0:\mu=\mu_0\) and an alternative \(H_1\), standardising the sample mean to the test statistic \(z=\dfrac{\bar x-\mu_0}{\sigma/\sqrt n}\). It sits in the Data analysis, probability and statistics area of study of the VCE Mathematics Study Design (VCAA), within the Statistical inference topic of Unit 4.

You will learn to state the hypotheses, choose a one- or two-tailed test, compute the test statistic and a p-value or critical value, then decide at a chosen level of significance \(\alpha\) and interpret the result as evidence in context — never as proof.

Create a free accountTrack your progress and save your work as you go.
Create free account

Theory

A hypothesis test for a population mean uses a single sample to decide between a null hypothesis \(H_0:\mu=\mu_0\) and an alternative \(H_1\). In Year 12 Specialist Mathematics you standardise the sample mean to the test statistic \(z=\dfrac{\bar x-\mu_0}{\sigma/\sqrt n}\), find a p-value or a critical value, then decide at a chosen level of significance \(\alpha\) and interpret the result in context.

A hypothesis test weighs the evidence a sample gives against a claim about the population mean \(\mu\). The null hypothesis \(H_0:\mu=\mu_0\) states that the mean equals a specified value; the alternative hypothesis \(H_1\) is what you test for.

The alternative sets the number of tails. A two-tailed test \(H_1:\mu\ne\mu_0\) asks whether the mean has changed in either direction; a one-tailed test \(H_1:\mu>\mu_0\) or \(H_1:\mu<\mu_0\) asks whether it has changed in one specific direction.

The test statistic standardises the sample mean under \(H_0\): \(z=\dfrac{\bar x-\mu_0}{\sigma/\sqrt n}\), where \(\sigma\) is the known population standard deviation and \(n\) the sample size. It measures how many standard errors \(\bar x\) sits from \(\mu_0\). (For a large sample, \(n\ge 30\), the sample standard deviation \(s\) may replace \(\sigma\).)

The level of significance \(\alpha\) (commonly \(0.05\) or \(0.01\)) is the risk you accept of rejecting a true \(H_0\). The p-value is the probability, if \(H_0\) were true, of a test statistic at least as extreme as the one observed. You reject \(H_0\) when \(p<\alpha\) (equivalently, when \(|z|\) exceeds the critical value); otherwise you do not reject \(H_0\).

Interpretation is evidence, not proof. Rejecting \(H_0\) means the sample gives sufficient evidence to reject it at the \(\alpha\) level — it never proves \(H_1\). Failing to reject \(H_0\) means there is insufficient evidence against it, not that \(H_0\) is true.

Two-tailed test rejection region at the 5% level A standard normal curve. Both tails, beyond z equals minus 1.96 and plus 1.96, are shaded as the rejection region for a two-tailed 5% test. A navy line marks the test statistic z equals 2.5, which lies inside the right tail, so the null hypothesis is rejected. z -1.96 1.96 z = 2.5
Two-tailed test at the \(5\%\) level: reject \(H_0\) when \(|z|>1.96\). Here the test statistic \(z=2.5\) falls in the right rejection tail, so \(H_0\) is rejected.
One-tailed (upper) test rejection region at the 5% level A standard normal curve. The right tail beyond z equals 1.645 is shaded as the rejection region for an upper one-tailed 5% test. A navy line marks the test statistic z equals 2.5, which lies inside the shaded tail, so the null hypothesis is rejected. z 1.645 z = 2.5
One-tailed (upper) test at the \(5\%\) level: reject \(H_0\) when \(z>1.645\). The single shaded tail holds all \(5\%\); here \(z=2.5\) lands inside it.

The standardised test statistic for a sample of size \(n\) with mean \(\bar x\), drawn from a population with mean \(\mu_0\) under \(H_0\) and known standard deviation \(\sigma\):

\[ z=\dfrac{\bar x-\mu_0}{\sigma/\sqrt n} \]
z=x¯μ0σ/n

The p-value is the tail area beyond the observed \(z\), using the standard normal cdf \(\Phi\). For a one-tailed test take one tail; for a two-tailed test double it:

\[ p_{\text{one-tail}}=1-\Phi(|z|),\qquad p_{\text{two-tail}}=2\bigl(1-\Phi(|z|)\bigr) \]
p=2(1Φ(|z|))

Decision. Reject \(H_0\) when \(p<\alpha\), or equivalently when \(|z|\) exceeds the critical value for the test:

Test\(5\%\) level\(1\%\) level
One-tailed critical value\(1.645\)\(2.326\)
Two-tailed critical value\(1.96\)\(2.576\)
Two-tailed splits the significance in half. A two-tailed \(5\%\) test puts \(2.5\%\) in each tail, so the cut-off \(1.96\) is larger than the one-tailed \(1.645\). The p-value route and the critical-value route always give the same decision.

How to run a hypothesis test for the mean

  1. State the hypotheses. Write \(H_0:\mu=\mu_0\) and choose \(H_1\) from the context: \(\ne\) for a change either way (two-tailed), \(>\) or \(<\) for a specific direction (one-tailed).
  2. Standardise. Compute the test statistic \(z=\dfrac{\bar x-\mu_0}{\sigma/\sqrt n}\) using the sample mean \(\bar x\), the value \(\mu_0\), the known \(\sigma\) and the size \(n\).
  3. Measure the evidence. Find the p-value — \(1-\Phi(|z|)\) for one tail, \(2\bigl(1-\Phi(|z|)\bigr)\) for two — or read off the critical value for the level \(\alpha\).
  4. Decide. Reject \(H_0\) if \(p<\alpha\) (equivalently \(|z|\) beyond the critical value); otherwise do not reject \(H_0\).
  5. Interpret in context. State the conclusion as evidence: e.g. "sufficient evidence at the \(5\%\) level that the mean differs from \(\mu_0\)" — never "this proves \(H_1\)".
Example 1 — A two-tailed test
A coffee machine dispenses cups whose volume is normally distributed with known standard deviation \(\sigma=8\) mL, set to a mean of \(\mu_0=250\) mL. A random sample of \(16\) cups has mean \(\bar x=246\) mL. Test, at the \(5\%\) level, whether the mean volume differs from \(250\) mL.
Solution

State the hypotheses (a two-tailed test):

\(H_0\)\(:\)\(\mu=250\)
\(H_1\)\(:\)\(\mu\ne250\)

Standardise \(z=\dfrac{\bar x-\mu_0}{\sigma/\sqrt n}\):

\(z\)\(=\)\(\dfrac{246-250}{8/\sqrt{16}}\)
\(=\)\(\dfrac{-4}{2}\)
\(=\)\(-2\)

Two-tailed p-value \(2\bigl(1-\Phi(|z|)\bigr)\):

\(p\)\(=\)\(2\bigl(1-\Phi(2)\bigr)\)
\(=\)\(2(1-0.9772)\)
\(=\)\(0.0455\)

Compare and decide at \(5\%\):

\(|z|\)\(=\)\(2>1.96\)
\(\Rightarrow\)\(\text{reject } H_0\)

Reject \(H_0\): there is sufficient evidence at the \(5\%\) level that the mean volume differs from \(250\) mL.

Two-tailed 5% test with z = -2 in the left rejection tail A standard normal curve with both tails beyond plus and minus 1.96 shaded. A navy line marks the test statistic z equals minus 2, which lies inside the left rejection tail, so the null hypothesis is rejected. z -1.96 1.96 z = -2
Example 2 — A one-tailed (lower) test
A supplier claims the mean breaking strength of a cable is at least \(\mu_0=500\) N, with strengths normally distributed and known \(\sigma=30\) N. Suspecting the mean is lower, an engineer tests a random sample of \(36\) cables and finds \(\bar x=490\) N. Test at the \(5\%\) level.
Solution

State the hypotheses (a one-tailed, lower test):

\(H_0\)\(:\)\(\mu=500\)
\(H_1\)\(:\)\(\mu<500\)

Standardise:

\(z\)\(=\)\(\dfrac{490-500}{30/\sqrt{36}}\)
\(=\)\(\dfrac{-10}{5}\)
\(=\)\(-2\)

One-tailed p-value \(1-\Phi(|z|)\):

\(p\)\(=\)\(1-\Phi(2)\)
\(=\)\(1-0.9772\)
\(=\)\(0.0228\)

Compare with the \(5\%\) critical value \(1.645\):

\(|z|\)\(=\)\(2>1.645\)
\(\Rightarrow\)\(\text{reject } H_0\)

Reject \(H_0\): there is sufficient evidence at the \(5\%\) level that the mean breaking strength is less than \(500\) N.

Example 3 — A one-tailed (upper) test
A training program claims to raise the mean score above \(\mu_0=70\); scores are normally distributed with known \(\sigma=12\). A random sample of \(36\) trained students has mean \(\bar x=75\). Test, at the \(5\%\) level, whether the mean score has increased.
Solution

State the hypotheses (a one-tailed, upper test):

\(H_0\)\(:\)\(\mu=70\)
\(H_1\)\(:\)\(\mu>70\)

Standardise:

\(z\)\(=\)\(\dfrac{75-70}{12/\sqrt{36}}\)
\(=\)\(\dfrac{5}{2}\)
\(=\)\(2.5\)

One-tailed p-value \(1-\Phi(|z|)\):

\(p\)\(=\)\(1-\Phi(2.5)\)
\(=\)\(1-0.9938\)
\(=\)\(0.0062\)

Compare with the \(5\%\) critical value \(1.645\):

\(z\)\(=\)\(2.5>1.645\)
\(\Rightarrow\)\(\text{reject } H_0\)

Reject \(H_0\): there is sufficient evidence at the \(5\%\) level that the mean score has increased above \(70\).

One-tailed (upper) test rejection region at the 5% level A standard normal curve. The right tail beyond z equals 1.645 is shaded as the rejection region for an upper one-tailed 5% test. A navy line marks the test statistic z equals 2.5, which lies inside the shaded tail, so the null hypothesis is rejected. z 1.645 z = 2.5
Example 4 — Do not reject at \(1\%\)
A food label states the mean sugar content is \(\mu_0=20\) g per serving. A consumer group tests a large random sample of \(100\) servings, finding \(\bar x=20.9\) g, with known \(\sigma=5\) g. Test, at the \(1\%\) level, whether the mean sugar content differs from \(20\) g.
Solution

State the hypotheses (a two-tailed test):

\(H_0\)\(:\)\(\mu=20\)
\(H_1\)\(:\)\(\mu\ne20\)

Standardise:

\(z\)\(=\)\(\dfrac{20.9-20}{5/\sqrt{100}}\)
\(=\)\(\dfrac{0.9}{0.5}\)
\(=\)\(1.8\)

Two-tailed p-value \(2\bigl(1-\Phi(|z|)\bigr)\):

\(p\)\(=\)\(2\bigl(1-\Phi(1.8)\bigr)\)
\(=\)\(2(1-0.9641)\)
\(=\)\(0.0719\)

Compare with the \(1\%\) critical value \(2.576\):

\(|z|\)\(=\)\(1.8<2.576\)
\(\Rightarrow\)\(\text{do not reject } H_0\)

Do not reject \(H_0\): there is insufficient evidence at the \(1\%\) level that the mean sugar content differs from \(20\) g.

Common pitfalls

Dividing by \(n\) instead of \(\sqrt n\). The standard error is \(\dfrac{\sigma}{\sqrt n}\), not \(\dfrac{\sigma}{n}\). For \(n=16\) you divide \(\sigma\) by \(4\), not by \(16\).
Mixing up the critical values. Match the value to the test: \(1.645\) (one-tailed \(5\%\)), \(1.96\) (two-tailed \(5\%\)), \(2.326\) (one-tailed \(1\%\)), \(2.576\) (two-tailed \(1\%\)). A two-tailed test splits \(\alpha\) into two tails, so its cut-off is larger.
Forgetting to double for two tails. A two-tailed p-value is \(2\bigl(1-\Phi(|z|)\bigr)\). Using the one-tailed \(1-\Phi(|z|)\) halves the p-value and can flip your decision.
Saying the test "proves" \(H_1\). Rejecting \(H_0\) gives sufficient evidence for \(H_1\) at the chosen level — it is evidence, never proof. And failing to reject \(H_0\) is not proof that \(H_0\) is true.

Frequently asked questions

What is the test statistic for a hypothesis test on a mean?

It is \(z=\dfrac{\bar x-\mu_0}{\sigma/\sqrt n}\), the standardised sample mean under \(H_0\). It counts how many standard errors the sample mean \(\bar x\) sits from the hypothesised mean \(\mu_0\).

How do I know whether to use a one-tailed or two-tailed test?

Read the alternative hypothesis. "Differs from" or "changed" means \(H_1:\mu\ne\mu_0\) (two-tailed); "increased", "greater than", "decreased" or "less than" means \(H_1:\mu>\mu_0\) or \(H_1:\mu<\mu_0\) (one-tailed).

How is the p-value calculated?

It is the tail area beyond the observed \(z\). For a one-tailed test \(p=1-\Phi(|z|)\); for a two-tailed test \(p=2\bigl(1-\Phi(|z|)\bigr)\), where \(\Phi\) is the standard normal cdf.

What are the critical values I need to remember?

For a \(5\%\) level: \(1.645\) one-tailed, \(1.96\) two-tailed. For a \(1\%\) level: \(2.326\) one-tailed, \(2.576\) two-tailed. Reject \(H_0\) when \(|z|\) exceeds the value.

When do I reject the null hypothesis?

Reject \(H_0\) when the p-value is less than the level of significance \(\alpha\), or equivalently when \(|z|\) exceeds the critical value. Both routes always give the same decision.

Does rejecting H0 prove the alternative is true?

No. Rejecting \(H_0\) means there is sufficient evidence, at the chosen level, to support \(H_1\). It is evidence, not proof; and not rejecting \(H_0\) does not prove \(H_0\) is true.