Resources For Teachers For Tutors For Students & Parents Pricing
Year 12 Maths Standard 2 (2027) Bivariate data analysis

Pearson's Correlation Coefficient

20 practice questions 0 video lessons Theory + worked examples

Master Pearson's correlation coefficient \(r\) for NSW Year 12 Mathematics Standard 2. You will learn that \(r\) is a single number between \(-1\) and \(1\) whose sign gives the direction of a linear correlation and whose size gives its strength — strong, moderate or weak.

This topic shows you how to read \(r\) from a scientific calculator's statistics mode, estimate it from a scatterplot, classify a correlation as positive or negative and strong, moderate or weak, and explain why a strong correlation does not prove causation — core Standard 2 skills for the bivariate data section of the HSC.

Create a free accountTrack your progress and save your work as you go.
Create free account

Theory

Pearson's correlation coefficient \(r\) measures how strongly two numerical variables are linearly related. This Year 12 Standard 2 (NSW) guide shows how to find \(r\) with a scientific calculator, read its sign as the direction and its size as the strength of the correlation, estimate \(r\) from a scatterplot, and why correlation does not prove causation.

Pearson's correlation coefficient \(r\) is a single number that measures the strength and direction of the linear relationship between two numerical variables. It always lies between \(-1\) and \(1\).

The sign of \(r\) gives the direction: \(r>0\) is a positive correlation (both variables tend to rise together) and \(r<0\) is a negative correlation (one rises as the other falls). The size of \(|r|\) gives the strength: values near \(1\) are strong, values near \(0\) are weak.

\(r=1\) or \(r=-1\) means the points lie exactly on a straight line (a perfect correlation), while \(r=0\) means there is no linear correlation. In this Year 12 Standard 2 (NSW) course you read \(r\) straight from a scientific calculator's statistics mode.

Strong positive correlationSeven points rising from lower left to upper right, close to a line, r about 0.95 x y 1 2 3 4 5 6 7 2 4 6 8
Points rise together and hug the line \(\Rightarrow\) strong positive, \(r\approx 0.95\).
Strong negative correlationSeven points falling from upper left to lower right, close to a line, r about -0.95 x y 1 2 3 4 5 6 7 2 4 6 8
Points fall as \(x\) rises \(\Rightarrow\) strong negative, \(r\approx -0.95\).

From a data set, the calculator computes \(r\) as

\[r = \frac{S_{xy}}{\sqrt{S_{xx}\,S_{yy}}}\]
r=SxySxxSyy

where \(S_{xy}\) measures how \(x\) and \(y\) vary together and \(S_{xx}, S_{yy}\) measure how each varies on its own. You never compute this by hand in Standard 2 — you enter the pairs and read \(r\) — but it shows why \(r\) is a pure number with

\[-1 \le r \le 1\]
-1r1
Strength bands. \(|r|\ge 0.75\) is strong, \(0.5\le|r|<0.75\) is moderate, and \(|r|<0.5\) is weak. The sign is read separately as the direction.

How to find and interpret \(r\)

  1. Enter the paired data into your calculator's statistics (or regression) mode as \((x,y)\) pairs.
  2. Read off the value labelled \(r\), rounded to \(2\) decimal places.
  3. Direction from the sign: \(r>0\) positive, \(r<0\) negative.
  4. Strength from the size: \(|r|\ge 0.75\) strong, \(0.5\le|r|<0.75\) moderate, \(|r|<0.5\) weak.
  5. Interpret in context, and remember a correlation does not prove causation.
Example 1 — Estimate r from a scatterplot
A swimming coach records each swimmer's weekly training hours \(t\) and the laps completed in a time trial \(L\). Estimate \(r\) and describe the correlation.
Solution

Read the direction from the trend, then the strength from how close the points are to a line.

Example 1 scatterplotSeven points rising from left to right, close to a line t L 1 2 3 4 5 6 7 2 4 6
\(\text{trend rises}\)\(\Rightarrow\)\(r>0\ (\text{positive})\)
\(\text{close to a line}\)\(\Rightarrow\)\(\text{strong}\)
\(r\)\(\approx\)\(0.9\)

A strong positive correlation: more training hours go with more laps.

Example 2 — Compute r from a table
Find \(r\) (to \(2\) d.p.) for the data below with your calculator, then describe the correlation.
Solution

Enter the five pairs, read \(r\), then classify the sign and size.

\(x\)12345
\(y\)65869
Example 2 scatterplotFive points drifting upward with noticeable scatter x y 1 2 3 4 5 2 4 6 8
\(r\)\(=\)\(\dfrac{S_{xy}}{\sqrt{S_{xx}\,S_{yy}}}\)
\(r\)\(=\)\(\dfrac{7}{\sqrt{10\times 10.8}}=\dfrac{7}{\sqrt{108}}\)
\(r\)\(\approx\)\(0.67\)

Since \(0.5\le 0.67<0.75\) and \(r>0\), it is a moderate positive correlation.

Example 3 — Strength vs sign
The scatterplot shows daily social-media time \(s\) (hours) and sleep \(h\) (hours) for seven students. Describe the correlation, and say whether it or a group with \(r=0.7\) is stronger.
Solution

Direction from the falling trend; strength from \(|r|\), compared as sizes.

Example 3 scatterplotSeven points falling from left to right, close to a line s h 1 2 3 4 5 6 7 2 4 6 8
\(\text{trend falls}\)\(\Rightarrow\)\(r<0\)
\(\text{close to a line}\)\(\Rightarrow\)\(\text{strong}\)
\(r\)\(\approx\)\(-0.89\)
\(|-0.89|=0.89\)\(>\)\(0.7\)

A strong negative correlation; and \(0.89>0.7\), so this group is the stronger of the two.

Example 4 — Correlation is not causation
Over six hot days a kiosk records ice-cream sales \(I\) (hundreds) and sunscreen bottles sold \(S\) (tens). (i) Find \(r\). (ii) Describe it. (iii) Does this prove ice-cream sales cause sunscreen sales?
Solution

Read \(r\) from the calculator, classify it, then judge the causation claim.

\(I\)123456
\(S\)436689
Example 4 scatterplotSix points rising steadily from left to right I S 1 2 3 4 5 6 2 4 6 8
\(r\)\(\approx\)\(0.94\)
\(r>0,\ |r|=0.94\)\(\ge\)\(0.75\)
\(\Rightarrow\)\(\)\(\text{strong positive}\)

(iii) No — correlation does not prove causation. Hot weather lifts both sales, so it is the common cause.

Common pitfalls

\(r\) only measures linear patterns. A strong curved relationship can still give \(r\approx 0\); a low \(r\) means no straight-line trend, not "no relationship at all".
\(r\) is not the gradient. The line of best fit can have any gradient, but \(r\) must stay in \([-1,1]\). They share a sign, nothing more — a gradient of \(2\) does not mean \(r=2\).
Strength is about \(|r|\), not the sign. \(r=-0.9\) is a stronger correlation than \(r=0.6\); a negative \(r\) is not "weaker" than a positive one.
Correlation is not causation. A high \(r\) shows two variables move together, not that one causes the other — a lurking third variable may drive both.

Frequently asked questions

What does Pearson's correlation coefficient r tell you?

It tells you two things about the linear relationship between two numerical variables: the direction (its sign, positive or negative) and the strength (how close its size is to 1). It is always a number between -1 and 1.

What value of r counts as a strong correlation?

A common guide is that |r| of 0.75 or more is strong, between 0.5 and 0.75 is moderate, and below 0.5 is weak. The sign is read separately: it only tells you the direction, not the strength.

Can the correlation coefficient be greater than 1?

No. Pearson's r can never be more than 1 or less than -1, so a value like 1.4 is impossible and usually means a calculation or data-entry error. The extremes r = 1 and r = -1 are perfect straight-line relationships.

Is r the same as the gradient of the line of best fit?

No. The gradient can be any size and carries the units of the data, while r is a pure number in the range -1 to 1. They always share the same sign, but a gradient of 2 does not mean r = 2.

Does a high correlation mean one variable causes the other?

No. A high r shows the variables tend to change together, but correlation does not prove causation. A third factor can drive both variables, so you cannot conclude cause and effect from r alone.

How do you find r on a calculator?

Put the calculator into statistics or regression mode, enter the data as (x, y) pairs, then read off the value labelled r. Round it to two decimal places and interpret its sign and size in context.