Predictions: Interpolation & Extrapolation
Learn how to make predictions from a line of best fit for NSW Year 12 Mathematics Standard 2. You substitute into the least-squares regression line to estimate one variable from another, then decide whether the prediction can be trusted.
This topic covers interpolation (predicting inside the range of the data, which is more reliable) and extrapolation (predicting outside the range, which is less reliable), and explains the limitations of extrapolating a linear trend beyond the collected data β a key Standard 2 skill for bivariate data analysis.
Theory
Making predictions from a line of best fit is a core Year 12 Standard 2 (NSW) skill. This guide shows how to substitute into the least-squares regression line to predict a value, solve for the other variable, and classify each prediction as interpolation (inside the data, more reliable) or extrapolation (outside the data, less reliable).
A prediction uses a line of best fit β or the least-squares regression line β to estimate one variable from another. Substitute a known \(x\) to predict \(y\), or substitute a known \(y\) and solve for \(x\).
A prediction is interpolation when the value lies inside the range of the collected data. Because the trend is supported by evidence, interpolation is more reliable.
A prediction is extrapolation when the value lies outside the data range. There is no evidence the linear trend continues there, so extrapolation is less reliable β the key limitation to watch for in Year 12 Standard 2 (NSW).
The line of best fit / least-squares regression line has the form:
To predict \(y\), substitute the known \(x\). To predict \(x\) from a given \(y\), rearrange:
How to make and judge a prediction
- Write down the line \(y=mx+c\) and the range of the collected data (smallest and largest values).
- Substitute the known value: put in \(x\) to find \(y\), or put in \(y\) and solve \(x=\dfrac{y-c}{m}\).
- Compare the value with the smallest and largest data values.
- Classify: inside the range β interpolation (reliable); outside β extrapolation (less reliable). If extrapolating, note that the trend may not continue.
Substitute \(d=10\), then check the value against the data range.
| \(F\) | \(=\) | \(1.2d+4.5\) |
| \(F\) | \(=\) | \(1.2(10)+4.5\) |
| \(F\) | \(=\) | \(16.5\) |
| \(1\le 10\) | \(\le\) | \(25\ \Rightarrow\ \text{interpolation}\) |
The predicted fare is \(\$16.50\) β a reliable interpolation.
Substitute \(B=141\) and solve the linear equation for \(u\).
| \(141\) | \(=\) | \(0.28u+15\) |
| \(126\) | \(=\) | \(0.28u\) |
| \(u\) | \(=\) | \(\dfrac{126}{0.28}=450\) |
| \(100\le 450\) | \(\le\) | \(900\ \Rightarrow\ \text{interpolation}\) |
The predicted usage is \(450\) kWh.
Substitute \(w=16\), then compare it with the data range.
| \(H\) | \(=\) | \(2.5w+8\) |
| \(H\) | \(=\) | \(2.5(16)+8\) |
| \(H\) | \(=\) | \(48\) |
| \(16\) | \(>\) | \(10\ \Rightarrow\ \text{extrapolation}\) |
As \(16>10\), this is extrapolation. A sunflower's growth slows as it matures, so the linear trend may not continue β the prediction of \(48\) cm may be unreliable.
Substitute each temperature, then compare with the \(5\)β\(20^{\circ}\text{C}\) data range.
| \(S\) | \(=\) | \(-4T+120\) |
| \(T=12:\ S\) | \(=\) | \(-4(12)+120=72\) |
| \(T=28:\ S\) | \(=\) | \(-4(28)+120=8\) |
\(5\le 12\le 20\Rightarrow\) interpolation (reliable). \(28>20\Rightarrow\) extrapolation (less reliable; beyond \(T=30^{\circ}\text{C}\) the line even predicts negative sales).
Common pitfalls
Frequently asked questions
What is the difference between interpolation and extrapolation?
Interpolation is making a prediction for a value inside the range of the collected data, while extrapolation is making a prediction for a value outside that range. Interpolation is more reliable because the trend is supported by the data; extrapolation is less reliable because there is no evidence the trend continues.
How do you use a line of best fit to make a prediction?
Substitute the known value into the equation of the line. To predict y, put in the x-value and evaluate. To predict x, put in the y-value and solve the linear equation, which rearranges to x equals (y minus c) divided by m.
Why is extrapolation less reliable than interpolation?
The line of best fit is only fitted to the data that was collected. Outside that range there is no evidence the linear relationship still holds; the real trend may curve, level off or stop, so an extrapolated prediction can be well off. That is the main limitation of extrapolation.
Is extrapolation always wrong?
No. Extrapolation just means predicting outside the data range, and the result can still be close. But because it is not backed by evidence, you should treat it with caution and say the prediction may be unreliable, especially far beyond the data or where it gives an impossible value such as a negative amount.
How do you decide if a prediction is interpolation or extrapolation?
Find the smallest and largest values of the data. If the value you are predicting for lies between them, the prediction is interpolation. If it lies below the smallest or above the largest, it is extrapolation.