Math · Introductory statistics · Concept
Linear regression, correlation and residuals
Linear regression fits the line of best fit: the line that makes the squared vertical distances from the points, the residuals, as small as possible. The slope estimates how much y changes per unit of x. The correlation r, between −1 and 1, measures the strength and direction of a linear pattern, and r² gives the fraction of the variation in y that the line accounts for.
The least-squares line
Among all straight lines, the least-squares line has the smallest sum of squared residuals. Its slope and intercept come from the deviations from the means, and the line always passes through the point (x̄, ȳ).
Reading the slope and intercept
The slope b is the predicted change in y for a one-unit increase in x. The intercept a is the prediction at x = 0, which means something only when x = 0 is inside or near the data.
Residuals
A residual is observed minus predicted, y − ŷ, so points above the line have positive residuals. The residuals of a least-squares line add to zero. A curved or fan-shaped pattern in them means a straight line is the wrong model.
Correlation
r ranges from −1 to 1. Its sign matches the slope, values near ±1 mean the points lie close to a line, and values near 0 mean no linear pattern, though there may be a curved one. r has no units and stays the same if x and y swap roles.
r squared
r² is the fraction of the variation in y about its mean that the line explains: one minus the ratio of the residual sum of squares to the total sum of squares. An r² of 0.98 means the line accounts for 98% of that variation, not that predictions are 98% accurate.
Prediction and extrapolation
Use the line inside the range of the data. Beyond it the pattern may change, and extrapolation can predict impossible values.
Correlation is not causation
A strong r shows association. A lurking variable, or chance in a small sample, can produce a correlation without any causal link; a designed experiment is what supports a causal claim.
Curved data
When y grows by a roughly constant factor per step, an exponential model y = a·e^(bx) fits better than a line. Chalk Inverse fits exponential and power models by least squares on the original scale, so its coefficients can differ slightly from a straight-line fit to ln y.
Common mistakes
- Reading r² as the share of predictions that are right.
- Extrapolating far outside the data.
- Treating a strong correlation as proof of cause and effect.
- Computing a residual as predicted minus observed: it is observed minus predicted.
Key terms
- Linear regression
- Fitting a straight line y = a + bx to paired data, usually by least squares. In statistics “linear” means linear in the fitted coefficients, so a polynomial fit counts too.
- Least squares
- Choosing the fit that makes the sum of the squared residuals as small as possible. Squaring stops positive and negative misses from canceling and weights big misses more.
- Residual
- Observed value minus predicted value, y − ŷ. A positive residual means the point lies above the fitted line.
- Residual plot
- A plot of the residuals against x or the predicted values. A random scatter suits the model; a curve or a fan shape shows the model is missing something.
- Pearson correlation
- r, a number from −1 to 1 measuring how closely paired data follow a straight line. r = 0 doesn’t rule out a curved relationship, and correlation doesn’t prove causation.
- Coefficient of determination
- R², the share of the variation in y that a model explains: 1 − SSE/SST. For a straight-line fit with an intercept it equals r²; for other models it can even be negative.
- Extrapolation
- Using a fitted model outside the range of the data. The formula still gives numbers, but the real relationship there may be different.
- Exponential regression
- Fitting a curve y = A·e^(bx) to data, where each equal step in x multiplies y by the same factor. This app fits the curve to y directly, which can differ from fitting a line to ln y.
Work through an example
Five students report hours studied and exam scores: (1, 55), (2, 62), (3, 64), (4, 71) and (5, 78). Find the least-squares line, r and r², and predict the score after 3.5 hours.
Fit a least-squares line and interpret r² →Sources and scope
Authored study material. Tool results depend on the stated inputs and model assumptions.
Try in the workspace
Open the example inputs, change a value and keep a useful result on your board.
Fit the line in Statistics Check the line in Math Open worked example on a board Statistics formulas in Math ReferenceYour existing work stays on this device. Examples open as editable copies.