Chalk−1

Math · Introductory statistics · Concept

Linear regression, correlation and residuals

Linear regression fits the line of best fit: the line that makes the squared vertical distances from the points, the residuals, as small as possible. The slope estimates how much y changes per unit of x. The correlation r, between −1 and 1, measures the strength and direction of a linear pattern, and r² gives the fraction of the variation in y that the line accounts for.

The least-squares line

Among all straight lines, the least-squares line has the smallest sum of squared residuals. Its slope and intercept come from the deviations from the means, and the line always passes through the point (x̄, ȳ).

b=∑(xi−x¯)(yi−y¯)∑(xi−x¯)2,a=y¯−bx¯
b=∑(xi−x¯)(yi−y¯)∑(xi−x¯)2a=y¯−bx¯

Reading the slope and intercept

The slope b is the predicted change in y for a one-unit increase in x. The intercept a is the prediction at x = 0, which means something only when x = 0 is inside or near the data.

Residuals

A residual is observed minus predicted, y − ŷ, so points above the line have positive residuals. The residuals of a least-squares line add to zero. A curved or fan-shaped pattern in them means a straight line is the wrong model.

ei=yi−y^i

Correlation

r ranges from −1 to 1. Its sign matches the slope, values near ±1 mean the points lie close to a line, and values near 0 mean no linear pattern, though there may be a curved one. r has no units and stays the same if x and y swap roles.

r=∑(xi−x¯)(yi−y¯)∑(xi−x¯)2∑(yi−y¯)2
r=Sx⁢ySx⁢xSy⁢ySx⁢y=∑(xi−x¯)(yi−y¯)

r squared

r² is the fraction of the variation in y about its mean that the line explains: one minus the ratio of the residual sum of squares to the total sum of squares. An r² of 0.98 means the line accounts for 98% of that variation, not that predictions are 98% accurate.

r2=1−SSESST

Prediction and extrapolation

Use the line inside the range of the data. Beyond it the pattern may change, and extrapolation can predict impossible values.

Correlation is not causation

A strong r shows association. A lurking variable, or chance in a small sample, can produce a correlation without any causal link; a designed experiment is what supports a causal claim.

Curved data

When y grows by a roughly constant factor per step, an exponential model y = a·e^(bx) fits better than a line. Chalk Inverse fits exponential and power models by least squares on the original scale, so its coefficients can differ slightly from a straight-line fit to ln y.

Common mistakes

  • Reading r² as the share of predictions that are right.
  • Extrapolating far outside the data.
  • Treating a strong correlation as proof of cause and effect.
  • Computing a residual as predicted minus observed: it is observed minus predicted.

Key terms

Linear regression
Fitting a straight line y = a + bx to paired data, usually by least squares. In statistics “linear” means linear in the fitted coefficients, so a polynomial fit counts too.
Least squares
Choosing the fit that makes the sum of the squared residuals as small as possible. Squaring stops positive and negative misses from canceling and weights big misses more.
Residual
Observed value minus predicted value, y − ŷ. A positive residual means the point lies above the fitted line.
Residual plot
A plot of the residuals against x or the predicted values. A random scatter suits the model; a curve or a fan shape shows the model is missing something.
Pearson correlation
r, a number from −1 to 1 measuring how closely paired data follow a straight line. r = 0 doesn’t rule out a curved relationship, and correlation doesn’t prove causation.
Coefficient of determination
R², the share of the variation in y that a model explains: 1 − SSE/SST. For a straight-line fit with an intercept it equals r²; for other models it can even be negative.
Extrapolation
Using a fitted model outside the range of the data. The formula still gives numbers, but the real relationship there may be different.
Exponential regression
Fitting a curve y = A·e^(bx) to data, where each equal step in x multiplies y by the same factor. This app fits the curve to y directly, which can differ from fitting a line to ln y.

Work through an example

Five students report hours studied and exam scores: (1, 55), (2, 62), (3, 64), (4, 71) and (5, 78). Find the least-squares line, r and r², and predict the score after 3.5 hours.

Fit a least-squares line and interpret r² →

Interpret a negative correlation →

Fit an exponential model to growth data →

Sources and scope

Authored study material. Tool results depend on the stated inputs and model assumptions.

Make it concrete

Try in the workspace

Open the example inputs, change a value and keep a useful result on your board.

Fit the line in Statistics Check the line in Math Open worked example on a board Statistics formulas in Math Reference

Your existing work stays on this device. Examples open as editable copies.