Chalk−1

Math · Introductory statistics · Concept

The chi-square goodness-of-fit test

A goodness-of-fit test asks whether observed counts in categories match a claimed distribution, such as a fair die or a 3:1 genetic ratio. Each category contributes (observed − expected)²/expected, and the total χ² follows a chi-square distribution with k − 1 degrees of freedom when the model is right. A small p-value means the counts fit the model poorly.

Expected counts

Multiply the total count by each category’s claimed proportion. Expected counts need not be whole numbers.

Ei=npi

The statistic

Each term measures one category’s mismatch relative to its expected size. χ² is never negative, and it is 0 only when every count matches exactly.

χ2=∑(Oi−Ei)2Ei

Degrees of freedom and the p-value

With k categories, df = k − 1: once k − 1 counts are known, the total fixes the last. The p-value is the right-tail area beyond χ² under the chi-square curve.

Conditions

The counts must come from independent observations, and every expected count should be at least 5 for the chi-square approximation to hold; combine sparse categories if needed.

What a result means

A small p-value says the model does not fit. A large p-value says the data are consistent with the model, not that the model is proven.

Common mistakes

  • Using percentages or proportions instead of counts.
  • Dividing by the observed count instead of the expected count.
  • Using k instead of k − 1 degrees of freedom.
  • Running the test with expected counts below 5.

Key terms

Chi-square goodness-of-fit test
A test of whether observed counts in categories match the counts a model expects, using χ² = Σ(observed − expected)²/expected. It needs expected counts that aren’t too small, usually at least 5 each.
Chi-square distribution
A right-skewed distribution of nonnegative values, set by its degrees of freedom. It is used for chi-square tests on counts and for inference about variances.
Expected count
The count a model predicts on average, such as total × category probability. It doesn’t have to be a whole number, and a real sample will usually differ from it.
Degrees of freedom
The number of values free to vary once estimates have been fixed, such as n − 1 for one sample’s standard deviation. It sets the shape of the t, chi-square and F distributions.
p-value
The probability, assuming the null hypothesis is true, of a result at least as extreme as the one observed. A small p-value is evidence against the null; it is not the probability that the null is true.
Null hypothesis
H₀, the claim a test assumes so it can compute probabilities, usually “no effect” or a specific value such as μ = 50. Failing to reject it doesn’t prove it true.

Work through an example

A die is rolled 120 times, giving 25 ones, 17 twos, 15 threes, 23 fours, 24 fives and 16 sixes. Is this consistent with a fair die at α = 0.05?

Test whether a die is fair →

Test a 3:1 genetic ratio →

Sources and scope

Authored study material. Tool results depend on the stated inputs and model assumptions.

Make it concrete

Try in the workspace

Open the example inputs, change a value and keep a useful result on your board.

Find the p-value in Statistics Check χ² in Math Open worked example on a board Statistics formulas in Math Reference

Your existing work stays on this device. Examples open as editable copies.