STAT 350 / WORKSHEET 10

Build a Normal QQ-plot

n = x̄ = s = IQR/s =
Choose a population pattern:

1 · Divide 0–1 into equal probability intervals

Use the midpoint, not the right endpoint. Neither 0 nor 1 is used.

2 · Transform probabilities into Normal quantiles

qnorm(p)

Equal probability spacing becomes unequal z-spacing, especially in the tails.

3 · Calculate the points on the reference line

y = x̄ + sz

The line uses transformed Normal quantiles—not the observed mpg values.

4 · Pair zᵢ with the ith smallest original value

Do not transform the data

Normal QQ-plot

Observed (zᵢ, x₍ᵢ₎)Reference (zᵢ, x̄ + szᵢ)y = x̄ + sz

Histogram with reference curves

HistogramSmooth curveFitted Normal curve

Original observations → ordered values → QQ coordinates

Rank iOriginal rowObservationx₍ᵢ₎pᵢzᵢx̄ + szᵢDifference

Click a row to highlight that observation in the demonstration. Ties keep their original row order; they are still separate observations.

Histogram and curve settings

These controls change the display, not the observations or their QQ coordinates.

8
1.00×

Less smoothing follows small bumps more closely. More smoothing blends nearby features. Neither setting is a test of Normality.

Reading the construction and the patterns

Two different vertical values at the same horizontal position

For rank i, divide the probability range [0,1] into n intervals of width 1/n. The ith interval runs from (i−1)/n to i/n, and its midpoint is pᵢ = (i−0.5)/n. Transform that midpoint using zᵢ = Φ⁻¹(pᵢ). This zᵢ is the horizontal coordinate.

Reference point: (zᵢ, x̄ + szᵢ). Transform the Normal quantile using the sample mean and standard deviation. These points lie exactly on the line y = x̄ + sz.
Observed point: (zᵢ, x₍ᵢ₎). Use the ith smallest original value as the vertical coordinate. Do not replace that value with x̄ + szᵢ, and do not standardize it in this version of the plot.

The vertical difference is x₍ᵢ₎ − (x̄ + szᵢ). A positive difference places the observed point above the line. Even a sample generated from a Normal population will not put every point on the line.

The midpoint convention avoids infinite endpoint quantiles. It is not the unique plotting-position formula and does not give the exact expected Normal order statistics. The straight reference is a mean/SD line, not a least-squares regression line.

Present without scrolling

Choose a rank or click a probability, quantile, ordered value, or QQ point. Use steps 1–4 to reveal the construction; “Build one point” plays those stages. “Sweep ranks” moves through the observations. Left/right arrow keys change rank; Home/End select an endpoint; Escape stops playback. On narrow viewers, the selected construction step is shown above a QQ/histogram tab switch.

All descriptions here use theoretical Normal quantiles horizontally and ordered data vertically. Focus on broad patterns across ranks, not one point. These are diagnostic clues, not one-to-one rules or tests.

Normal: approximately straight without sustained curvature. The fitted mean/SD line need not be a 45-degree line.
Right skew: broad upward curvature; the upper ordered values grow more rapidly. A long right tail should also be visible in the histogram. Left skew mirrors this with downward curvature.
Heavier tails: lower tail tends below the reference and upper tail above it. Both ends are more extreme than the Normal reference. Exact crossings depend on the fitted line and the finite sample.
Lighter tails: lower tail tends above and upper tail below. The endpoints are less extreme; bounded Uniform data are an example.
Two clusters: a central gap or changing slopes can appear. Inspect the histogram before calling this multimodality; the QQ-plot alone does not identify the number of modes.
An outlier: an isolated endpoint may separate from neighboring points. The outlier also changes x̄ and s, so even the other points may shift relative to the mean/SD line.

Use “New sample” repeatedly at a fixed n, then change n. A small sample can hide population features or create apparent ones. The population names describe the simulation mechanism, not a verdict from the plotted sample.

Backward Empirical Rule

IQR divided by sample standard deviation

The exact Normal population benchmark is approximately 1.34898 (about 1.35; the worksheet uses 1.34). Sample proportions and sample ratios fluctuate, and a match is not proof of Normality. Inclusive endpoints are used for the observed counts.

mtcars

The built-in mpg values show mild right-skewed structure rather than a perfect bell shape. With only 32 values and rounded, repeated measurements, avoid a definitive population claim based on these displays. A marginal mpg plot also does not check regression errors or independence.

Data and conventions

The mtcars option uses the 32 mpg values and car names supplied for Worksheet 10, in their original row order. The full calculation table identifies each original row. Other options are simulated illustrations; their units are not MPG. The seeded browser generator is reproducible within this page, not identical to R’s random generator.

All QQ points use (i−0.5)/n. Sample SD uses n−1, and sample quartiles use R’s type-7 interpolation. The smooth curve is a data-driven estimate; it is not a fitted Normal density. The fitted Normal curve and mean/SD QQ line both use the same x̄ and s. The histogram uses density = count/(n × bin width).

R’s default ppoints() uses the midpoint convention for n > 10 and a different offset for smaller samples. Here we keep the worksheet’s midpoint convention for every n. Numerical inverse-Normal calculations use a rational approximation, checked against reference values. The page runs locally with no external scripts or network requests.

Reference documentation

R: plotting positions · R: Normal functions · R: sample quantiles · ggplot2: QQ points and lines

Only these optional documentation links use the internet. All data, drawing code, and controls are embedded in this file.