Scatterplots, Statistics, & Probability
SAT Math Prep · Problem Solving & Data AnalysisPreview
1. Introduction
This cluster of topics — scatterplots, summary statistics, and basic probability — is where the SAT tests whether you can read data, not just crunch it. The math is light: means, medians, simple fractions. The challenge is interpretation. You must extract the right number from a graph or two-way table, understand what a slope or standard deviation means in context, and avoid the carefully designed answer traps.
The good news is that the question types are extremely predictable. The SAT recycles the same handful of structures: read a value off a line of best fit, decide how an outlier moves the mean versus the median, compare the spread of two data sets, and compute a conditional probability from a table. Master these and you turn a whole section of the test into nearly free points.
Beyond the basics, the SAT also tests residuals (how far a point lies from the model), marginal distributions in two-way tables, and "at least one" probability via complements. Students who rush to compute without reading the question stem carefully — especially the phrases "given that" and "of the" — lose easy points every test.
This article builds genuine understanding of each statistic, then layers SAT-specific strategy on top — including which questions reward the calculator's statistics functions and which are faster done by eye.
What the SAT actually tests. Expect – questions per test touching this material: one scatterplot or line-of-best-fit item, one center/spread comparison, one mean or median computation, and one or two probability questions from a table. The difficulty is almost never arithmetic — it is choosing the correct statistic or the correct denominator. Students who score well here read carefully and compute minimally.
Study priority. Master conditional probability wording first (the highest-error subtopic), then the mean-versus-median outlier distinction, then line-of-best-fit slope interpretation. Standard deviation comparison by eye is a quick skill that pays off on nearly every practice test.
Diagnostic checklist. Before submitting: Did I use the correct denominator (grand total vs. row/column)? Did the question ask for mean or median? Is the slope interpretation stated as a rate (per unit), not a total? Did I account for the new count when computing a revised mean? For "at least one," did I use the complement?
2. Core Concepts
2.1 Scatterplots and Association
A scatterplot plots paired data as points. The overall pattern reveals the association:
- Positive association: as increases, tends to increase (points rise left to right).
- Negative association: as increases, tends to decrease.
- No association: no clear trend.
- Linear vs. nonlinear: points hug a straight line, or curve.
Strength describes how tightly the points cluster around the trend — tightly clustered is "strong," loosely scattered is "weak." The SAT rarely asks you to compute a correlation coefficient; instead you judge strength visually.
2.2 The Line of Best Fit
The line of best fit (least-squares regression line) is the straight line that best summarizes a linear trend. Two SAT-critical interpretations:
- The slope is the predicted change in for each one-unit increase in . Always state slope with units in context ("each additional hour of study predicts more points").
- The -intercept is the predicted value of when — sometimes meaningful, sometimes just a mathematical artifact.
You predict by substituting an -value into the equation. Interpolation (predicting inside the data range) is reliable; extrapolation (far outside it) is risky and the SAT often flags it.
2.3 Residuals
A residual for a data point is the difference between the actual -value and the predicted value from the model:
Positive residuals mean the point lies above the line; negative residuals mean below. The line of best fit minimizes the sum of squared residuals — you will not compute that on the SAT, but you may be asked which point has the largest residual or whether a point is above or below the model.
2.4 Measures of Center
- Mean: — the balance point; sensitive to every value.
- Median: the middle value when data is ordered (average the two middle values if is even) — resistant to outliers.
- Mode: the most frequent value — useful for categorical or repeated data.
The defining SAT fact: an outlier pulls the mean toward itself but barely moves the median. In a right-skewed set, mean median; left-skewed, mean median.
2.5 Measures of Spread
- Range maximum minimum. Easy to compute but driven entirely by extremes.
- Standard deviation measures the typical distance of values from the mean. You are never asked to compute it by hand on the SAT — only to compare. More spread out larger standard deviation; tightly bunched smaller.
- Interquartile range (IQR) , the spread of the middle of data. Like the median, IQR is resistant to outliers.
2.6 Probability Foundations
The probability of an event is
a number between (impossible) and (certain). Most SAT probability comes from two-way tables. The entire difficulty is choosing the correct denominator: the grand total, a single row total, or a single column total, depending on the wording ("of all students," "of the seniors," "given that...").
2.7 Marginal and Conditional Distributions
In a two-way table, marginal totals appear in the rightmost column and bottom row — they sum across one category. Conditional probability restricts attention to a single row or column. The phrase "given that" always tells you to shrink the denominator to that row or column total, not the grand total.
2.8 Dot Plots, Histograms, and Box Plots (Conceptual)
The SAT may present data in forms other than scatterplots:
- A dot plot stacks dots above each value on a number line — useful for spotting clusters, gaps, and outliers at a glance.
- A histogram groups data into bins; bar height represents frequency. The SAT tests whether you can read which bin contains the median or which range has the most data.
- A box plot shows the five-number summary: minimum, , median, , maximum. The box spans the IQR; whiskers extend to the extremes. Two box plots side by side let you compare center and spread without computing anything.
You will not construct these graphs on the test, but you must read them fluently.
2.9 Sampling and Generalization (Light Touch)
Occasionally the SAT asks whether a conclusion is justified based on how data was collected. A random sample from a population supports generalization; a convenience sample (surveying only people at one location) does not. You are not tested on formal statistics vocabulary, but you should recognize when a sample is too narrow to support a broad claim.
Continue reading with Premium
Upgrade to read the full article and unlock all Premium features.
Free
- Unlimited practice — all difficulties
- 3 hints / day
- Community solutions
- 2 timed mocks / month
Premium
- ✓Full article + all 57+ theory guides
- ✓Unlimited hints on practice problems
- ✓Unlimited timed mock exams & PDF worksheets