What Each Statistic Claims (and Does Not)
Where this fits
You produced a Table 1 on the previous page. Before you write an interpretation of it, pause and ask: what does each cell in that table actually claim? This page is a short tour through the five most common descriptive statistics you will see — what each one says, what it leaves out, and the most common ways beginners (and AI tools) mislabel them.
The five common statistics
| Statistic | What it answers | What it does not answer |
|---|---|---|
| N (count) | How many records are in this group? | Whether the group is representative of any population |
| Mean | What is the average value? | Whether the distribution is symmetric, skewed, or has outliers |
| SD (standard deviation) | How spread out are values around the mean? | Whether the spread is symmetric or driven by a tail |
| Median | What value falls in the middle? | How extreme the tails are |
| IQR (interquartile range) | What is the spread of the middle 50%? | The tails outside the middle 50% |
A Table 1 typically reports N plus either mean (SD) or median (IQR) depending on the variable.
When to prefer mean(SD) vs median(IQR)
Use mean (SD) when the variable is roughly symmetric and well-behaved. Age in the adult cohort is a reasonable example.
Use median (IQR) when the variable is skewed or has outliers. Healthcare costs, hospital length of stay, and household income are usual suspects. The mean of a skewed variable is pulled by the long tail; the median is not.
If you are unsure, plot a histogram. If you see a long tail on one side, prefer median (IQR).
For this course’s NHANES classroom Table 1, BMI and age are commonly reported as mean (SD). That is a reasonable convention for adult anthropometric variables. State the choice and the reason in your methods note.
What “N” alone does not tell you
A cell that says N = 47 is informative only if the reader knows:
- which cohort the 47 came from (cohort definition, eda02);
- how many of those 47 were excluded for missing values (missingness table, also eda02);
- what the comparison group sizes are (the rest of the Table 1).
If any of those three is invisible, the N is not yet defensible.
Three common beginner mistakes
These are the mistakes that show up most often in AI-drafted Table 1 code. The studio revisits them on a deliberately flawed starter.
1. Label says “mean” but code computes median()
Either change the label to Median BMI or change the function to mean(). Decide which one you want — do not let the AI choose silently.
2. SD reported without na.rm = TRUE
If a single NA is present, sd(BMI) returns NA. Use na.rm = TRUE and report the missingness count separately (see eda02).
3. Percent without a denominator
A percent is only meaningful if the reader can see the denominator. Always include the N in the same row or in an adjacent cell.
Small-cell suppression: when not to report a number
If a stratified cell has very few records (say, fewer than 10), publishing the exact statistic can be misleading or, in some health data settings, a privacy risk. The convention is to either:
- combine the small group with a neighbor and label it clearly, or
- replace the cell with a “<10” marker.
Your classroom Table 1 with n >= 30 per group (as in the worked example) is already cautious. The principle still applies: do not interpret cells where N is too small to support the claim.
The audit habit for descriptive statistics
After you build a Table 1, look at each column heading and each row:
- Does the column heading describe the statistic actually computed (mean, median, count, percent)?
- Is the denominator visible for every percent?
- Are missing values either filtered out (with the count reported elsewhere) or handled with
na.rm = TRUE? - Are small cells either combined, marked, or excluded?
If any answer is “no”, revise the table or its labels before you write the methods note.
Where this goes next
eda05 introduces the four AI-audit categories that organize every check you have seen so far this week, ready for the studio on the planted-error starter.