library(tidyverse) # AI-EDIT(2026-06-23): tidyverse-default — consolidated core library() calls into library(tidyverse)
library(knitr)
candidate_paths <- c(
"examples/nhanes-equity/data/nhanes_equity_v6.csv",
"../../examples/nhanes-equity/data/nhanes_equity_v6.csv"
)
data_path <- candidate_paths[file.exists(candidate_paths)][1]
if (is.na(data_path)) {
stop("Could not find examples/nhanes-equity/data/nhanes_equity_v6.csv")
}
nhanes <- read_csv(data_path, show_col_types = FALSE)Worked Example: EDA and Table 1
Where this fits
You met the four-artifact workflow on eda01 and practiced cohort definition and missingness reporting on eda02. This page is the first end-to-end EDA of the term: load the cached CSV, restrict to an adult cohort, report missingness, produce a Table 1, and write a short methods note that survives audit. The same workflow underwrites Assignment 6 and Milestone 2.
Goal
This worked example creates a small Table 1-style descriptive summary from the NHANES Health Equity CSV snapshot. The goal is transparent EDA: define the cohort, check missingness, summarize key variables, and write a methods note that does not overclaim.
Define the Analysis Cohort
Decision: restrict to adults age 20-80. This keeps the example focused on adult BMI summaries and avoids mixing children, adolescents, and the oldest age-coded records into one classroom table.
Missingness Check
| variable | missing_n | missing_percent |
|---|---|---|
| BMI | 1009 | 1.8 |
| Age | 0 | 0.0 |
| IncomeGroup | 5395 | 9.4 |
| Gender | 0 | 0.0 |
| Race | 0 | 0.0 |
| Cycle | 0 | 0.0 |
Table 1-Style Summary
This example stratifies by IncomeGroup and reports N plus mean (SD) for BMI and age. It is unweighted and descriptive.
| IncomeGroup | N | BMI, mean (SD) | Age, mean (SD) |
|---|---|---|---|
| High Income (>3.5) | 16520 | 28.6 (6.3) | 49.7 (16.0) |
| Low Income (<1.3) | 15646 | 29.4 (7.5) | 47.4 (18.1) |
| Middle Income | 19576 | 29.3 (7.0) | 50.0 (18.4) |
Methods Note Template
We used the prepared NHANES Health Equity classroom CSV snapshot. The analysis cohort included records with age 20-80. We summarized BMI and age by income group using unweighted descriptive statistics. Missing values were checked before summarizing. These summaries describe the prepared class dataset and should not be interpreted as causal effects or population estimates.
Stewardship Note Template
The dataset is a prepared classroom snapshot derived from public-use NHANES files. It should be used for descriptive learning activities, not for individual-level inference or small-cell claims. Reports should cite the source, describe exclusions, and avoid uploading row-level records to external AI tools.
Survey Design Note
NHANES has a complex survey design. This classroom Table 1 is intentionally unweighted so that the workflow is easy to audit. A population-representative NHANES analysis would require the appropriate weights, strata, and primary sampling units.
Reproducibility Notes
- The data path is relative to the repository.
- The cohort restriction is stated before the table.
- Missingness is reported before complete-case summaries.
- The table labels distinguish N, mean, and SD.
- The methods note states that the analysis is descriptive and non-causal.
Where this goes next
eda04 takes a step back to ask what each statistic in this table actually claims. eda05 introduces the four audit categories you will use to evaluate a deliberately flawed version of this analysis in the eda06 studio.