Skip to main content

Worked Example: EDA and Table 1

Where this fits

You met the four-artifact workflow on eda01 and practiced cohort definition and missingness reporting on eda02. This page is the first end-to-end EDA of the term: load the cached CSV, restrict to an adult cohort, report missingness, produce a Table 1, and write a short methods note that survives audit. The same workflow underwrites Assignment 6 and Milestone 2.

Goal

This worked example creates a small Table 1-style descriptive summary from the NHANES Health Equity CSV snapshot. The goal is transparent EDA: define the cohort, check missingness, summarize key variables, and write a methods note that does not overclaim.

library(tidyverse)  # AI-EDIT(2026-06-23): tidyverse-default — consolidated core library() calls into library(tidyverse)
library(knitr)

candidate_paths <- c(
  "examples/nhanes-equity/data/nhanes_equity_v6.csv",
  "../../examples/nhanes-equity/data/nhanes_equity_v6.csv"
)

data_path <- candidate_paths[file.exists(candidate_paths)][1]
if (is.na(data_path)) {
  stop("Could not find examples/nhanes-equity/data/nhanes_equity_v6.csv")
}

nhanes <- read_csv(data_path, show_col_types = FALSE)

Define the Analysis Cohort

Decision: restrict to adults age 20-80. This keeps the example focused on adult BMI summaries and avoids mixing children, adolescents, and the oldest age-coded records into one classroom table.

analysis_df <- nhanes |>
  filter(Age >= 20, Age <= 80)

cohort_counts <- tibble(
  step = c("Prepared CSV rows", "Rows after age 20-80 restriction"),
  n = c(nrow(nhanes), nrow(analysis_df))
)

kable(cohort_counts)
step n
Prepared CSV rows 105626
Rows after age 20-80 restriction 57137

Missingness Check

key_vars <- c("BMI", "Age", "IncomeGroup", "Gender", "Race", "Cycle")

missingness <- tibble(variable = key_vars) |>
  mutate(
    missing_n = vapply(analysis_df[key_vars], function(x) sum(is.na(x)), numeric(1)),
    missing_percent = round(100 * missing_n / nrow(analysis_df), 1)
  )

kable(missingness)
variable missing_n missing_percent
BMI 1009 1.8
Age 0 0.0
IncomeGroup 5395 9.4
Gender 0 0.0
Race 0 0.0
Cycle 0 0.0

Table 1-Style Summary

This example stratifies by IncomeGroup and reports N plus mean (SD) for BMI and age. It is unweighted and descriptive.

mean_sd <- function(x) {
  sprintf("%.1f (%.1f)", mean(x, na.rm = TRUE), sd(x, na.rm = TRUE))
}

table1 <- analysis_df |>
  filter(!is.na(IncomeGroup)) |>
  group_by(IncomeGroup) |>
  summarise(
    N = n(),
    `BMI, mean (SD)` = mean_sd(BMI),
    `Age, mean (SD)` = mean_sd(Age),
    .groups = "drop"
  )

kable(table1)
IncomeGroup N BMI, mean (SD) Age, mean (SD)
High Income (>3.5) 16520 28.6 (6.3) 49.7 (16.0)
Low Income (<1.3) 15646 29.4 (7.5) 47.4 (18.1)
Middle Income 19576 29.3 (7.0) 50.0 (18.4)

Methods Note Template

We used the prepared NHANES Health Equity classroom CSV snapshot. The analysis cohort included records with age 20-80. We summarized BMI and age by income group using unweighted descriptive statistics. Missing values were checked before summarizing. These summaries describe the prepared class dataset and should not be interpreted as causal effects or population estimates.

Stewardship Note Template

The dataset is a prepared classroom snapshot derived from public-use NHANES files. It should be used for descriptive learning activities, not for individual-level inference or small-cell claims. Reports should cite the source, describe exclusions, and avoid uploading row-level records to external AI tools.

Survey Design Note

NHANES has a complex survey design. This classroom Table 1 is intentionally unweighted so that the workflow is easy to audit. A population-representative NHANES analysis would require the appropriate weights, strata, and primary sampling units.

Reproducibility Notes

  • The data path is relative to the repository.
  • The cohort restriction is stated before the table.
  • Missingness is reported before complete-case summaries.
  • The table labels distinguish N, mean, and SD.
  • The methods note states that the analysis is descriptive and non-causal.

Where this goes next

eda04 takes a step back to ask what each statistic in this table actually claims. eda05 introduces the four audit categories you will use to evaluate a deliberately flawed version of this analysis in the eda06 studio.