Income and BMI Patterns in NHANES Adults — Draft

Audience and question

This draft is for municipal public-health staff preparing an internal health-equity briefing. It asks how mean BMI differs across income groups among adults age 20-80 in the prepared classroom dataset.

Data and methods

The file is a prepared classroom snapshot derived from CDC/NCHS NHANES public-use data (Centers for Disease Control and Prevention, National Center for Health Statistics 2026). We restricted the file to ages 20-80, reported missingness, and then used complete records for age, BMI, and income group. All summaries are unweighted and descriptive.

library(tidyverse)
library(knitr)

nhanes <- read_csv("data/nhanes_equity_v6.csv", show_col_types = FALSE)

adult_cohort <- nhanes |>
  filter(Age >= 20, Age <= 80)

missingness <- tibble(
  variable = c("BMI", "Age", "IncomeGroup"),
  missing_n = c(
    sum(is.na(adult_cohort$BMI)),
    sum(is.na(adult_cohort$Age)),
    sum(is.na(adult_cohort$IncomeGroup))
  )
)

kable(missingness, caption = "Missingness before complete-case exclusions.")
Missingness before complete-case exclusions.
variable missing_n
BMI 1009
Age 0
IncomeGroup 5395
analysis_df <- adult_cohort |>
  filter(!is.na(Age), !is.na(BMI), !is.na(IncomeGroup))

Current results

income_summary <- analysis_df |>
  group_by(IncomeGroup) |>
  summarise(
    n = n(),
    mean_bmi = mean(BMI),
    sd_bmi = sd(BMI),
    `BMI, mean (SD)` = sprintf("%.1f (%.1f)", mean_bmi, sd_bmi),
    .groups = "drop"
  )

kable(
  income_summary |> select(IncomeGroup, n, `BMI, mean (SD)`),
  caption = "Unweighted BMI summary by income group."
)
Unweighted BMI summary by income group.
IncomeGroup n BMI, mean (SD)
High Income (>3.5) 16330 28.6 (6.3)
Low Income (<1.3) 15293 29.4 (7.5)
Middle Income 19228 29.3 (7.0)

Mean BMI is slightly lower in the high-income group than in the low- and middle-income groups in this prepared dataset.

cycle_summary <- analysis_df |>
  filter(!is.na(Cycle)) |>
  group_by(Cycle, IncomeGroup) |>
  summarise(n = n(), mean_bmi = mean(BMI), .groups = "drop") |>
  filter(n >= 30)

ggplot(
  cycle_summary,
  aes(Cycle, mean_bmi, color = IncomeGroup, group = IncomeGroup)
) +
  geom_line(linewidth = 0.9) +
  geom_point(size = 2) +
  scale_color_brewer(palette = "Dark2") +
  expand_limits(y = 0) +
  labs(
    title = "Mean BMI by income group across NHANES cycle",
    subtitle = "Adults age 20-80 in the prepared classroom dataset",
    x = "NHANES cycle",
    y = "Mean BMI (kg/m²)",
    color = "Income group"
  ) +
  theme_minimal(base_size = 12) +
  theme(
    axis.text.x = element_text(angle = 45, hjust = 1),
    legend.position = "bottom"
  )

Mean BMI by income group across available NHANES cycles. Values are unweighted and descriptive; the vertical scale begins at zero so small differences are not exaggerated.

Draft interpretation

The overall difference is modest. The figure is useful for seeing that the three groups do not have identical cycle-specific averages, but it does not identify a cause or show a population trend.

Limitations

  • The file is a prepared classroom snapshot rather than a new raw-data retrieval.
  • The summaries are unweighted and do not account for the NHANES survey design.
  • Missing BMI and income-group records are excluded from grouped results.
  • Cross-sectional descriptive differences do not establish causality.

Draft next step

For a real decision, the next analysis would use the appropriate survey design and examine whether missingness changes the comparison. That extension is not part of this course project.

References

Centers for Disease Control and Prevention, National Center for Health Statistics. 2026. “National Health and Nutrition Examination Survey.” https://wwwn.cdc.gov/nchs/nhanes/.