Income and BMI Patterns in NHANES Adults

Author

Cedar Equity Lab

Purpose and audience

This report is for municipal public-health staff preparing an internal health-equity briefing. It asks how mean body mass index (BMI) differs across income groups among adults age 20-80 in a prepared NHANES classroom dataset. The result is intended to identify a descriptive pattern worth discussing and a next analytic question, not to support a causal or population-level decision.

Data source

The report uses data/nhanes_equity_v6.csv, a cached classroom snapshot derived from CDC/NCHS National Health and Nutrition Examination Survey public-use files (Centers for Disease Control and Prevention, National Center for Health Statistics 2026). The committed copy allows the report to render without a network data retrieval.

library(tidyverse)
library(knitr)

nhanes <- read_csv("data/nhanes_equity_v6.csv", show_col_types = FALSE)

Methods

We restricted the prepared file to records with age from 20 through 80, inclusive. We reported missingness before selecting complete records for age, BMI, and income group. We then calculated unweighted counts, means, and standard deviations by income group. The cycle figure also excludes missing cycle values and groups with fewer than 30 records, although no supplied cycle-income group fell below that threshold.

adult_cohort <- nhanes |>
  filter(Age >= 20, Age <= 80)

cohort_counts <- tibble(
  step = c("Prepared file", "Adult cohort, age 20-80"),
  n = c(nrow(nhanes), nrow(adult_cohort))
)

kable(cohort_counts, caption = "Cohort definition.")
Cohort definition.
step n
Prepared file 105626
Adult cohort, age 20-80 57137
key_vars <- c("BMI", "Age", "IncomeGroup", "Gender", "Race", "Cycle")

missingness <- tibble(variable = key_vars) |>
  mutate(
    missing_n = vapply(
      adult_cohort[key_vars],
      function(x) sum(is.na(x)),
      numeric(1)
    ),
    missing_percent = round(100 * missing_n / nrow(adult_cohort), 1)
  )

kable(missingness, caption = "Missingness before complete-case exclusions.")
Missingness before complete-case exclusions.
variable missing_n missing_percent
BMI 1009 1.8
Age 0 0.0
IncomeGroup 5395 9.4
Gender 0 0.0
Race 0 0.0
Cycle 0 0.0
analysis_df <- adult_cohort |>
  filter(!is.na(Age), !is.na(BMI), !is.na(IncomeGroup))

Results

income_summary <- analysis_df |>
  group_by(IncomeGroup) |>
  summarise(
    N = n(),
    mean_bmi = mean(BMI),
    sd_bmi = sd(BMI),
    mean_age = mean(Age),
    sd_age = sd(Age),
    `BMI, mean (SD)` = sprintf("%.1f (%.1f)", mean_bmi, sd_bmi),
    `Age, mean (SD)` = sprintf("%.1f (%.1f)", mean_age, sd_age),
    .groups = "drop"
  )

kable(
  income_summary |>
    select(IncomeGroup, N, `BMI, mean (SD)`, `Age, mean (SD)`),
  caption = "Table 1-style unweighted summary by income group."
)
Table 1-style unweighted summary by income group.
IncomeGroup N BMI, mean (SD) Age, mean (SD)
High Income (>3.5) 16330 28.6 (6.3) 49.6 (15.9)
Low Income (<1.3) 15293 29.4 (7.5) 47.2 (18.1)
Middle Income 19228 29.3 (7.0) 49.8 (18.3)

Among complete adult records, mean BMI was approximately 28.6 in the high-income group, 29.4 in the low-income group, and 29.3 in the middle-income group. The difference is modest: the low- and middle-income averages were less than one BMI unit above the high-income average.

cycle_summary <- analysis_df |>
  filter(!is.na(Cycle)) |>
  group_by(Cycle, IncomeGroup) |>
  summarise(
    n = n(),
    mean_bmi = mean(BMI),
    .groups = "drop"
  ) |>
  filter(n >= 30)

ggplot(
  cycle_summary,
  aes(Cycle, mean_bmi, color = IncomeGroup, group = IncomeGroup)
) +
  geom_line(linewidth = 0.9) +
  geom_point(size = 2) +
  scale_color_brewer(palette = "Dark2") +
  expand_limits(y = 0) +
  labs(
    title = "Mean BMI by income group across NHANES cycle",
    subtitle = "Adults age 20-80 in the prepared classroom dataset",
    x = "NHANES cycle",
    y = "Mean BMI (kg/m²)",
    color = "Income group"
  ) +
  theme_minimal(base_size = 12) +
  theme(
    axis.text.x = element_text(angle = 45, hjust = 1),
    legend.position = "bottom"
  )

Mean BMI by income group across available NHANES cycles among complete adult records. The vertical scale begins at zero. Values are unweighted and descriptive, not population estimates.

The cycle-specific averages are not identical, but the figure should be read as a description of the prepared records rather than evidence of a causal income effect or a national time trend.

Plain-language takeaway

Average BMI was slightly lower in the high-income group than in the low- and middle-income groups in this classroom dataset. The difference is small and does not explain why the groups differ. For an internal briefing, the useful next question is whether the pattern remains after accounting for the NHANES survey design and examining who is missing from the comparison.

Limitations and stewardship

  • The analysis uses a prepared classroom snapshot rather than a new retrieval from the raw NHANES source files.
  • Results are unweighted and do not account for weights, strata, or primary sampling units.
  • Missing BMI and income-group records are excluded from the grouped results after missingness is reported.
  • Cross-sectional descriptive differences do not establish that income causes BMI differences.
  • Aggregate group comparisons should be communicated without stigmatizing people in any income category.
  • Row-level records were not uploaded to external AI tools or displayed in the communication products.

Reproducibility and AI audit

Render this report from the repository root with quarto render report/report.qmd. The final AI-use audit is at milestones/final-portfolio/submission/ai-use-note.md. The group checked the cohort and complete-case counts, the displayed statistics, the figure against its source data, and the CDC/NCHS citation after the final render.

Conclusion

This project supports a modest, audience-facing descriptive claim and makes its boundaries visible. A survey-aware analysis would be the appropriate next step before using the pattern for a public-health decision.

References

Centers for Disease Control and Prevention, National Center for Health Statistics. 2026. “National Health and Nutrition Examination Survey.” https://wwwn.cdc.gov/nchs/nhanes/.