library(tidyverse)
library(knitr)
nhanes <- read_csv("data/nhanes_equity_v6.csv", show_col_types = FALSE)Income and BMI Patterns in NHANES Adults
Purpose and audience
This report is for municipal public-health staff preparing an internal health-equity briefing. It asks how mean body mass index (BMI) differs across income groups among adults age 20-80 in a prepared NHANES classroom dataset. The result is intended to identify a descriptive pattern worth discussing and a next analytic question, not to support a causal or population-level decision.
Data source
The report uses data/nhanes_equity_v6.csv, a cached classroom snapshot derived from CDC/NCHS National Health and Nutrition Examination Survey public-use files (Centers for Disease Control and Prevention, National Center for Health Statistics 2026). The committed copy allows the report to render without a network data retrieval.
Methods
We restricted the prepared file to records with age from 20 through 80, inclusive. We reported missingness before selecting complete records for age, BMI, and income group. We then calculated unweighted counts, means, and standard deviations by income group. The cycle figure also excludes missing cycle values and groups with fewer than 30 records, although no supplied cycle-income group fell below that threshold.
adult_cohort <- nhanes |>
filter(Age >= 20, Age <= 80)
cohort_counts <- tibble(
step = c("Prepared file", "Adult cohort, age 20-80"),
n = c(nrow(nhanes), nrow(adult_cohort))
)
kable(cohort_counts, caption = "Cohort definition.")| step | n |
|---|---|
| Prepared file | 105626 |
| Adult cohort, age 20-80 | 57137 |
key_vars <- c("BMI", "Age", "IncomeGroup", "Gender", "Race", "Cycle")
missingness <- tibble(variable = key_vars) |>
mutate(
missing_n = vapply(
adult_cohort[key_vars],
function(x) sum(is.na(x)),
numeric(1)
),
missing_percent = round(100 * missing_n / nrow(adult_cohort), 1)
)
kable(missingness, caption = "Missingness before complete-case exclusions.")| variable | missing_n | missing_percent |
|---|---|---|
| BMI | 1009 | 1.8 |
| Age | 0 | 0.0 |
| IncomeGroup | 5395 | 9.4 |
| Gender | 0 | 0.0 |
| Race | 0 | 0.0 |
| Cycle | 0 | 0.0 |
analysis_df <- adult_cohort |>
filter(!is.na(Age), !is.na(BMI), !is.na(IncomeGroup))Results
income_summary <- analysis_df |>
group_by(IncomeGroup) |>
summarise(
N = n(),
mean_bmi = mean(BMI),
sd_bmi = sd(BMI),
mean_age = mean(Age),
sd_age = sd(Age),
`BMI, mean (SD)` = sprintf("%.1f (%.1f)", mean_bmi, sd_bmi),
`Age, mean (SD)` = sprintf("%.1f (%.1f)", mean_age, sd_age),
.groups = "drop"
)
kable(
income_summary |>
select(IncomeGroup, N, `BMI, mean (SD)`, `Age, mean (SD)`),
caption = "Table 1-style unweighted summary by income group."
)| IncomeGroup | N | BMI, mean (SD) | Age, mean (SD) |
|---|---|---|---|
| High Income (>3.5) | 16330 | 28.6 (6.3) | 49.6 (15.9) |
| Low Income (<1.3) | 15293 | 29.4 (7.5) | 47.2 (18.1) |
| Middle Income | 19228 | 29.3 (7.0) | 49.8 (18.3) |
Among complete adult records, mean BMI was approximately 28.6 in the high-income group, 29.4 in the low-income group, and 29.3 in the middle-income group. The difference is modest: the low- and middle-income averages were less than one BMI unit above the high-income average.
cycle_summary <- analysis_df |>
filter(!is.na(Cycle)) |>
group_by(Cycle, IncomeGroup) |>
summarise(
n = n(),
mean_bmi = mean(BMI),
.groups = "drop"
) |>
filter(n >= 30)
ggplot(
cycle_summary,
aes(Cycle, mean_bmi, color = IncomeGroup, group = IncomeGroup)
) +
geom_line(linewidth = 0.9) +
geom_point(size = 2) +
scale_color_brewer(palette = "Dark2") +
expand_limits(y = 0) +
labs(
title = "Mean BMI by income group across NHANES cycle",
subtitle = "Adults age 20-80 in the prepared classroom dataset",
x = "NHANES cycle",
y = "Mean BMI (kg/m²)",
color = "Income group"
) +
theme_minimal(base_size = 12) +
theme(
axis.text.x = element_text(angle = 45, hjust = 1),
legend.position = "bottom"
)The cycle-specific averages are not identical, but the figure should be read as a description of the prepared records rather than evidence of a causal income effect or a national time trend.
Plain-language takeaway
Average BMI was slightly lower in the high-income group than in the low- and middle-income groups in this classroom dataset. The difference is small and does not explain why the groups differ. For an internal briefing, the useful next question is whether the pattern remains after accounting for the NHANES survey design and examining who is missing from the comparison.
Limitations and stewardship
- The analysis uses a prepared classroom snapshot rather than a new retrieval from the raw NHANES source files.
- Results are unweighted and do not account for weights, strata, or primary sampling units.
- Missing BMI and income-group records are excluded from the grouped results after missingness is reported.
- Cross-sectional descriptive differences do not establish that income causes BMI differences.
- Aggregate group comparisons should be communicated without stigmatizing people in any income category.
- Row-level records were not uploaded to external AI tools or displayed in the communication products.
Reproducibility and AI audit
Render this report from the repository root with quarto render report/report.qmd. The final AI-use audit is at milestones/final-portfolio/submission/ai-use-note.md. The group checked the cohort and complete-case counts, the displayed statistics, the figure against its source data, and the CDC/NCHS citation after the final render.
Conclusion
This project supports a modest, audience-facing descriptive claim and makes its boundaries visible. A survey-aware analysis would be the appropriate next step before using the pattern for a public-health decision.