R Deepening and Parity Check
Week 5 is polyglot-aware, not Python-first. Your strongest analysis path remains R and the tidyverse. Python enters as a comparison language: it helps you read code from other sources, ask better AI translation questions, and verify that a translated result still means the same thing.
A parity check means confirming that the R result and Python result match on the important parts: row counts, column counts, grouping labels, missing values, and rounded summary values.
Learning Goals
By the end of this activity, you should be able to:
- write a clear R
dplyrpipeline for a descriptive health-data summary; - refactor a repeated summary into a small R function;
- translate the same task to Python/pandas one step at a time;
- compare outputs across languages using row counts, grouping levels, and summary values;
- document any mismatch before deciding which code to change.
Dataset
Use the NHANES case-study snapshot:
Part A: R Pipeline
Write Parts A and B as R chunks in assignments/assignment04-polyglot/submission/polyglot-parity.qmd in your personal workspace. Work in the assessed file from the start rather than creating a second script to move later.
Start with R. The goal is a descriptive, unweighted summary of BMI by income group and gender.
library(tidyverse)
nhanes <- read_csv("examples/nhanes-equity/data/nhanes_equity_v6.csv")
r_summary <- nhanes |>
filter(!is.na(BMI), !is.na(IncomeGroup), !is.na(Gender)) |>
group_by(IncomeGroup, Gender) |>
summarise(
n = n(),
mean_bmi = round(mean(BMI), 2),
sd_bmi = round(sd(BMI), 2),
.groups = "drop"
) |>
arrange(IncomeGroup, Gender)
r_summaryPart B: R Function
Refactor the summary so the value and grouping variables are easy to change later.
summarise_mean_sd <- function(data, value, group1, group2) {
data |>
filter(!is.na({{ value }}), !is.na({{ group1 }}), !is.na({{ group2 }})) |>
group_by({{ group1 }}, {{ group2 }}) |>
summarise(
n = n(),
mean_value = round(mean({{ value }}), 2),
sd_value = round(sd({{ value }}), 2),
.groups = "drop"
) |>
arrange({{ group1 }}, {{ group2 }})
}
r_function_summary <- summarise_mean_sd(nhanes, BMI, IncomeGroup, Gender)
r_function_summaryBefore moving to Python, check that the function output has the same number of rows and the same grouping labels as the original pipeline.
More R Mastery Exercises
The pipeline and function above are the worked example. Complete at least two of these short exercises in R chunks in assignments/assignment04-polyglot/submission/polyglot-parity.qmd before moving to Python:
- Swap the value column. Run
summarise_mean_sd(nhanes, Waist, IncomeGroup, Gender). Check that the grouping labels match the BMI summary. Did the row count change? Why might it? - Swap a grouping column. Run
summarise_mean_sd(nhanes, BMI, Race, Gender). How many grouped rows do you get now, and why? - Extend the pipeline. Add
mean_age = round(mean(Age), 1)insidesummarise()in the Part A pipeline, then rerun it. Confirm the number of rows did not change — adding a summary column should never change the grouping.
AI Audit Prompt
Use this prompt only after you have a working R result:
Translate this R
dplyrpipeline into Python/pandas one step at a time. For each step, explain the equivalent Python operation, the output I should inspect, and the parity check I should run against my R result. Do not combine steps or add new analysis choices.
Part C: Python Translation
Create the Python translation at assignments/assignment04-polyglot/submission/translation.ipynb in your personal workspace — the same notebook you set up in Python Environment and use in Bilingual Translation. Attempt each step yourself first with the cheat sheet; use AI to check or fix one step at a time.
Translate one step at a time. Do not ask AI for a full script until you have verified each cell.
import pandas as pd
nhanes = pd.read_csv("../../../examples/nhanes-equity/data/nhanes_equity_v6.csv")
py_summary = (
nhanes
.dropna(subset=["BMI", "IncomeGroup", "Gender"])
.groupby(["IncomeGroup", "Gender"], as_index=False)
.agg(
n=("BMI", "size"),
mean_bmi=("BMI", "mean"),
sd_bmi=("BMI", "std")
)
.sort_values(["IncomeGroup", "Gender"])
)
py_summary["mean_bmi"] = py_summary["mean_bmi"].round(2)
py_summary["sd_bmi"] = py_summary["sd_bmi"].round(2)
py_summaryPart D: Parity Audit
Record the following checks in a short Quarto note:
| Check | R command | Python command | What to record |
|---|---|---|---|
| Raw row count | nrow(nhanes) |
len(nhanes) |
Same number of rows |
| Column names | names(nhanes) |
nhanes.columns.tolist() |
Same expected variables |
| Analysis rows after missingness filter | nrow(filter(nhanes, !is.na(BMI), !is.na(IncomeGroup), !is.na(Gender))) |
len(nhanes.dropna(subset=["BMI", "IncomeGroup", "Gender"])) |
Same filtered row count |
| Group count | nrow(r_summary) |
len(py_summary) |
Same number of grouped rows |
| Summary values | inspect mean_bmi and sd_bmi |
inspect mean_bmi and sd_bmi |
Values match after rounding |
Debug on your own for up to 15 minutes first: recheck the translation traps, rerun the checks above, and write down what you tried and your best guess at the cause. After 15 minutes, ask for help — and bring your notes. Documented debugging attempts are part of the expected workflow, not a sign of failure.
Student Output
Submit a short note with:
- the R pipeline;
- the R function;
- the Python translation;
- a parity table showing the checks above;
- two sentences explaining any mismatch or confirming that the outputs agree.
These are the Assignment 4 submission files; do not also save copies under a weeks/ folder in your personal workspace.