library(tidyverse) # AI-EDIT(2026-06-23): tidyverse-default — consolidated core library() calls into library(tidyverse)
candidate_paths <- c(
"examples/nhanes-equity/data/nhanes_equity_v6.csv",
"../../examples/nhanes-equity/data/nhanes_equity_v6.csv"
)
nhanes_path <- candidate_paths[file.exists(candidate_paths)][1]
if (is.na(nhanes_path)) {
stop("Could not find examples/nhanes-equity/data/nhanes_equity_v6.csv")
}
nhanes_analysis <- read_csv(nhanes_path, show_col_types = FALSE)
nhanes_analysis |>
summarise(
rows = n(),
missing_bmi = sum(is.na(BMI)),
missing_income = sum(is.na(IncomeGroup))
)Polyglot Awareness & R Deepening
Overview
Week 5 is about becoming a careful bilingual reader of data analysis code. R remains the core language for required course work. Python is introduced so you can read short pandas examples, translate small tasks with AI assistance, and verify that a translated result still answers the same question.
Polyglot awareness means being able to recognize and compare similar data steps in more than one programming language. R remains the required language; Python is used here for comparison and translation practice.
The week has two linked goals:
- deepen R fluency with
dplyr, functions, and clear summaries; - compare a small R result with a Python/pandas translation using explicit parity checks.
Learning Objectives
By the end of this week, you should be able to:
- explain when R is the better default and when Python may be useful;
- write a clear R
dplyrpipeline for a descriptive health-data summary; - refactor repeated R code into a small function;
- create and run a Python notebook in Codespaces;
- load and inspect a CSV with
pandas; - identify common R-to-Python translation traps;
- use AI one step at a time while auditing each translated output;
- document approved Python dependencies in the workspace-root
.devcontainer/requirements.txt; - use Git to commit a clean translation workflow.
Connection
Week 4 made small changes reviewable through Git. This week uses that review mindset inside the analysis itself: first build a trusted R result, then translate one step at a time and compare. Next week you will turn those checked summaries into clearer visual explanations.
Case Study Data Analysis
What to do with this code: read it. The code block below runs automatically when this page is built — you do not need to run or change it here. You will write your own version in the studio.
The NHANES Health Equity data spine now supports code-reading, EDA, visualization, and reproducibility checks. Use the cached CSV/RDS for routine class work; a CDC retrieval script exists for advanced users.
- Cached RDS:
examples/nhanes-equity/data/nhanes_equity_v6.rds - CSV snapshot:
examples/nhanes-equity/data/nhanes_equity_v6.csv - Case-study README:
examples/nhanes-equity/README.md
The default classroom path is to use the cached data so the analysis work is reproducible without internet access.
What to do here: you may run the code below — it loads the cached CSV and prints two small summaries; reading it carefully is enough unless your week’s page asks for more.
R Deepening Studio
Start in R. Use the NHANES case-study CSV:
Where To Write Your R Code
Work directly in your personal workspace. Create assignments/assignment04-polyglot/submission/polyglot-parity.qmd and put the R studio code in R chunks there. This avoids creating a second week5.R file that would only need to be copied later.
R Pipeline and R Function
The studio code — a dplyr pipeline and a small reusable function — lives in one place so there is a single copy to keep correct: work through Parts A and B of R Deepening and Parity Check. Type or paste each block into R chunks in polyglot-parity.qmd and run it.
Before translating, compare the pipeline output and the function output. They should have the same grouping labels and the same number of rows.
In-Class Activity: R First, Python Second
Work in pairs. One partner reads the R output; the other reads the Python output.
A parity check means confirming that the R result and Python result match on the important parts: row counts, column counts, grouping labels, missing values, and rounded summary values.
R steps happen in polyglot-parity.qmd. Python steps happen in translation.ipynb in the same A4 submission folder (see Python Environment). pandas comes preinstalled in the course Codespace.
- Load the NHANES CSV in R and Python.
- Record raw row count, column count, and missing
BMIcount in both languages. - Produce mean and SD of
BMIbyIncomeGroupandGenderin R. - Translate the same task to Python/pandas one step at a time.
- Fill in a parity table comparing row counts, grouping labels, and rounded summary values.
- If something does not match, write the likely reason before changing code.
If the R and Python results do not match, follow the 15-minute rule: spend up to 15 minutes debugging on your own — recheck the translation traps and write down what you tried and your best guess at the cause. After 15 minutes, ask your partner, a TA, or the instructor for help, and show your notes.
Student output: a parity-check note with the R pipeline, R function, Python translation, parity table, and a short AI-use note.
Page Map
Use the child pages in this order:
- Python vs. R
- Python Environment
- Packages and Reproducibility
- Translation Traps
- Load and Inspect Data
- Python Tracebacks
- One Step, One Cell
- Mixing R and Python
- Reusable Python Scripts
- Git Progress
- Proposal Readiness
- Knowledge Check
- Bilingual Translation
- R To Python Cheatsheet
For the explicit R mastery exercises — the worked pipeline, the reusable function, and the extra dplyr and function practice — use R Deepening and Parity Check.
Assignment 4
Complete Assignment 4: Polyglot Parity.
Submit a repository link on Canvas after committing:
-
assignments/assignment04-polyglot/submission/polyglot-parity.qmd; -
assignments/assignment04-polyglot/submission/translation.ipynb; -
assignments/assignment04-polyglot/submission/helpers.py.
Knowledge Check
Before leaving Week 5, answer these:
- What is one analysis task you would keep in R?
- What is one task where Python may be useful?
- What are three parity checks you ran?
- What did AI help translate, and how did you verify it?
- What changed in your Git history from the beginning to the end of the week?