Skip to main content

Polyglot Awareness & R Deepening

Overview

Week 5 is about becoming a careful bilingual reader of data analysis code. R remains the core language for required course work. Python is introduced so you can read short pandas examples, translate small tasks with AI assistance, and verify that a translated result still answers the same question.

Polyglot awareness means being able to recognize and compare similar data steps in more than one programming language. R remains the required language; Python is used here for comparison and translation practice.

The week has two linked goals:

  • deepen R fluency with dplyr, functions, and clear summaries;
  • compare a small R result with a Python/pandas translation using explicit parity checks.

Learning Objectives

By the end of this week, you should be able to:

  • explain when R is the better default and when Python may be useful;
  • write a clear R dplyr pipeline for a descriptive health-data summary;
  • refactor repeated R code into a small function;
  • create and run a Python notebook in Codespaces;
  • load and inspect a CSV with pandas;
  • identify common R-to-Python translation traps;
  • use AI one step at a time while auditing each translated output;
  • document approved Python dependencies in the workspace-root .devcontainer/requirements.txt;
  • use Git to commit a clean translation workflow.

Connection

Week 4 made small changes reviewable through Git. This week uses that review mindset inside the analysis itself: first build a trusted R result, then translate one step at a time and compare. Next week you will turn those checked summaries into clearer visual explanations.

Case Study Data Analysis

What to do with this code: read it. The code block below runs automatically when this page is built — you do not need to run or change it here. You will write your own version in the studio.

The NHANES Health Equity data spine now supports code-reading, EDA, visualization, and reproducibility checks. Use the cached CSV/RDS for routine class work; a CDC retrieval script exists for advanced users.

The default classroom path is to use the cached data so the analysis work is reproducible without internet access.

What to do here: you may run the code below — it loads the cached CSV and prints two small summaries; reading it carefully is enough unless your week’s page asks for more.

library(tidyverse) # AI-EDIT(2026-06-23): tidyverse-default — consolidated core library() calls into library(tidyverse)

candidate_paths <- c(
  "examples/nhanes-equity/data/nhanes_equity_v6.csv",
  "../../examples/nhanes-equity/data/nhanes_equity_v6.csv"
)

nhanes_path <- candidate_paths[file.exists(candidate_paths)][1]
if (is.na(nhanes_path)) {
  stop("Could not find examples/nhanes-equity/data/nhanes_equity_v6.csv")
}

nhanes_analysis <- read_csv(nhanes_path, show_col_types = FALSE)

nhanes_analysis |>
  summarise(
    rows = n(),
    missing_bmi = sum(is.na(BMI)),
    missing_income = sum(is.na(IncomeGroup))
  )

nhanes_analysis |>
  filter(!is.na(BMI), !is.na(IncomeGroup), Age >= 20, Age <= 80) |>
  group_by(IncomeGroup) |>
  summarise(
    n = n(),
    mean_bmi = round(mean(BMI), 1),
    .groups = "drop"
  )

R Deepening Studio

Start in R. Use the NHANES case-study CSV:

examples/nhanes-equity/data/nhanes_equity_v6.csv

Where To Write Your R Code

Work directly in your personal workspace. Create assignments/assignment04-polyglot/submission/polyglot-parity.qmd and put the R studio code in R chunks there. This avoids creating a second week5.R file that would only need to be copied later.

R Pipeline and R Function

The studio code — a dplyr pipeline and a small reusable function — lives in one place so there is a single copy to keep correct: work through Parts A and B of R Deepening and Parity Check. Type or paste each block into R chunks in polyglot-parity.qmd and run it.

Before translating, compare the pipeline output and the function output. They should have the same grouping labels and the same number of rows.

In-Class Activity: R First, Python Second

Work in pairs. One partner reads the R output; the other reads the Python output.

A parity check means confirming that the R result and Python result match on the important parts: row counts, column counts, grouping labels, missing values, and rounded summary values.

R steps happen in polyglot-parity.qmd. Python steps happen in translation.ipynb in the same A4 submission folder (see Python Environment). pandas comes preinstalled in the course Codespace.

  1. Load the NHANES CSV in R and Python.
  2. Record raw row count, column count, and missing BMI count in both languages.
  3. Produce mean and SD of BMI by IncomeGroup and Gender in R.
  4. Translate the same task to Python/pandas one step at a time.
  5. Fill in a parity table comparing row counts, grouping labels, and rounded summary values.
  6. If something does not match, write the likely reason before changing code.

If the R and Python results do not match, follow the 15-minute rule: spend up to 15 minutes debugging on your own — recheck the translation traps and write down what you tried and your best guess at the cause. After 15 minutes, ask your partner, a TA, or the instructor for help, and show your notes.

Student output: a parity-check note with the R pipeline, R function, Python translation, parity table, and a short AI-use note.

Page Map

Use the child pages in this order:

  1. Python vs. R
  2. Python Environment
  3. Packages and Reproducibility
  4. Translation Traps
  5. Load and Inspect Data
  6. Python Tracebacks
  7. One Step, One Cell
  8. Mixing R and Python
  9. Reusable Python Scripts
  10. Git Progress
  11. Proposal Readiness
  12. Knowledge Check
  13. Bilingual Translation
  14. R To Python Cheatsheet

For the explicit R mastery exercises — the worked pipeline, the reusable function, and the extra dplyr and function practice — use R Deepening and Parity Check.

Assignment 4

Complete Assignment 4: Polyglot Parity.

Submit a repository link on Canvas after committing:

  • assignments/assignment04-polyglot/submission/polyglot-parity.qmd;
  • assignments/assignment04-polyglot/submission/translation.ipynb;
  • assignments/assignment04-polyglot/submission/helpers.py.

Knowledge Check

Before leaving Week 5, answer these:

  1. What is one analysis task you would keep in R?
  2. What is one task where Python may be useful?
  3. What are three parity checks you ran?
  4. What did AI help translate, and how did you verify it?
  5. What changed in your Git history from the beginning to the end of the week?

Video walkthrough (optional)