library(tidyverse) # AI-EDIT(2026-06-23): tidyverse-default — consolidated core library() calls into library(tidyverse)
candidate_paths <- c(
"examples/nhanes-equity/data/nhanes_equity_v6.csv",
"../../examples/nhanes-equity/data/nhanes_equity_v6.csv"
)
nhanes_path <- candidate_paths[file.exists(candidate_paths)][1]
if (is.na(nhanes_path)) {
stop("Could not find examples/nhanes-equity/data/nhanes_equity_v6.csv")
}
nhanes_analysis <- read_csv(nhanes_path, show_col_types = FALSE)
nhanes_analysis |>
summarise(
rows = n(),
missing_bmi = sum(is.na(BMI)),
missing_income = sum(is.na(IncomeGroup))
)Week 3: R with AI
Overview
Summary: In Week 2 you learned how to use your Codespace, edit a Quarto document, and make your first commit. Now you learn how to speak R—not by memorizing syntax, but by directing an AI co-pilot and auditing what it writes. Your job is to be the pilot: you decide what the analysis should do, the AI drafts the code, and you verify every line before it flies.
What is R? R is the main programming language we use in this course for cleaning, summarizing, and visualizing health data. GitHub and Codespaces provide the workspace; R is the tool we use inside that workspace to work with data.
Learning Objectives:
- Understand that R is a calculator that needs specific, case-sensitive instructions.
- Load packages with
library()and explain why this is needed each session. - Explain how the course devcontainer and dependency notes make package use visible.
- Use core constructs: variables (assignment), functions (with arguments), and the pipe (
%>%/|>). - Perform basic data transformations using
dplyrverbs:filter(),select(),mutate(),summarise(),group_by(). - Inspect a dataset on first contact: structure, dimensions, missingness.
- Identify control flow (
if/else, loops) in AI-generated code and audit whether it is correct. - Write and adapt a simple custom function.
- Apply an AI safety checklist to catch hallucinations, wrong column names, and unsafe operations.
- Document your analysis in Quarto by weaving code chunks with Markdown narrative.
- Commit incrementally as you work (reinforcing Week 2 Git basics).
Connection
Week 2 made the project structure reproducible. This week you use that structure to load, inspect, summarize, and audit R code with AI support. Next week you will put those edits under version control.
Before you start, make sure your tools are ready: The Setup (The Cockpit).
For each code block below, follow the label above it: Read means inspect only, Run means execute it as written, Modify means make the requested change, and Copy means place it in the assignment file named in the instructions.
Case Study Data Analysis
What to do with this code (the case-study block below): Run — it loads the cached NHANES CSV and prints two summaries. You do not need to modify it this week.
The NHANES Health Equity data spine now supports code-reading, EDA, visualization, and reproducibility checks. Use the cached CSV/RDS for routine class work; a CDC retrieval script exists for advanced users.
- Cached RDS:
examples/nhanes-equity/data/nhanes_equity_v6.rds - CSV snapshot:
examples/nhanes-equity/data/nhanes_equity_v6.csv - Case-study README:
examples/nhanes-equity/README.md
The default classroom path is to use the cached data so the analysis work is reproducible without internet access.
What to do here: you may run the code below — it loads the cached CSV and prints two small summaries; reading it carefully is enough unless your week’s page asks for more.
In-Class Activity
- Load the CSV snapshot and inspect column names, row count, and missingness.
- Summarize one numeric variable and one grouping variable with clear
NAhandling. - Ask an AI assistant to comment on your code, then audit whether the comments match the actual operations.
- Note one place where an AI assistant could hallucinate a variable name or overstate what the descriptive summary means.
Student output: a short Quarto note with one data inspection, one summary, and one AI-audit comment. Assignment link: Assignment 2.
Packages: Opening the Toolbox
Deeper dive: Packages (The Toolbox).
Install vs Load
-
install.packages("tidyverse")— buying the toolbox. Done once per environment (already done in your course Codespace). -
library(tidyverse)— opening the toolbox. Required every time you restart R or reopen your Codespace.
What to do with this code: Run — execute it as written every time you reopen your Codespace.
Where to run this: First start R so you are inside an R session. Open a terminal in VS Code (Terminal → New Terminal), type R, and press Enter. You should now see the R prompt, which looks like >. Then type the line below at that > prompt (you can also run it from the R Console or a code chunk in your .qmd).
If you see an error like there is no package called 'tidyverse', the package is not installed in your Codespace. Use the Canvas support channel or office hours and include the full error message.
Course Environment and Dependency Note
The personal workspace devcontainer is the supported course environment. It installs the R and Python packages used by the required pathway, so you do not need to initialize renv, create a lockfile, or install packages for Assignment 2.
For every submitted analysis:
- Load each R package explicitly with
library(...)in the Quarto source. - Name the packages you used in
dependency-note.md. - State whether you used only the course defaults.
- If you think you need another package, ask the instructor or TA before adding it. Record an approved addition in the workspace devcontainer setup and in your dependency note so a fresh Codespace can reproduce it.
For Assignment 2, a complete dependency note can be as short as:
The test is practical: your document must render in a fresh course Codespace.
Core Constructs: Boxes, Toasters, and the Assembly Line
Deeper dives: Basics: Boxes & Toasters; the pipe is also covered in Control Flow.
Variables: The Labeled Box
R has no memory unless you store results in a named box using <-.
What to do with this code: Run — type it in the Console, then check the Environment pane to see your new “boxes”.
Case sensitivity: Age, age, and AGE are three different boxes. This is one of the most common bugs in AI-generated code. Always check that column names match your data exactly.
Functions: The Toaster
A function takes input and returns output. Arguments (inside the parentheses) control its behaviour.
What to do with this code: Run — execute it as written.
-
c(72, 85, 90)is the input (a vector of numbers). -
na.rm = TRUEtells R to ignore missing values. Always look for this when auditing AI code that works with health data — missingness is the norm, not the exception.
The Pipe: The Assembly Line
The pipe passes the result of one step into the next. Think of it as “and then…”
What to do with this code: Read — data here is a placeholder, so this block will not run as-is.
Read this as: “Take the data, and then keep only patients over 65, and then calculate the average BMI.”
You will see two pipe symbols in the wild: %>% (from the magrittr / tidyverse packages) and |> (built into R 4.1+). They do the same thing for our purposes. If AI generates the other one, don’t panic — both work in your Codespace.
Data Import and First Contact
Deeper dive: The Loading Dock (Data Import).
Loading Data
Use relative paths so your code runs on any machine (including your grader’s Codespace).
What to do with this code: Modify — replace the path with the relative path to your own file, then run it.
How to get the path right: Right-click the file in the VS Code Explorer → Copy Relative Path → paste it into your code.
AI Prompt:
Read the CSV file at “data/patients.csv” and save it as
my_data. Use thereadrpackage.
First Contact: Inspect Before You Analyze
Never start analysis without looking at your data. Run these four checks:
What to do with this code: Run — execute it as written once your data is loaded.
Run these in your assignment .qmd (as a code chunk) or at the R > prompt, using the course workspace repo. glimpse() comes from tidyverse, so if you get could not find function "glimpse", run library(tidyverse) first.
Open the Environment pane (top-right in RStudio or via the R extension) to see your variables and click a data frame to view it in a grid.
Health data reality: Real datasets like NHANES have substantial missingness. Knowing where and how much is missing is not optional — it shapes every downstream decision.
A concrete example: in NHANES, income questions are skipped far more often than age questions. If you silently drop rows with missing income, your “average” quietly excludes many of the very people a health-equity analysis is trying to describe. Handling missingness carefully is part of responsible data stewardship, not just a coding chore.
Data Transformation: The Five dplyr Verbs
These five verbs are the core of data wrangling in R. You will use them in nearly every assignment and in your term project.
What to do with the code in this section: Modify — swap in your own dataset and column names, then run each block.
filter() — Keep rows that match a condition
select() — Keep (or drop) columns
mutate() — Create or modify a column
summarise() — Collapse rows into a summary
group_by() + summarise() — Summaries by group
AI Prompt:
Using
dplyr, filter my dataset to patients over 65, group by sex, and calculate the mean BMI for each group. Handle missing values.
Audit the AI output: After running the AI-generated code, check that (a) the column names match your actual data, (b) na.rm = TRUE is present where needed, and (c) the row counts make sense (e.g., the filtered data should be smaller than the original).
Git checkpoint: You have now loaded, inspected, and transformed data. Stage your .qmd, write a commit message describing what you did, and commit. Don’t wait until the end.
Control Flow: Auditing AI Decisions
Deeper dive: Control Flow.
You will not write many if/else blocks or loops by hand. But AI frequently generates them, so you need to read them.
What to do with the code in this section: Read — these are patterns to recognize in AI output; you do not need to run them.
If/Else: The Fork in the Road
Audit question: Is the threshold (50) appropriate for my research question, or did the AI pick an arbitrary number?
Loops: The Treadmill
Audit questions: Does the loop have a clear end condition? Could this be replaced with a simpler vectorized operation (e.g., summarise(across(...)))?
In R, loops are rarely the best tool. If AI gives you a loop, ask: “Can this be done without a loop using dplyr or purrr?” The answer is usually yes.
Writing a Simple Function
The syllabus asks you to write at least one custom function. A function packages a repeated task so you don’t copy-paste code.
Example: A Reusable Summary
What to do with this code: Run, then Modify — run it as written on your data, then change one thing as described in “Your Task” below.
summarise_column <- function(data, column_name) {
data %>%
summarise(
n_total = n(),
n_missing = sum(is.na(.data[[column_name]])),
mean_val = mean(.data[[column_name]], na.rm = TRUE),
sd_val = sd(.data[[column_name]], na.rm = TRUE)
)
}
# Use it
summarise_column(my_data, "bmi")
summarise_column(my_data, "age")Your Task
- Ask your AI co-pilot to write a function that takes a data frame and a column name and returns summary statistics.
-
Audit the code: Does it handle
NA? Does it work on a numeric column? What happens if you pass a categorical column? -
Adapt it: Change one thing (e.g., add a median, or add a
group_byargument). This is the difference between accepting code blindly and understanding it.
AI Prompt:
Write an R function called
summarise_columnthat takes a data frame and a column name (as a string) and returns the count, number of missing values, mean, and standard deviation. Then show me how to call it on the “bmi” column ofmy_data.
Debugging: Reading the Red Text
Deeper dive: The Detective (Debugging).
Errors are normal. AI makes mistakes, you make typos, and data surprises you.
The “Fix It” Loop
-
Read the error message. Look for keywords:
not found,unexpected,non-numeric. - Copy the full error text.
- Paste it into your AI chat and ask:
Here is an R error from my Codespace: [paste]. Explain what went wrong in one sentence and give corrected code.
Common Errors and What They Mean
| Error message | Likely cause |
|---|---|
object 'X' not found |
Typo in variable name, or you forgot to run an earlier chunk |
could not find function "X" |
Forgot library(tidyverse) or misspelled the function |
non-numeric argument to binary operator |
Trying to do math on a text/factor column |
unexpected symbol |
Missing comma, parenthesis, or pipe |
AI debugging guardrail: When AI suggests a fix, re-read the corrected code before running it. Does the fix address the actual error, or did AI rewrite the whole chunk and change your logic?
When asking AI for help, explicitly say: “Do not change the logic or intended result of my code. Explain any suggested change before rewriting it.”
AI Safety: The Audit Checklist
Deeper dive: AI Safety: The Audit Checklist.
Every time you accept AI-generated code, run through this checklist:
-
Column names: Are they spelled exactly as they appear in
glimpse(my_data)? - Data types: Is the code doing math only on numeric columns?
-
Missingness: Is
na.rm = TRUEpresent where needed? - Privacy: Does the code avoid printing individual patient rows that could be identifiable?
- Logic: Does the code do what you intended, or what the AI assumed you meant?
The Paper Trail Rule
Deeper dive: The Paper Trail (Comments).
Every time you use AI to write a code chunk, add a # comment above it explaining what you asked for and what you verified:
What to do with this code: Read — then imitate this comment pattern above your own AI-assisted chunks.
This comment trail is your audit log. It demonstrates accountability (required by the course AI policy) and helps you debug later.
Exporting Clean Data
After cleaning and transforming, save your work so future weeks can pick up where you left off.
What to do with this code: Run — it creates output/clean_data.csv; create the output/ folder first if it does not exist.
Reproducibility habit: Your Quarto document should contain every step from raw data to clean export. Anyone (including your grader) should be able to render it from scratch and get the same clean_data.csv.
Git checkpoint: Your analysis is complete. Stage, commit with a clear message (e.g., "Complete Week 3: load, inspect, transform, and export patient data"), and Sync.
Literate Analysis in Quarto
Your .qmd file is not just code — it is a document. Weave narrative around your code chunks to explain what you are doing and why.
Pattern
What to do with this code: Read — a pattern to imitate in your own document, not text to paste as-is.
## Data Inspection
We begin by examining the structure of the dataset to understand
variable types and missingness before any transformations.
```{r, eval=FALSE}
glimpse(my_data)
colSums(is.na(my_data))
```
The dataset contains `nrow(my_data)` observations across
`ncol(my_data)` variables. Notably, BMI is missing for
approximately 12% of records, which we handle with `na.rm = TRUE`
in downstream summaries.This interleaving of code and prose is literate programming — the methodology that underpins reproducible research in this course. Practice it now; you will need it for your term project report (Week 10) and final portfolio.
Knowledge Check
- What is the difference between
install.packages()andlibrary()? - Why is
na.rm = TRUEimportant when working with health data? - What does
my_data %>% filter(age > 65) %>% summarise(n = n())do? Read it aloud using “and then.” - Name the five core
dplyrverbs and what each does in one sentence. - You see the error
object 'BMI' not foundbut your column is calledbmi. What went wrong? - What is the difference between
%>%and|>? Does it matter for this course? - Why should you write
#comments above AI-generated code? - What four checks should you run immediately after loading a new dataset?
Assignment 2: R With AI
Full instructions: Assignment 2 README. Short version: Assignment 2 walkthrough.
Task: Build a short Quarto report that loads, inspects, transforms, and summarizes a health dataset — using your AI co-pilot for code generation and your own judgment for verification.
Step-by-Step Workflow
- Open the Assignment 2 README and follow its setup steps: copy the starter file (
assignments/assignment02-r-with-ai/starter-analysis.qmd) into your submission folder asanalysis-note.qmd. The dataset isexamples/nhanes-equity/data/nhanes_equity_v6.csv(view on GitHub). -
Load packages and data: Use AI to write the
library()andread_csv()calls. Verify the data loaded correctly withdim()andglimpse(). -
Inspect missingness: Run
colSums(is.na(...)). Add a Markdown paragraph interpreting what you see. -
Transform: Use
dplyrverbs to:- Filter to a meaningful subgroup (e.g., age ≥ 18).
- Create at least one new column with
mutate(). - Produce a grouped summary with
group_by()+summarise().
- Write a function: Create (or adapt from AI) a reusable summary function. Call it on at least two different columns.
-
Export: Save your summary table to
output/summary_table.csv. -
Document: Every code chunk should have:
- A
#comment stating what you asked the AI and what you verified. - A Markdown paragraph below interpreting the result.
- A
- Render the document to verify it runs cleanly from top to bottom.
- Git: Stage, write a descriptive commit message, Commit, and Sync.
Deliverable
A rendered .qmd that a reader can follow as a narrative: what data you started with, what you did to it, what you found, and why your decisions were justified.
Quick Reference
See also: Final Summary Checklist.
| Concept | R code | Notes |
|---|---|---|
| Load packages | library(tidyverse) |
Every session |
| Read CSV | read_csv("data/file.csv") |
Use relative paths |
| Assignment | x <- 5 |
The <- arrow stores a value |
| Inspect structure | glimpse(df) |
Column names, types, preview |
| Dimensions | dim(df) |
Rows × columns |
| Summary stats | summary(df) |
Min, median, mean, max, NAs |
| Check missingness | colSums(is.na(df)) |
Count NAs per column |
| Filter rows | filter(df, age > 65) |
Keep matching rows |
| Select columns | select(df, id, age, bmi) |
Keep named columns |
| New column | mutate(df, bmi_cat = ...) |
Create or modify |
| Summarize | summarise(df, m = mean(x, na.rm = TRUE)) |
Collapse to one row |
| Group + summarize | group_by(df, sex) %>% summarise(...) |
Per-group summaries |
| Export | write_csv(df, "output/clean.csv") |
Save for later |
| Pipe |
%>% or |>
|
“And then…” |