Skip to main content

Week 6: Data Visualization

Overview

This week turns a descriptive summary into a visual claim. You will use ggplot2 as the core tool, learn how to spot the most common ways a plot misleads a reader, practice prompting an AI co-pilot for plot code you can defend, and finish with a polished figure for Assignment 5.

The emphasis is not decoration. A useful plot makes the statistic, grouping, exclusions, and limitation visible enough that a reader can understand what the figure does and does not support.

ImportantRequired / Draft / Optional This Week
  • Required (graded): the seven Assignment 5 files, explained one by one in viz07, plus your part of the M1 proposal initial-visual-plan.md.
  • Draft (studio work that feeds A5, not graded on its own): the corrected plot, caption, and audit note from viz06.
  • Optional (no penalty for skipping): the alt-text practice in viz07.

Tags like [required] and [optional] are defined in How To Read Submission Lists.

Objectives

By the end of the week you can:

  • build a rerunnable ggplot2 visualization from the NHANES Health Equity CSV snapshot;
  • identify visual choices that can exaggerate, hide, or mislabel a descriptive result;
  • revise title, scale, color, grouping, and caption choices to reduce misinterpretation;
  • document whether a plot is weighted or unweighted, descriptive or causal, and complete-case or missingness-aware;
  • use AI assistance as a draft partner while checking the code and claims yourself.

Connection

Week 5 checked whether simple summaries agree across R and Python. Week 6 asks whether a visual summary is faithful to the evidence. Week 7 then documents the cohort, missingness, and Table 1 behind those same descriptive claims.

Case Study Data Analysis

The NHANES Health Equity data spine now supports code-reading, EDA, visualization, and reproducibility checks. Use the cached CSV/RDS for routine class work; a CDC retrieval script exists for advanced users.

The default classroom path is to use the cached data so the analysis work is reproducible without internet access.

What to do here: you may run the code below — it loads the cached CSV and prints two small summaries; reading it carefully is enough unless your week’s page asks for more.

library(readr)
library(dplyr)

candidate_paths <- c(
  "examples/nhanes-equity/data/nhanes_equity_v6.csv",
  "../../examples/nhanes-equity/data/nhanes_equity_v6.csv"
)

nhanes_path <- candidate_paths[file.exists(candidate_paths)][1]
if (is.na(nhanes_path)) {
  stop("Could not find examples/nhanes-equity/data/nhanes_equity_v6.csv")
}

nhanes_analysis <- read_csv(nhanes_path, show_col_types = FALSE)

nhanes_analysis |>
  summarise(
    rows = n(),
    missing_bmi = sum(is.na(BMI)),
    missing_income = sum(is.na(IncomeGroup))
  )

nhanes_analysis |>
  filter(!is.na(BMI), !is.na(IncomeGroup), Age >= 20, Age <= 80) |>
  group_by(IncomeGroup) |>
  summarise(
    n = n(),
    mean_bmi = round(mean(BMI), 1),
    .groups = "drop"
  )

Reading order

Work through the pages in order. Each page is short and self-contained.

  1. viz01 — Why Visualize, and What For
  2. viz02 — Grammar of a ggplot, and Your First Plot
  3. viz03 — Worked Example: Descriptive BMI Visualization
  4. viz04 — How a Plot Can Mislead
  5. viz05 — AI Prompts You Can Defend
  6. viz06 — In-Class Studio: AI Visual Audit
  7. viz07 — Assignment 5 Walkthrough and Reference

Class plan

  1. Read the worked example (viz03) and name the statistic, grouping variable, exclusions, and limitation.
  2. Inspect the flawed plot code referenced in the studio.
  3. Identify at least three ways the plot could mislead a reader.
  4. Revise the plot so the scale, geometry, title, labels, and caption match the evidence.
  5. Pair-review the corrected caption and check that it does not imply causality or population inference.

Student output

By the end of class each student has a corrected plot, a caption, a short audit note, and an AI-use note if AI helped draft code or wording.

Definition of done

  • The worked example renders from source using a relative path.
  • The revised figure has clear labels, an honest scale, and a limitation-aware caption.
  • The audit note identifies at least three issues across correctness, design, interpretation, or reproducibility.
  • The M1 proposal uses the same standard of transparency for dataset choice, feasibility, and stewardship.

What students leave with

One polished descriptive visualization and a reusable checklist for auditing AI-assisted plots before they appear in a report, dashboard, or slide.