Skip to main content

Week 7: EDA, Table 1 & AI Auditing

Overview

This week builds the audit trail behind a descriptive analysis. You will define an analysis cohort, check missingness, create a Table 1-style summary, and audit a planted-error starter that mimics common AI-assisted mistakes.

A Table 1 is a descriptive summary of the cohort before the main analysis. It usually shows who is included, key variables, missingness, and basic group summaries. It is not a causal result. Worked examples of every piece — cohort definition, exclusions, the missingness table, complete-case analysis, and a sample Table 1 — are in eda02 and eda03.

EDA is treated as evidence stewardship: every exclusion, statistic, and interpretation should be visible enough for another person to rerun and challenge.

This week also revisits data provenance and Indigenous Data Sovereignty. The stewardship audit category (eda05) asks the same questions you met with OCAP and CARE in Week 1: who has rights over this data, and does our use respect them?

ImportantRequired / Draft / Optional This Week
  • Required (graded): the six Assignment 6 files, explained one by one in eda07, and the group M2 preliminary analysis due the same Monday.
  • Draft (studio work that feeds A6, not graded on its own): the EDA note, missingness table, Table 1, and planted-error audit from eda06.
  • Optional: nothing extra this week — every listed file is required.

Tags like [required] and [optional] are defined in How To Read Submission Lists.

Objectives

By the end of the week you can:

  • define and document an analysis cohort before summarizing;
  • report missingness for key variables rather than dropping rows silently;
  • create a Table 1-style descriptive summary with N, mean, and SD;
  • distinguish what descriptive statistics claim from what they cannot claim;
  • audit code and prose for correctness, reproducibility, interpretation, and stewardship errors.

Connection

Week 6 focused on whether visual design makes a descriptive claim honest and clear. Week 7 documents the cohort, missingness, and summary statistics behind those claims. Week 8 then turns one audited descriptive finding into a dashboard-style KT prototype.

Case Study Data Analysis

The NHANES Health Equity data spine now supports code-reading, EDA, visualization, and reproducibility checks. Use the cached CSV/RDS for routine class work; a CDC retrieval script exists for advanced users.

The default classroom path is to use the cached data so the analysis work is reproducible without internet access.

What to do here: you may run the code below — it loads the cached CSV and prints two small summaries; reading it carefully is enough unless your week’s page asks for more.

library(readr)
library(dplyr)

candidate_paths <- c(
  "examples/nhanes-equity/data/nhanes_equity_v6.csv",
  "../../examples/nhanes-equity/data/nhanes_equity_v6.csv"
)

nhanes_path <- candidate_paths[file.exists(candidate_paths)][1]
if (is.na(nhanes_path)) {
  stop("Could not find examples/nhanes-equity/data/nhanes_equity_v6.csv")
}

nhanes_analysis <- read_csv(nhanes_path, show_col_types = FALSE)

nhanes_analysis |>
  summarise(
    rows = n(),
    missing_bmi = sum(is.na(BMI)),
    missing_income = sum(is.na(IncomeGroup))
  )

nhanes_analysis |>
  filter(!is.na(BMI), !is.na(IncomeGroup), Age >= 20, Age <= 80) |>
  group_by(IncomeGroup) |>
  summarise(
    n = n(),
    mean_bmi = round(mean(BMI), 1),
    .groups = "drop"
  )

Reading order

  1. eda01 — EDA as Evidence Stewardship
  2. eda02 — Cohort and Missingness
  3. eda03 — Worked Example: EDA and Table 1
  4. eda04 — What Each Statistic Claims
  5. eda05 — Four Audit Categories for AI-Drafted EDA
  6. eda06 — In-Class Studio: Planted-Error Audit
  7. eda07 — Assignment 6 Walkthrough and Reference

Class plan

  1. Work through the worked example using the CSV snapshot.
  2. Name the cohort definition before inspecting any group summaries.
  3. Build a missingness table for the variables used in the summary.
  4. Produce a Table 1-style summary with N, BMI mean (SD), and age mean (SD).
  5. Audit the planted-error starter in the studio and categorize at least five issues.
  6. Revise the methods note so provenance, missingness, and non-causal framing are explicit.

Student output

By the end of class each student has a rerunnable EDA note, a Table 1-style summary, a planted-error audit, a provenance/stewardship note, and an AI-use note if AI helped draft code or prose.

Definition of done

  • Cohort restrictions are stated before results are interpreted.
  • Missingness is checked and reported before complete-case summaries are used.
  • Table labels match the statistic actually computed.
  • The audit identifies at least five planted issues across correctness, reproducibility, interpretation, and stewardship.
  • M2 preliminary analysis renders from source and explains what the descriptive results do not prove.

What students leave with

A reproducible EDA note and a habit of treating AI-generated summaries as drafts that must be checked against the data, code, and stewardship context.