Skip to main content

Assignment 6 Walkthrough and Reference

Where this fits

You have produced the four artifacts in class. This page is the bridge from studio drafts to the Assignment 6 submission and the Milestone 2 preliminary analysis. The deliverables here mirror the A6 README exactly — this page explains what goes into each file and how to verify it.

The six A6 deliverables, explained

Each file goes in assignments/assignment06-eda-ai-audit/submission/.

1. eda-note.qmd

A rerunnable Quarto file that loads the CSV with a relative path, defines the analysis cohort, reports missingness, produces the Table 1, and includes a short methods note. Pattern after the worked example.

  • Verify: open it in a fresh Codespace, render top to bottom, no errors.

2. eda-note.html

The rendered output of the .qmd above.

  • Verify: open the HTML in the Codespaces preview and confirm the cohort flow, missingness table, and Table 1 all appear.

3. table1.csv

The Table 1 exported as a CSV. Use write_csv(table1, "table1.csv") from inside the .qmd so it is generated by code, not by copy-paste.

  • Verify: the CSV opens, has the same row counts as the rendered table, and contains N, BMI mean (SD), and age mean (SD) by IncomeGroup.

4. audit-note.md

A short note listing at least five issues you identified in the planted-error starter (studio page). Use the structure from eda05:

Issue 1. [What it is]. This is a [category] problem because [why it matters]. Corrected: [how].
  • Verify: five issues, each tagged with one of the four categories (correctness, reproducibility, interpretation, stewardship), and each with a corrected approach.

5. provenance-stewardship-note.md

A short note suitable for reuse in M2. Use the five-sentence template from eda05: data source, source of source, access method, classroom-use boundary, one privacy caution.

  • Verify: five named items, each in a complete sentence.

6. ai-use-note.md

If you used AI to draft code or prose, name what it helped with and how you verified the result. If you did not use AI, write one sentence saying so.

Wondering how this note relates to the inline # AI prompt: / # Verified: comments from eda01? Both are expected, but only one is graded: ai-use-note.md is the graded artifact, and the inline comments are the required working practice that feeds it — copy your comment pairs into the note’s audit trail. The template and a filled example are on the assignments index.

  • Verify: specific to chunks or prompts, not generic.

Reproducibility checklist for A6

Before you click submit:

  • The .qmd renders in a fresh Codespace with no manual edits.
  • The data path is relative.
  • table1.csv is written by code (write_csv(...)).
  • Missingness is reported before the Table 1, not after.
  • The methods note avoids causal language and names the data as descriptive and unweighted.
  • Use the course devcontainer packages. If an additional package was approved, record it in your dependency note and the workspace setup so a fresh Codespace can install it.
  • All six files are committed and pushed.

A methods-note template you can reuse

We analyzed the prepared NHANES Health Equity classroom CSV snapshot. The cohort included records with age 20-80. We summarized N, BMI mean (SD), and age mean (SD) by IncomeGroup using unweighted descriptive statistics. Records with missing BMI, age, or income were excluded after the missingness counts were reported. The summary describes the prepared classroom dataset only and should not be interpreted as causal effects or population estimates.

That is five sentences. Use it as a starting point, edit to match your team’s choices.

A provenance/stewardship sentence you can reuse

The dataset is a prepared classroom snapshot derived from public-use NHANES files at the CDC, cached in the course repository for offline classroom use. It supports descriptive learning activities only — not individual-level inference, small-cell claims, or upload of row-level records to external AI tools.

Five quick knowledge checks

  1. What is a cohort, in one sentence?
  2. Why report missingness before dropping rows?
  3. When is median (IQR) more honest than mean (SD)?
  4. Name the four AI-audit categories.
  5. Which A6 file feeds the M2 methods-note.md?

EDA quick reference

Task Code Note
Load CSV read_csv("examples/nhanes-equity/data/nhanes_equity_v6.csv") Relative path
Cohort filter filter(Age >= 20, Age <= 80) Document the reason
Cohort counts tibble(step, n) table Flow diagram in text
Missingness vapply(df[vars], function(x) sum(is.na(x)), numeric(1)) Use the cohort denominator
Mean (SD) sprintf("%.1f (%.1f)", mean(x, na.rm=TRUE), sd(x, na.rm=TRUE)) One readable string
Group summary group_by(IncomeGroup) %>% summarise(...) One row per group
Export write_csv(table1, "table1.csv") Generated by code

Where this goes next

Next week (Week 8) turns the audited Table 1 from this assignment into a small dashboard-style KT product. The cleaner your A6/M2 work is, the easier the dashboard will be.