Skip to main content

AI Prompts You Can Defend

Where this fits

You can build a ggplot from scratch (viz02) and you can spot six common failure modes (viz04). This page connects the two: how to ask an AI co-pilot for plot code without inheriting its bad defaults, and how to audit what it gives you.

Why this matters

If you paste “make me a nice plot of NHANES” into a chat window, the model will produce something — possibly with a flashy palette, a trend line, and a confident title. None of those choices were yours. None of them were audited. By the time the figure is in your report, you own them anyway.

The fix is to write prompts that constrain the model instead of decorate the request.

Three prompt templates that work

Template 1 — Start from the data, not the picture

I have a data frame called bmi_by_income_cycle with columns Cycle, IncomeGroup, n, and mean_bmi. Write a ggplot2 figure that shows mean_bmi on the y-axis and Cycle on the x-axis, with one line per IncomeGroup. The summary is unweighted and descriptive. Do not add a trend line or smoothing.

What this prompt does well: it names the data, the variables, the geometry, and what not to add.

Template 2 — Ask for an honest scale and labels

Add labs(...) to the figure above so the title describes the comparison without making a causal claim, the subtitle states that the summary is unweighted and uses adults age 20-80, and the caption discloses that records with missing BMI, income, or cycle were excluded.

What this prompt does well: it tells the model what each label should do, not just that there should be labels.

Template 3 — Ask for a caption that names the limitation

Write a one-sentence caption for this figure that names the statistic (mean BMI), the comparison (across NHANES cycles, by income group), the data source (prepared NHANES Health Equity classroom CSV), the exclusions (missing values, ages outside 20-80), and one limitation (descriptive only, not a population estimate).

What this prompt does well: the model now has a checklist of what the caption must contain.

Auditing what the model gives you

After you paste in any AI-drafted plot code, walk through this checklist before you accept it:

  1. Column names match the data. Run glimpse(your_data) and compare.
  2. Geometry is appropriate. No smoothing across categorical cycles. No regression lines on a descriptive summary.
  3. Scale is honest. Does the y-axis start at 0, or is there a stated reason it does not?
  4. Title and subtitle are descriptive. No “X causes Y” wording.
  5. Caption discloses exclusions and limitation. Missing values, cohort restrictions, weighted vs unweighted.
  6. na.rm = TRUE is present wherever the summary uses a numeric column that may have NA.

If any item fails, edit the code or the labels before rendering.

The paper-trail rule (still in force)

From Week 3: every AI-drafted chunk in your .qmd gets a two-line comment.

# AI prompt: "Plot mean_bmi by Cycle, one line per IncomeGroup, no smoothing"
# Verified: columns match glimpse(); no smoother; descriptive title; caption lists exclusions
bmi_plot <- ggplot(bmi_by_income_cycle, aes(...)) + ...

This is the same audit log you used in Week 3 for data wrangling. The course AI policy requires it for graded work.

A note on what NOT to ask the AI

  • Do not paste row-level NHANES data into an external chat window. Aggregate first.
  • Do not ask the AI to “decide” whether income causes BMI. The figure is descriptive; causal language is your responsibility to remove.
  • Do not accept a “looks fine to me” from the model as audit evidence. The audit is yours.

Where this goes next

viz06 is the in-class studio where you apply all of this to a deliberately flawed plot.