Skip to main content

One Step, One Cell

Why This Rule Exists

Large AI-generated code blocks are hard to audit. A single hidden mistake in loading, filtering, grouping, or missingness handling can cascade into a polished but incorrect result.

Use one cell per analytic step:

  1. load data;
  2. inspect rows and columns;
  3. filter missing values;
  4. group and summarise;
  5. compare with the R result.

Good AI Prompt

Translate only the data-loading step from R to Python/pandas. Tell me what output I should inspect before moving to the next step.

Risky AI Prompt

Translate my full analysis, make the plots, and write the interpretation.

The risky prompt gives up too much control. It may change the analytic question without telling you.

Parity Audit Checks

Check R Python
Row count nrow(df) len(df) or df.shape[0]
Column names names(df) df.columns.tolist()
Unique levels unique(df$IncomeGroup) df["IncomeGroup"].unique()
Missing values sum(is.na(df$BMI)) df["BMI"].isna().sum()
Summary table rows nrow(summary_df) len(summary_df)

Notebook Habit

After each Python cell, add a short Markdown note:

Audit check: this cell matches the R output because ...

If you cannot fill in that sentence, do not move to the next cell yet.