One Step, One Cell
Why This Rule Exists
Large AI-generated code blocks are hard to audit. A single hidden mistake in loading, filtering, grouping, or missingness handling can cascade into a polished but incorrect result.
Use one cell per analytic step:
- load data;
- inspect rows and columns;
- filter missing values;
- group and summarise;
- compare with the R result.
Good AI Prompt
Translate only the data-loading step from R to Python/pandas. Tell me what output I should inspect before moving to the next step.
Risky AI Prompt
Translate my full analysis, make the plots, and write the interpretation.
The risky prompt gives up too much control. It may change the analytic question without telling you.
Parity Audit Checks
| Check | R | Python |
|---|---|---|
| Row count | nrow(df) |
len(df) or df.shape[0] |
| Column names | names(df) |
df.columns.tolist() |
| Unique levels | unique(df$IncomeGroup) |
df["IncomeGroup"].unique() |
| Missing values | sum(is.na(df$BMI)) |
df["BMI"].isna().sum() |
| Summary table rows | nrow(summary_df) |
len(summary_df) |
Notebook Habit
After each Python cell, add a short Markdown note:
If you cannot fill in that sentence, do not move to the next cell yet.