Skip to main content

Health Data Science: AI and Knowledge Translation

Authors
Affiliations

M. Ehsan. Karim

School of Population and Public Health, The University of British Columbia

AI-KT team

AI-KT team includes - Manya Jain, Rainie Fu, Md. Belal Hossain

Published

September 15, 2026

The Project

Welcome to the course notes for Health Data Science: AI and Knowledge Translation.

This platform bridges the gap between traditional epidemiology and modern data science. Unlike standard textbooks that focus on a single language, this resource adopts a “Polyglot” approach—teaching R and Python side-by-side—and integrates Artificial Intelligence directly into the workflow.

Whether you are a health researcher looking to modernize your skills or a data scientist seeking to understand epidemiological rigor, you are in the right place.

Here, we offer:

  • “Zero to AI Co-Pilot” Training: Learn to use Large Language Models (LLMs) to generate, debug, and translate code while critically auditing them for bias and hallucinations.
  • Real-World Data: Guides on navigating “messy” open data sources such as NHANES and WHO, rather than sanitized “toy” datasets.
  • Knowledge Translation: Step-by-step walkthroughs on building a Quarto dashboard-style KT product and dynamic reports; Shiny is shown as an optional demonstration.

What We Aim to Achieve

We are on a mission to:

  • Democratize Compute Access: Eliminate hardware barriers by using GitHub Codespaces — a cloud-native environment that runs on any modern web browser, with no local installation required.
  • Polyglot Awareness: R is the required core. Week 5 reads short Python/pandas examples and runs AI-assisted translation parity checks; Python is awareness, not a parallel curriculum.
  • Teach AI Audit Habits: Shift the focus from accepting AI-generated code to rigorous code review, security validation, and reproducibility.
  • Enable Professional Portfolios: Help students transform static analyses into public-facing data science portfolios.

Dive into Our Modules

Embark on a journey through our structured modules, designed to take you from data literacy to advanced AI workflows.

Module Topics Descriptions
1 Week 0: Onboarding Account setup, orientation, and pre-course survey.
2 Week 1: Health Data, KT & Ethics Open data, NHANES, knowledge translation, ethics, privacy, and data stewardship.
3 Week 2: Modern Research Workflows Codespaces, VS Code, Quarto, project structure, rendering, commits, and syncing.
4 Week 3: R with AI R fundamentals, tidyverse wrangling, functions, debugging with AI, and reproducible notes.
5 Week 4: Git & Collaboration Git provenance, branches, pull requests, peer review, repo hygiene, and group setup.
6 Week 5: Polyglot Awareness & R Deepening Python/pandas awareness, R-to-Python translation, AI parity checks, and R mastery practice.
7 Week 6: Data Visualization Clear health data graphics, ggplot2 workflows, visual audits, and misleading plot repair.
8 Week 7: EDA, Table 1 & AI Auditing EDA, missingness, Table 1 summaries, planted-error audits, and descriptive interpretation.
9 Week 8: Dashboard Prototypes for KT Dashboard-style knowledge translation products from one audited descriptive finding; Quarto starter is the required core, Shiny is shown only as an instructor demonstration.
10 Week 9: Communication of Scientific Findings Plain-language communication, revealjs slides, citations, and peer feedback.
11 Week 10: Writing & Publishing Reports Dynamic Quarto reports, figures, tables, citations, limitations, and publishing workflow.
12 Week 11: Portfolio Surgery Fresh-environment reproducibility testing, dependency repair, portfolio polish, and static dashboard extension.
13 Project Presentation (with Final Portfolio) Narrated slide deck submitted online with the Final Portfolio.
Note

The tutorials are designed with a consistent structure to provide a cohesive learning experience. Here is what you can expect in each chapter:

  1. Overview: A concise summary of learning objectives and topics.
  2. Concepts: Core theoretical materials (slides, video lessons).
  3. R Tutorials: In-depth R walkthroughs with AI-assisted code drafting and auditing. Week 5 adds short Python/pandas examples for polyglot awareness only.
  4. AI Auditing: Specific sections dedicated to verifying and “stress-testing” AI-generated code for the chapter’s topic.
  5. Knowledge Check: Ungraded self-assessments to reinforce key concepts.
  6. Practice Exercises: Hands-on tasks using real-world health data to build your portfolio.

How Our Content is Presented

All our resources are hosted on an easy-to-access GitHub page. The format? Engaging text, reproducible code, clear outputs, and crisp videos that distill complex topics.

This document is created using Quarto, allowing us to weave analysis, code, and interpretation into a single reproducible document.

How to Cite

Style Citation
APA Karim, M.E. et al. (2026). Health Data Science: AI and Knowledge Translation. Retrieved from https://ehsanx.github.io/HDSx/
Vancouver Karim, M.E. et al. Health Data Science: AI and Knowledge Translation. [Internet]. 2026 [cited August 31, 2026]. Available from: https://ehsanx.github.io/HDSx/
IEEE Karim, M.E. et al., "Health Data Science: AI and Knowledge Translation," 2026. [Online]. Available: https://ehsanx.github.io/HDSx/. Accessed on: August 31, 2026.

The BibTeX format can be downloaded from here.