Stanford Root

Schedule

Stanford Root

Schedule

BMDS 218

Data Centric AI for Healthcare (CS 287)

UNITS:3
GRADING:Medical Option (Med-Ltr-CR/NC)
LEVEL:Graduate
GER:—

This course explores how healthcare data is generated in practice - and how these processes influence the development of robust, trustworthy machine learning models. Students will examine how clinical workflows, documentation habits, and institutional practices introduce bias, noise, missingness, and spurious correlations into datasets. The course emphasizes the often-overlooked reality that many model failures stem not from algorithmic design, but from subtle flaws in how data is collected, labeled, and interpreted. Through lectures, real-world case studies, and observation of clinical data workflows, students will develop practical tools for understanding and improving data quality and model reliability. Topics include assessing data quality, addressing missingness and label leakage, developing scalable and accurate annotation strategies, leveraging synthetic data, and monitoring data and models for drift over time. Students will also learn to evaluate how these data challenges impact model generalization and robustness.

Syllabus for selected term:
View Winter 2027 Syllabus

Sections

1 Term
Lecture 1Open
ID: 23329
0 / 50 enrolled
DAYS:Monday, Wednesday
TIME:10:30 AM – 11:50 AM
LOCATION:TBD
INSTRUCTOR:
Alsentzer, Emily, Fries, Jason, Shen, Andrew
3units

BMDS 218: Data Centric AI for Healthcare (CS 287)

3 units · Medical Option (Med-Ltr-CR/NC)

This course explores how healthcare data is generated in practice - and how these processes influence the development of robust, trustworthy machine learning models. Students will examine how clinical workflows, documentation habits, and institutional practices introduce bias, noise, missingness, and spurious correlations into datasets. The course emphasizes the often-overlooked reality that many model failures stem not from algorithmic design, but from subtle flaws in how data is collected, labeled, and interpreted. Through lectures, real-world case studies, and observation of clinical data workflows, students will develop practical tools for understanding and improving data quality and model reliability. Topics include assessing data quality, addressing missingness and label leakage, developing scalable and accurate annotation strategies, leveraging synthetic data, and monitoring data and models for drift over time. Students will also learn to evaluate how these data challenges impact model generalization and robustness.

Offered in Winter 2027 at Stanford University.

Winter 2027 sections

  • Lecture — Monday Wednesday 10:30 AM – 11:50 AM — Alsentzer, Emily, Fries, Jason, Shen, Andrew (Graduate)

More BMDS courses

  • BMDS 210: Modeling Biomedical Systems (CS 270)
  • BMDS 212: Introduction to Biomedical Informatics Research Methodology (BIOE 212, CS 272, GENE 212)
  • BMDS 214: Representations and Algorithms for Computational Molecular Biology (BIOE 214, CS 274, GENE 214)
  • BMDS 215: Data Science for Medicine
  • BMDS 216: Representations and Algorithms for Molecular Biology: Lectures
  • BMDS 217: Translational Bioinformatics (BIOE 217, CS 275, GENE 217)
  • BMDS 219: Mathematical Models and Medical Decisions
  • BMDS 221: Machine Learning Approaches for Data Fusion in Biomedicine
  • BMDS 222: Cloud Computing for Biology and Healthcare (CS 273C, GENE 222)
  • BMDS 223: Deploying and Evaluating Fair AI in Healthcare (CSRE 323, EPI 220)
  • BMDS 224: Principles of Pharmacogenomics (GENE 224)
  • BMDS 236: Introduction to Cost-Effectiveness Analysis: Evaluating Benefits and Costs of Health Interventions (HRP 392)

All BMDS courses · All departments