Stanford Root

Schedule

Stanford Root

Schedule

CS 287

Data Centric AI for Healthcare (BMDS 218)

UNITS:3
GRADING:Medical Option (Med-Ltr-CR/NC)
LEVEL:Graduate
GER:—

This course explores how healthcare data is generated in practice - and how these processes influence the development of robust, trustworthy machine learning models. Students will examine how clinical workflows, documentation habits, and institutional practices introduce bias, noise, missingness, and spurious correlations into datasets. The course emphasizes the often-overlooked reality that many model failures stem not from algorithmic design, but from subtle flaws in how data is collected, labeled, and interpreted. Through lectures, real-world case studies, and observation of clinical data workflows, students will develop practical tools for understanding and improving data quality and model reliability. Topics include assessing data quality, addressing missingness and label leakage, developing scalable and accurate annotation strategies, leveraging synthetic data, and monitoring data and models for drift over time. Students will also learn to evaluate how these data challenges impact model generalization and robustness.

Syllabus for selected term:
View Winter 2027 Syllabus

Sections

1 Term
Lecture 1Open
ID: 6621
0 / 50 enrolled
DAYS:Monday, Wednesday
TIME:10:30 AM – 11:50 AM
LOCATION:TBD
INSTRUCTOR:
Alsentzer, Emily, Fries, Jason, Shen, Andrew
3units

CS 287: Data Centric AI for Healthcare (BMDS 218)

3 units · Medical Option (Med-Ltr-CR/NC)

This course explores how healthcare data is generated in practice - and how these processes influence the development of robust, trustworthy machine learning models. Students will examine how clinical workflows, documentation habits, and institutional practices introduce bias, noise, missingness, and spurious correlations into datasets. The course emphasizes the often-overlooked reality that many model failures stem not from algorithmic design, but from subtle flaws in how data is collected, labeled, and interpreted. Through lectures, real-world case studies, and observation of clinical data workflows, students will develop practical tools for understanding and improving data quality and model reliability. Topics include assessing data quality, addressing missingness and label leakage, developing scalable and accurate annotation strategies, leveraging synthetic data, and monitoring data and models for drift over time. Students will also learn to evaluate how these data challenges impact model generalization and robustness.

Offered in Winter 2027 at Stanford University.

Winter 2027 sections

  • Lecture — Monday Wednesday 10:30 AM – 11:50 AM — Alsentzer, Emily, Fries, Jason, Shen, Andrew (Graduate)

More CS courses

  • CS 278: Social Computing (SOC 174, SOC 274)
  • CS 279: Computational Biology: Structure and Organization of Biomolecules and Cells (BIOE 279, BIOPHYS 279, BMDS 245, CME 279)
  • CS 281: Ethics of Artificial Intelligence
  • CS 282: Computer Systems Architecture (EE 282)
  • CS 283: Governing Artificial Intelligence: Law, Policy, and Institutions (COMM 152A, COMM 252A, GLOBAL 245B, INTLPOL 245B, POLISCI 145B, POLISCI 445B)
  • CS 286: Advanced Topics in Computer Vision and Biomedicine (BMDS 276)
  • CS 288: Applied Causal Inference with Machine Learning and AI (MS&E 228)
  • CS 293: Empowering Educators via Language Technology (EDUC 473)
  • CS 295: Software Engineering
  • CS 298: Seminar on Teaching Introductory Computer Science (EDUC 298)
  • CS 300: Departmental Lecture Series
  • CS 309A: Cloud Computing Seminar

All CS courses · All departments