
Tina Hernandez-Boussard
[intermediate/advanced] The AI Lifecycle in Healthcare: Evaluation, Deployment, and Responsible Implementation
Summary
Artificial intelligence in healthcare is rapidly evolving beyond predictive models into generative, multimodal, and agentic systems. This evolution fundamentally changes not only what AI can do, but also how it must be evaluated. Traditional measures of model performance remain essential, but are insufficient to determine whether an AI system will perform safely and effectively across patients, clinical settings, and real-world environments. In this course, we will develop a practical and scientific framework for evaluating healthcare AI, moving from model performance to system-level evaluation and from static benchmarks to evidence generated across the AI lifecycle. Through clinical examples and interactive exercises, participants will learn how to identify consequential failure modes, determine what evidence is needed for a specific intended use, and design rigorous evaluation strategies for the AI systems they develop, study, or deploy.
Syllabus
We will cover core principles of AI evaluation, including discrimination and calibration, internal and external validation, robustness, generalizability, and transportability across populations and healthcare settings. We will then examine how evaluation changes as predictive models become components of larger clinical systems, where data, workflow, human interaction, and context can influence performance and safety. Building on these foundations, we will consider the evaluation of generative, multimodal, and agentic AI, including reasoning, information sufficiency, uncertainty, reliability, failure modes, human-AI interaction, and trajectories of actions. Finally, we will examine approaches for generating evidence across the AI lifecycle, from retrospective and external validation to simulation, silent evaluation, prospective studies, real-world monitoring, and re-evaluation as systems and environments change.
References
Van Calster B, Collins GS, Vickers AJ, et al. “Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: overview and guidance.” The Lancet Digital Health, vol. 7, no. 12, 100916, 2025.
Bedi S, Cui H, Fuentes M, et al. “Holistic evaluation of large language models for medical tasks with MedHELM.” Nature Medicine, vol. 32, pp. 943–951, 2026.
Handler R, Sharma S, Hernandez-Boussard T. “The fragile intelligence of GPT-5 in medicine.” Nature Medicine, vol. 31, pp. 3968–3970, 2025.
Bielick CG, Awwad A, Ellen J, et al. “Moving beyond the benchmarks: Five foundational principles for meaningful AI evaluation in healthcare.” PLOS Digital Health, vol. 5, no. 5, e0001115, 2026.
Koul A, Duran D, Hernandez-Boussard T. “Synthetic data, synthetic trust: navigating data challenges in the digital revolution.” The Lancet Digital Health, vol. 7, no. 11, 100924, 2025.
Pre-requisites
Basic familiarity with artificial intelligence or machine learning and fundamental concepts in statistics is recommended. No prior background in clinical AI deployment will be assumed.
Short bio
Dr. Tina Hernandez-Boussard is Associate Dean of Research and Professor of Medicine, Biomedical Data Science, Surgery, and Epidemiology & Population Health at Stanford University School of Medicine. An internationally recognized leader in biomedical informatics and clinical AI, her research focuses on the development, evaluation, and translation of artificial intelligence in healthcare. Her work leverages real-world clinical data to develop rigorous methods for evaluating how AI systems perform across populations, healthcare settings, and clinical environments, with an emphasis on health equity and safe and effective clinical implementation. She has published more than 250 scientific articles, is a Fellow of the American College of Medical Informatics, and serves on the Board of the Coalition for Health AI (CHAI).


















