I work on health AI: choosing meaningful research questions, defining credible evidence of progress, and building the systems to test them. My work spans foundation models, health agents, and clinical evaluation, combining medical expertise with hands-on research and engineering. I’m interested in shaping what we ask of AI in health and developing environments where models can learn to meet those demands.
Research & experience
Pheno.AI
November 2023–present
Staff Research ScientistMarch 2025–present
Data ScientistNovember 2023–February 2025
Research on the Human Phenotype Project (HPP), a prospective, deeply phenotyped cohort with approximately 28,000 enrolled participants.
PhenoBenchClinical questions as executable evaluations
- Conceived and built PhenoBench end to end: research design, task ingestion, implementation, tests, evaluations, and manuscript. Defined 90 clinically grounded tasks with explicit populations, inputs, splits, metrics, and baselines.
- Built PhenoBench-LLM to evaluate 14 language models; benchmarked tabular foundation models with collaborators, analyzing effect sizes and failure modes against conventional baselines.
- Used by all Pheno data science teams to define tasks, evaluate models, and document results that inform which models to retain.
Health AgentClinical evaluation that guides development
- Defined clinical tasks, conducted clinical review, and designed the evaluation methodology and system, with checks for numerical accuracy, evidence grounding, missing-data handling, and clinical language.
- Co-first author of a study comparing five system conditions. The research became the basis of a health-agent system now in beta with HPP participants.
- GluFormer: identified external cohorts and designed evaluations of clinical value and generalization. The study tested transfer across 19 cohorts; co-wrote the Nature paper (second author).
- HealthFormer: defined the evaluation agenda for a generative model of multimodal physiology, including UK Biobank comparisons and clinical-trial simulations assessed against published intervention outcomes.
Samsung Research collaborationFrom research proposal to pilot delivery
- Led the Galaxy Health pilot from proposal to delivery, integrating smartwatch, CGM, dietary, and clinical data from approximately 200 participants. Met or exceeded agreed health-indicator and glucose-prediction KPIs.