Research Program
Mechanism-informed, uncertainty-aware AI for clinical & biomedical data.
A single question organizes my research: how do we build models for clinical and biomedical decisions that are accurate at the edges that matter most — and that report what they do not know?
Overview
My program develops machine-learning, causal-inference, and multiscale-modeling methods for high-dimensional clinical and biomedical data — electronic health records, physiological signals, pharmacometric dynamics, and biomolecular measurements. Four themes share a common methodological core: causal reasoning, mechanistic structure, and calibrated uncertainty.
The throughline across all four is reliability: methods are most useful in medicine when they respect the mechanisms they model, generalize across sites and instruments, and quantify what they do not know. The four themes below move from clinical-AI reliability (my own forward programme) through the causal, multiscale, and reproducibility methods that support it.
Theme 01
Trustworthy & Reliable Clinical AI
Forward programme
My independent research direction — distinct from the Duke–Weill Cornell CMV project I contribute to — concerns the reliability and potential harm of perioperative and multimodal clinical AI. As Project Co-Investigator and Visiting Scholar at MIT Critical Data, I work on multi-task deep-learning models over clinical data (e.g., MIMIC-IV, MIMIC-CXR), consensus "harm" scoring, care-phenotype linkage, and subgroup and fairness auditing.
The methodological emphasis is calibration and conformal prediction under distribution shift: a model that is accurate on average but silently unreliable for specific subgroups or at clinical decision points — induction, decompensation, weaning, recovery — is not yet safe to deploy. My aim is decision support that reports its own uncertainty honestly.
Theme 02
Causal Inference & Bayesian Networks for Biomedical Discovery
BaMANI
BaMANI (Bayesian Multi-Algorithm Causal Network Inference) is my open-source framework for discovering reproducible causal structure in high-dimensional biomedical data. It combines constraint-based and score-based structure learning under ensemble Bayesian model averaging, with MCMC posterior exploration over graph spaces, conditional-independence testing with multiple-hypothesis correction, and bootstrapping for edge stability.
The result is a causal graph with calibrated confidence on every edge — not a single brittle point estimate. This directly addresses the instability that makes single-algorithm causal discovery unreliable on high-dimensional gene-expression and clinical data, and it grew out of my doctoral dissertation on inferring heterocellular networks in cancer.
Ensemble causal-network inference
BaMANI output on high-dimensional gene-expression data, with bootstrap-stability confidence on inferred edges. The published figure will be placed here.
Theme 03
Multiscale Scientific ML for Translational Medicine
Duke–Weill Cornell
As a Postdoctoral Associate at Duke, I build reproducible frameworks that couple mechanistic simulators (agent-based models, ODEs) with machine learning: gradient-boosted trees with Bayesian hyperparameter optimization, Gaussian-process surrogates for uncertainty-aware emulation, and physics-informed neural networks that embed domain constraints.
Within the NIH-funded Duke–Weill Cornell collaboration (R01 AI173333), these methods support in-silico modeling of maternal immunity and vaccine efficacy in congenital CMV transmission. The same mechanism-informed surrogates are directly transferable to perioperative pharmacology, hemodynamics, and physiological-signal emulation.
Theme 04
Reproducible Clinical-AI Pipelines & Cross-Dataset Harmonization
Infrastructure
Models that fail to transfer across sites and instruments are not deployable. I build end-to-end harmonization pipelines that align distributions via optimal transport with relaxed marginals on Gaussian-mixture embeddings, integrate Leiden / igraph graph clustering with Wasserstein-grounded distances, and handle quality assurance, batch correction, and feature engineering.
The work is engineered for reproducibility: containerized with Docker, scaled on SLURM / HPC, and grounded in PhysioNet-credentialed data practice with subject-level linkage and temporal alignment — consistent with multi-site clinical-data harmonization needs.
Methodological backbone: BaMANI ensemble Bayesian causal-discovery; OT-RMC relaxed-marginal cross-domain alignment; SciML neural emulators for mechanistic systems.
See the publications behind this program →