← Home

Research Program

Mechanism-informed, uncertainty-aware AI for clinical & biomedical data.

A single question organizes my research: how do we build models for clinical and biomedical decisions that are accurate at the edges that matter most — and that report what they do not know?

Overview

My program develops machine-learning, causal-inference, and multiscale-modeling methods for high-dimensional clinical and biomedical data — electronic health records, physiological signals, pharmacometric dynamics, and biomolecular measurements. Four themes share a common methodological core: causal reasoning, mechanistic structure, and calibrated uncertainty.

The throughline across all four is reliability: methods are most useful in medicine when they respect the mechanisms they model, generalize across sites and instruments, and quantify what they do not know. The four themes below move from clinical-AI reliability (my own forward programme) through the causal, multiscale, and reproducibility methods that support it.

Theme 01

Trustworthy & Reliable Clinical AI

Forward programme

My independent research direction — distinct from the Duke–Weill Cornell CMV project I contribute to — concerns the reliability and potential harm of perioperative and multimodal clinical AI. As Project Co-Investigator and Visiting Scholar at MIT Critical Data, I work on multi-task deep-learning models over clinical data (e.g., MIMIC-IV, MIMIC-CXR), consensus "harm" scoring, care-phenotype linkage, and subgroup and fairness auditing.

The methodological emphasis is calibration and conformal prediction under distribution shift: a model that is accurate on average but silently unreliable for specific subgroups or at clinical decision points — induction, decompensation, weaning, recovery — is not yet safe to deploy. My aim is decision support that reports its own uncertainty honestly.

Multi-task Deep LearningConformal PredictionCalibrationCare PhenotypingSubgroup & Fairness Auditing

Theme 02

Causal Inference & Bayesian Networks for Biomedical Discovery

BaMANI

BaMANI (Bayesian Multi-Algorithm Causal Network Inference) is my open-source framework for discovering reproducible causal structure in high-dimensional biomedical data. It combines constraint-based and score-based structure learning under ensemble Bayesian model averaging, with MCMC posterior exploration over graph spaces, conditional-independence testing with multiple-hypothesis correction, and bootstrapping for edge stability.

The result is a causal graph with calibrated confidence on every edge — not a single brittle point estimate. This directly addresses the instability that makes single-algorithm causal discovery unreliable on high-dimensional gene-expression and clinical data, and it grew out of my doctoral dissertation on inferring heterocellular networks in cancer.

Bayesian Structure LearningMCMCConstraint + Score-BasedConditional-Independence TestingBootstrap Stability
Placeholder · figure forthcoming

Ensemble causal-network inference

BaMANI output on high-dimensional gene-expression data, with bootstrap-stability confidence on inferred edges. The published figure will be placed here.

Theme 03

Multiscale Scientific ML for Translational Medicine

Duke–Weill Cornell

As a Postdoctoral Associate at Duke, I build reproducible frameworks that couple mechanistic simulators (agent-based models, ODEs) with machine learning: gradient-boosted trees with Bayesian hyperparameter optimization, Gaussian-process surrogates for uncertainty-aware emulation, and physics-informed neural networks that embed domain constraints.

Within the NIH-funded Duke–Weill Cornell collaboration (R01 AI173333), these methods support in-silico modeling of maternal immunity and vaccine efficacy in congenital CMV transmission. The same mechanism-informed surrogates are directly transferable to perioperative pharmacology, hemodynamics, and physiological-signal emulation.

Physics-Informed Neural NetworksNeural ODEsGaussian-Process SurrogatesAgent-Based ModelsBayesian Optimization

Theme 04

Reproducible Clinical-AI Pipelines & Cross-Dataset Harmonization

Infrastructure

Models that fail to transfer across sites and instruments are not deployable. I build end-to-end harmonization pipelines that align distributions via optimal transport with relaxed marginals on Gaussian-mixture embeddings, integrate Leiden / igraph graph clustering with Wasserstein-grounded distances, and handle quality assurance, batch correction, and feature engineering.

The work is engineered for reproducibility: containerized with Docker, scaled on SLURM / HPC, and grounded in PhysioNet-credentialed data practice with subject-level linkage and temporal alignment — consistent with multi-site clinical-data harmonization needs.

Optimal Transport (Relaxed-Marginal)Wasserstein MetricsLeiden / igraphDocker + SLURM/HPCPhysioNet-Credentialed

Methodological backbone: BaMANI ensemble Bayesian causal-discovery; OT-RMC relaxed-marginal cross-domain alignment; SciML neural emulators for mechanistic systems.

See the publications behind this program →