Research
I grew up in Wuhan, split by the Yangtze River, and earned my PhD in Pittsburgh, where the Allegheny and Monongahela meet to form the Ohio. In both cities, bridges shape daily life, keeping people, goods, and information moving.
I see my work in biostatistics in a similar way. Statistics and artificial intelligence provide tools for learning from complex data, while biology and clinical medicine give us the questions that guide this learning. Our research connects the two. We aim to produce findings that researchers can investigate and predictions that clinicians can understand, evaluate, and use.
My research focuses on generalizable statistical learning for biomedical data science. We study how to learn disease-related information from heterogeneous data and determine when it remains informative across populations, measurement platforms, and clinical settings. Our work in cancer biology and metabolic disease informs the statistical questions we pursue. The methods we develop, in turn, help us investigate disease across different sources of evidence.
Research Interests
Generalizable Statistical Learning.
Biomedical data reflect both disease and how it is observed. Differences between datasets may arise from biology, patient selection, measurement, or clinical practice. These sources of variation can be difficult to separate, particularly when biological differences coincide with differences in how studies were conducted. We develop statistical methods to account for these sources of variation while preserving information relevant to the biomedical question. We examine which distinctions the data support and how conclusions change under different assumptions.
We study how information can be shared across studies at the level of individual measurements, relationships among measurements, or their associations with outcomes. We ask what should be learned for a particular purpose, which evidence can be combined across sources, and whether the resulting findings or predictions remain useful in new samples and settings. When they do not, we seek to understand why and what needs to change. Across this work, we state assumptions clearly, quantify uncertainty, evaluate generalizability, and make our analyses reproducible.
Cancer Biology and Progression.
Our cancer research focuses on the biological processes involved in recurrence and metastasis, with a particular interest in breast cancer. We study why tumors with similar features at diagnosis follow different courses after treatment. We examine how cancer cells persist, interact with surrounding tissue, and grow at distant sites.
We use omics data to characterize the molecular programs associated with these processes. A molecular signal may reflect changes within cancer cells, differences in the cells present in a sample, or both. We ask which programs are shared across tumors and how their relationships with disease depend on biological context. These questions motivate methods for comparing evidence across studies and assessing how biological and measurement differences affect the interpretation of molecular findings.
We also investigate which aspects of patient tumor biology are preserved in cell lines, organoids, and animal models. The suitability of a model depends on the biological process being studied and the cellular context it requires. We use these comparisons to guide the selection of experimental models for testing hypotheses about progression and treatment response.
Metabolic Health and Disease.
Our metabolic research examines how disturbances in metabolic regulation develop, persist, and resolve, and why disease courses differ between individuals. We study how metabolic abnormalities occur together and change over time, and how these changes relate to disease onset, progression, and recovery. We seek to understand how these disease courses are reflected in clinical measurements.
We use population cohorts and longitudinal clinical data to investigate these questions through relationships among measurements and changes within individuals. We ask which patterns reflect common features of disease progression and which vary with the population or clinical context. Our methodological work supports these comparisons by accounting for differences in observation and care and assessing whether the resulting patterns remain informative across clinical settings.