The increasing availability of complex data from multiple sources presents both new opportunities and methodological challenges for statistical analysis. This webinar brings together researchers working on the statistical foundations and emerging applications of data integration and data fusion, with a focus on how information from heterogeneous sources can be effectively combined while accounting for uncertainty, dependence, missingness, and potential sources of bias.
Prof Paul Kirk and Prof Shu Yang will discuss recent methodological developments and applications illustrating how statistical modelling can address these challenges in different settings. The talks will highlight the opportunities and challenges associated with combining complementary sources of information, as well as the role of principled statistical modelling in extracting meaningful and reliable insights from integrated data.
The webinar will provide an opportunity to explore common methodological themes across different applications and to discuss emerging directions in statistical data integration and fusion.
First talk: 2–3pm — Professor Paul Kirk
Integrative Bayesian clustering for multi-omics and clinical data
Using omics datasets to identify meaningful subgroups of patients, genes, or other biological units remains a central task in statistical omics and molecular medicine. The growing availability of diverse data types presents both challenges and opportunities for subgroup identification. In this talk I will consider how Bayesian mixture modelling can be used both to identify subgroups and to integrate multiple datasets, including approaches that share clustering structure across omics layers and approaches that guide the clustering towards clinically relevant structure by incorporating outcome information. I will also discuss how we can assess, and attempt to maximise, the clinical relevance of the clusters identified, and touch on practical considerations for applying these methods at scale.
Second talk: 3–4pm — Professor Shu Yang
Integrating diverse evidence sources in clinical research: bridging RCTs and RWD - The interface of statistics, AI, and real-world evidence
Randomized clinical trials remain the gold standard for estimating treatment effects, but they are often small, selective, and short in follow-up. Real-world data from electronic health records, claims, and registries can fill some of these gaps, and AI/ML tools are increasingly used to extract features and build flexible models. The hard problem is not access to more data; it is combining sources without letting bias from observational data undermine valid inference.
This talk discusses how to integrate randomized trials and real-world data with statistical care. I will start with the roles of trials, real-world evidence, and AI in today’s regulatory landscape, then focus on hybrid controlled trials that borrow external or real-world controls to improve efficiency. A central question is how to borrow comparable external controls and down-weight or discard those that are not comparable. I will present a set of bias-aware methods, from simple test-then-pool rules to selective borrowing and randomization-based tests, that aim to gain power while protecting type-I error. I will close with what is ready for practice, and what remains open.
Paul Kirk is Research Professor in Biostatistical Machine Learning at the MRC Biostatistics Unit, University of Cambridge, where he co-leads the Biostatistical Machine Learning research theme. His research is at the intersection of Bayesian machine learning, multi-omics data integration and molecular precision medicine, with a current focus on reproductive and perinatal medicine. He has previously held positions at Oxford, Imperial and Warwick, and has made contributions to statistical systems biology, integrative clustering and proteomics applications.
Shu Yang is a Professor of Statistics, Goodnight Early Career Innovator, and University Faculty Scholar at North Carolina State University. She received her Ph.D. in Applied Mathematics and Statistics from Iowa State University and completed her postdoctoral training at the Harvard T.H. Chan School of Public Health. Her research focuses on causal inference, real-world evidence, and data integration, particularly in the context of comparative effectiveness research in health studies. Dr. Yang has served as Principal Investigator on multiple large-scale research grants from the NSF, NIH (R01), and FDA (U01). She is a recipient of the Committee of Presidents of Statistical Societies (COPSS) Emerging Leader Award and an elected Fellow of the American Statistical Association (ASA). Website:
https://shuyang.wordpress.ncsu.edu/.
Free to RSS members
£10 for non-members