• Uncertainty Quantification for Interpretable Machine Learning | Lili Zheng

    Special Seminar Series

    Interpretable machine learning has been widely deployed for scientific discoveries and decision-making, while its reliability hinges on the critical role of uncertainty quantification (UQ). In this talk, I will discuss UQ in two challenging scenarios motivated by scientific and societal applications: selective inference for large-scale graph learning and UQ for model-agnostic machine learning interpretations. Specifically, the first part concerns graphical model inference when only irregular, patchwise observations are available, a common setting in neuroscience, healthcare, genomics, and econometrics. To filter out low-confidence edges due to the irregular measurements, I will present a novel inference method that quantifies the uneven edgewise uncertainty levels over the graph as well as an FDR control procedure; this is achieved by carefully disentangling the dependencies across the graph and consequently yields more reliable graph selection. In the second part, I will discuss the computational and statistical challenges associated with UQ for feature importance of any machine learning model. I will take inspiration from recent advances in conformal inference and utilize an ensemble framework to address these challenges. This leads to an almost computationally free, assumption-light, and statistically powerful inference approach for occlusion-based feature importance. For both parts of the talk, I will highlight the potential applications of my research in science and society as well as how it contributes to more reliable and trustworthy data science.

  • Inference and Decision-Making amid Social Interactions | Shuangning Li

    Special Seminar Series
    Computer Science & Engineering Building (CSE), Room 4140 3234 Matthews Ln, La Jolla, CA, United States

    From social media trends to family dynamics, social interactions shape our daily lives. In this talk, I will present tools I have developed for statistical inference and decision-making in light of these social interactions.

    (1) Inference: I will talk about estimation of causal effects in the presence of interference. In causal inference, the term “interference” refers to a situation where, due to interactions between units, the treatment assigned to one unit affects the observed outcomes of others. I will discuss large-sample asymptotics for treatment effect estimation under network interference where the interference graph is a random draw from a graphon. When targeting the direct effect, we show that popular estimators in our setting are considerably more accurate than existing results suggest. Meanwhile, when targeting the indirect effect, we propose a consistent estimator in a setting where no other consistent estimators are currently available.

    (2) Decision-Making: Turning to reinforcement learning amid social interactions, I will focus on a problem inspired by a specific class of mobile health trials involving both target individuals and their care partners. These trials feature two types of interventions: those targeting individuals directly and those aimed at improving the relationship between the individual and their care partner. I will present an online reinforcement learning algorithm designed to personalize the delivery of these interventions. The algorithm's effectiveness is demonstrated through simulation studies conducted on a realistic test bed, which was constructed using data from a prior mobile health study. The proposed algorithm will be implemented in the ADAPTS HCT clinical trial, which seeks to improve medication adherence among adolescents undergoing allogeneic hematopoietic stem cell transplantation.

  • Statistical insight for biomedical data science, with applications to single-cell RNA-sequencing data | Yiqun Chen

    Special Seminar Series

    My research centers around bringing statistical insights and understanding to the practice of modern data science, and I will cover two projects related to this research vision in this talk.

    The first part of the talk is motivated by the practice of testing data-driven hypotheses. In the biomedical sciences, it has become increasingly common to collect massive datasets without a pre-specified research question. In this setting, a data analyst might use the data both to generate a research question, and to test the associated null hypothesis. For example, in single-cell RNA-sequencing analyses, researchers often first cluster the cells, and then test for differences in the expected gene expression levels between the clusters to quantify up- or down-regulation of genes, annotate known cell types, and identify new cell types. However, this popular practice is invalid from a statistical perspective: once we have used the data to generate hypotheses, standard statistical inference tools are no longer valid. To tackle this problem, I developed a conditional selective approach to test for a difference in means between pairs of clusters obtained via k-means clustering.

    The proposed approach has appropriate statistical guarantees (e.g., selective Type 1 error control). In the second part of the talk, I will consider how to leverage large language models (LLMs) such as ChatGPT for biomedical discovery. While significant progress has been made in customizing large language models for biomedical data, these models often require extensive data curation and resource-intensive training. In the context of single-cell RNA-sequencing data, I will show that we can achieve surprisingly competitive results on many downstream tasks via a much simpler alternative: I input textual descriptions of genes into an off-the-shelf LLM, such as ChatGPT, to obtain low-dimensional representations of the genes, or “embeddings.” I then use these embeddings as features in downstream tasks. A similar approach enables LLM-derived embeddings of cells. This work highlights the potential of LLMs to provide meaningful and concise representations for biomedical data, and also raises a number of challenging statistical questions. Addressing these questions requires bringing principled statistical thinking to the practice of modern data science.

  • Algorithm Dynamics in Modern Statistical Learning: Universality and Implicit Regularization | Tianhao Wang

    Special Seminar Series

    Modern statistical learning is featured by the high-dimensional nature of data and over-parameterization of models. In this regime, analyzing the dynamics of the used algorithms is challenging but crucial for understanding the performance of learned models. This talk will present recent results on the dynamics of two pivotal algorithms: Approximate Message Passing (AMP) and Stochastic Gradient Descent (SGD). Specifically, AMP refers to a class of iterative algorithms for solving large-scale statistical problems, whose dynamics admit asymptotically a simple but exact description known as state evolution. We will demonstrate the universality of AMP's state evolution over large classes of random matrices, and provide illustrative examples of applications of our universality results. Secondly, for SGD, a workhorse for training deep neural networks, we will introduce a novel mathematical framework for analyzing its implicit regularization. This is essential for SGD's ability to find solutions with strong generalization performance, particularly in the case of over-parameterization. Our framework offers a general method to characterize the implicit regularization induced by gradient noise. Finally, in the context of underdetermined linear regression, we will show that both AMP and SGD can provably achieve sparse recovery, yet they do so from markedly different perspectives.

  • Women in Data Science Social | RSVP – Space is Limited

    Women in Data Science
    Zanzibar 9500 Gilman Drive, La Jolla, CA, United States

    Come to the 1st WDS Lunch Social to connect with other women researchers and students working in machine learning, AI, and data science related fields. Lunch will be provided. RSVP is required as space is limited. RSVP at QR code in image.

  • Universal Learning for Decision-Making | Moise Blanchard

    Special Seminar Series

    We provide general-use decision-making algorithms under provably minimal assumptions on the data, using the universal learning framework. Classically, learning guarantees typically require two types of assumptions: (1) restrictions on target policies to be learned and (2) assumptions on the data-generating process. Instead, we show that we can provide consistent algorithms with vanishing regret compared to the best policy in hindsight, (1) irrespective of the optimal policy, known as universal consistency, and (2) well beyond standard i.i.d. or stationary assumptions on the data. We present our results for the classical online regression problem as well as for the contextual bandit problem, where the learner's rewards depend on their selected actions and an observable context. This generalizes the standard multi-armed bandit to the case where side information is available, e.g., patients' records or customers' history, which allows for personalized treatment. Precisely, we give necessary and sufficient conditions on the context-generating process for universal consistency to be possible. Surprisingly, for finite action spaces, universally learnable processes are the same for contextual bandits as for the supervised learning setting, suggesting that going from full feedback (supervised learning) to partial feedback (contextual bandits) came at no extra cost in terms of learnability. We then show that there always exists an algorithm that guarantees universal consistency whenever this is achievable. In particular, such an algorithm is universally consistent under provably minimal assumptions: if it fails to be universally consistent for some context-generating process, then no other algorithm would succeed either. In the case of finite action spaces, this algorithm balances a fine trade-off between generalization (similar to structural risk minimization) and personalization (tailoring actions to specific contexts).

  • Promoting fairness, inclusivity, and excitement in undergraduate education through teaching and research | Josh Grossman

    Special Seminar Series

    This talk has three parts. First, I'll conduct an interactive, undergraduate-level introduction to k-means clustering. I'll then outline my pedagogy, making specific reference to strategies employed in the teaching demo. Finally, I'll discuss my research on fairness in college admissions and increasing access to higher education. 

  • Lessons from the deep: engineering biosensors, workflows, and visualizations for communication and collaboration in comparative medicine and climate science | Jessica Kendall-Bar

    Special Seminar Series
    Powell-Focht Bioengineering Hall (PFBH), FUNG Auditorium

    Abstract: Effective conservation and management relies on an in-depth understanding of the health of marine ecosystems. Dr. Kendall-Bar's interdisciplinary approach combines engineering, visualization, and computation to study ocean resilience in terms of the extreme physiology and behavior of marine animals, establishing eco-physiological baselines to track over time in the face of climate change. This seminar and chalk talk will review her work to create innovative tools to detect, visualize, and analyze the physiology and behavior of animals in extreme environments that showcase their biological resilience to oxygen and sleep deprivation. From individuals to ecosystems, Kendall-Bar conducts multidisciplinary physiological studies that combine basic and applied science with potential to advance conservation and comparative medicine. This seminar reviews Kendall-Bar's dissertation research on sleep in seals and presents some current and ongoing projects to combine high-performance computing, automation, and visualization to assess diving physiology in human freedivers, epilepsy in sea lions, and cardiac performance in some of the largest (blue whales) and smallest (emperor penguins) divers. Kendall-Bar’s newest projects involve novel data visualizations and science communication to inform research as well as international policy in domains ranging from marine mammal conservation to traditional ecological knowledge and coral reef restoration.

  • Algebraic vision: A gentle introduction | Jessie Loucks-Tavitas

    Special Seminar Series

    Abstract:
    My talk will be broken into three parts:
    Part I: Meet Jessie.
    Part II: Assessing Deep Learning Models. A short lesson on assessment criteria for deep learning models, such as LLMs and image segmentation models.
    Part III: Algebraic Vision, a Gentle Introduction. Algebraic vision, lying in the intersection of computer vision and projective geometry, is the study of 3D objects being photographed by multiple cameras, using techniques found in computational algebraic geometry. Two natural questions arise: (1) Given a 3D object and multiple images of it, can we determine the relative camera positions? And, (2) given multiple images as well as relative camera locations, can we reconstruct the object being photographed? Carlsson and Weinshall showed in 1998 that the algorithms to solve these problems are intrinsically connected. A beneficial corollary of recent joint work with Erin Connelly and Timothy Duff is a formalization of this “duality” mechanism. We will discuss this formalization, along with some future directions that we hope to venture down.