Filters

Changing any of the form inputs will cause the list of events to refresh with the filtered results.

  • Lessons from the deep: engineering biosensors, workflows, and visualizations for communication and collaboration in comparative medicine and climate science | Jessica Kendall-Bar

    Special Seminar Series
    Powell-Focht Bioengineering Hall (PFBH), FUNG Auditorium

    Abstract: Effective conservation and management relies on an in-depth understanding of the health of marine ecosystems. Dr. Kendall-Bar's interdisciplinary approach combines engineering, visualization, and computation to study ocean resilience in terms of the extreme physiology and behavior of marine animals, establishing eco-physiological baselines to track over time in the face of climate change. This seminar and chalk talk will review her work to create innovative tools to detect, visualize, and analyze the physiology and behavior of animals in extreme environments that showcase their biological resilience to oxygen and sleep deprivation. From individuals to ecosystems, Kendall-Bar conducts multidisciplinary physiological studies that combine basic and applied science with potential to advance conservation and comparative medicine. This seminar reviews Kendall-Bar's dissertation research on sleep in seals and presents some current and ongoing projects to combine high-performance computing, automation, and visualization to assess diving physiology in human freedivers, epilepsy in sea lions, and cardiac performance in some of the largest (blue whales) and smallest (emperor penguins) divers. Kendall-Bar’s newest projects involve novel data visualizations and science communication to inform research as well as international policy in domains ranging from marine mammal conservation to traditional ecological knowledge and coral reef restoration.

  • Scaling Data-Constrained Language Model

    EnCORE Series
    Virtual

    Extrapolating scaling trends suggest that training dataset size for LLMs may soon be limited by the amount of text data available on the internet. In this talk we investigate scaling language models in data-constrained regimes. Specifically, we run a set of empirical experiments varying the extent of data repetition and compute budget. From these experiments we propose and empirically validate a scaling law for compute optimality that accounts for the decreasing value of repeated tokens and excess parameters. Finally, we discuss and experiment with approaches for mitigating data scarcity.

  • Algebraic vision: A gentle introduction | Jessie Loucks-Tavitas

    Special Seminar Series

    Abstract:
    My talk will be broken into three parts:
    Part I: Meet Jessie.
    Part II: Assessing Deep Learning Models. A short lesson on assessment criteria for deep learning models, such as LLMs and image segmentation models.
    Part III: Algebraic Vision, a Gentle Introduction. Algebraic vision, lying in the intersection of computer vision and projective geometry, is the study of 3D objects being photographed by multiple cameras, using techniques found in computational algebraic geometry. Two natural questions arise: (1) Given a 3D object and multiple images of it, can we determine the relative camera positions? And, (2) given multiple images as well as relative camera locations, can we reconstruct the object being photographed? Carlsson and Weinshall showed in 1998 that the algorithms to solve these problems are intrinsically connected. A beneficial corollary of recent joint work with Erin Connelly and Timothy Duff is a formalization of this “duality” mechanism. We will discuss this formalization, along with some future directions that we hope to venture down.

  • Inference in context: Statistical theory and thinking | Jeffrey Bye

    Special Seminar Series

    Abstract: The likelihood function plays a foundational role in statistical theory. I will demonstrate my teaching philosophy and approach through a lesson on maximum likelihood estimation and its connection to Neyman-Pearson, Bayesian, and other approaches to statistical and scientific inference. I will then expand on the role of context in statistical thinking, particularly how it informs my scholarship on how people learn about data, math, statistics, and programming.

  • EnCORE : Theoretical Exploration of Foundation Model Adaptation, Kangwook Lee, UW Madison, Feb 9th, 1-2pm

    EnCORE Series
    Atkinson Hall, Fourth Floor

    Abstract: Due to the enormous size of foundation models, various new methods for efficient model adaptation have been developed. Parameter-efficient fine-tuning (PEFT) is an adaptation method that updates only a tiny fraction of the model parameters, leaving the remainder unchanged. In-context Learning (ICL) is a test-time adaptation method, which repurposes foundation models by providing them with labeled samples as part of the input context. Given the growing importance of this emerging paradigm, developing theoretical foundations for the new paradigm is of utmost importance.

  • Integrating Longitudinal Multimodal Data To Realize Precision Medicine | Samantha Piekos

    Special Seminar Series
    Halıcıoğlu Data Science Institute (HDSI), Room 123 3234 Matthews Ln, La Jolla, CA, United States

    Abstract: The interplay of biology, environment, and lifestyle direct the development and progression of complex diseases and other health outcomes. Therefore, integration of longitudinal multimodal data is needed to understand the mechanisms underpinning major molecular transitions. Previously during my doctoral work at Stanford, I integrated multiomics data to elucidate the epigenetic mechanism of human surface ectoderm differentiation. I also built a pipeline to investigate the role of polymorphism, particularly non-coding genetic variants, in complex diseases. To address the common pain point of data silos limiting the interpretation of multimodal data integration, I formed a collaboration with Google Data Commons to build a free, open-source biomedical knowledge graph with a common schema and API. Currently it is composed of approximately 130 million nodes and 1.7 trillion triples (node-edge-node) from 22 publicly available biomedical datasets. Knowledge graphs are a key tool for hypothesis generation, data interpretation, and dimensionality reduction required for systems medicine research. Upon starting my postdoctoral work at the Institute for Systems Biology, I identified pregnancy as an excellent model system for prototyping precision medicine approaches. I used electronic healthcare records (EHR) from Providence St. Joseph Healthcare to investigate the impact of COVID-19 maternal infection and vaccination on maternal-fetal outcomes. In addition, I integrated multiomics placental data to investigate molecular network changes (interomics and intraomics) in common obstetric disorders. In a follow-up study (enrollment complete) we have longitudinal deep-phenotyping data of 435 people throughout pregnancy 80 of which have pregnancy complications. This includes multiomics, survey, EHR, and air quality data collected from first prenatal visit through delivery. My lab will use this data to define major molecular transition states throughout pregnancy. I will also investigate the disease mechanisms of common obstetric disorders including identifying for an individual the earliest possible point of deviation from a healthy trajectory. This interdisciplinary approach will identify potential drug targets, biomarker panels, and individualized clinical interventions.

  • Principled Approaches for Trustworthy Algorithms, Statistics, and Machine Learning | Gautam Kamath

    Special Seminar Series
    Halıcıoğlu Data Science Institute (HDSI), Room 123 3234 Matthews Ln, La Jolla, CA, United States

    Abstract: Despite impressive recent advances, machine learning models exhibit a number of critical deficiencies. They are prone to leaking sensitive information about their training data. They remain alarmingly brittle to attacks by malicious parties. Troublingly, these issues stem from more fundamental statistical vulnerabilities, which remain unresolved even decades later, highlighting significant gaps in our understanding of how to deal with these important considerations. As long as these problems remain, our models will not be appropriate for use beyond deployment in toy settings. In this talk, I will discuss recent advances on a number of these problems, which give key new algorithmic insights into how to address these considerations, and enable real-world deployments that were previously thought infeasible. In a first vignette, we will explore how to guarantee individual privacy in machine learning models, with a particular focus on large language models and the important role played by public data in the training pipeline. In a second vignette, we focus on how to robustly perform mean estimation, giving the first efficient and accurate algorithms for multivariate settings. We will go on to discuss connections to robustness against data poisoning attacks, robust exploratory data analysis, and surprising conceptual and technical connections with privacy.

  • On Data Ecology, Data Markets, the Value of Data, and Dataflow Governance | Raul Castro Fernandez

    Seminar Series

    Abstract:
    Data shapes our social, economic, cultural, and technological environments. Data is valuable, so people seek it, inducing data to flow. The resulting dataflows distribute data and thus value. For example, large Internet companies profit from accessing data from their users, and engineers of large language models seek large and diverse data sources to train powerful models. It is possible to judge the impact of data in an environment by analyzing how the dataflows in that environment impact the participating agents. My research hypothesizes that it is also possible to design (better) data environments by controlling what dataflows materialize; not only can we analyze environments but also synthesize them. In this talk, I present the research agenda on “data ecology,” which seeks to build the principles, theory, algorithms, and systems to design beneficial data environments. I will also present examples of data environments my group has designed, including data markets for machine learning, data-sharing, and data integration. I will conclude by discussing the impact of dataflows in data governance and how the ideas are interwoven with the concepts of trust, privacy, and the elusive notion of “data value.” As part of the technical discussion, I will complement the data market designs with the design of a data escrow system that permits controlling dataflows.

  • Enabling Performant and Trustworthy Learning-enabled CPS-IoT Systems | Mani Srivastava

    Special Seminar Series
    Halıcıoğlu Data Science Institute (HDSI), Room 123 3234 Matthews Ln, La Jolla, CA, United States

    Abstract: "The previously discrete technologies of IoT and AI have now entered a tight virtuous embrace. IoT allows sensing and actuation in our physical, social, and urban spaces with unimaginable ubiquity. AI allows sophisticated inferences and decisions to be made algorithmically using deep neural networks, even from unstructured and high- dimensional data, with uncanny performance. Together they seek to perform sophisticated perception-cognition-communication-action loops in diverse applications. However, designers of learning-enabled IoT systems face the challenge of extremely resource-constrained edge platforms operating in uncertain environments while assuring performance and trustworthiness. Moreover, in many applications, the systems go beyond taking actions based on rich inferences about the world state to perform long-term reasoning about complex events and obey the underlying physics, rules, and constraints. Based on our experience in designing such systems in applications including mHealth, ocean animal health, agriculture robotics, and military, This talk explores meeting these challenges through a combination of (i) neurosymbolic architectures that allow the incorporation of physics awareness and human knowledge while enhancing user trust, (ii) automatic platform-aware architecture search and code generation, and (iii) techniques to efficiently adapt to the deployment environment."