Filters

Changing any of the form inputs will cause the list of events to refresh with the filtered results.

  • HDSI Alumni Celebration

    Ridgewalk Social University of California San Diego, Rimac Annex, San Diego

    You're invited to join us in celebrating and socializing with HDSI alumni this summer when we gather for our annual celebration! Alumni are always eager to see their former professors, so if you happen to be in town, we would love to have you!

  • Causal Inference symposium

    Causality is increasingly a part of AI, data science, robotics, and more, but it is not always clear how we can learn causality from data. This symposium will be featuring leading HDSI Faculty who will be providing an introductory overview on these methods, followed by domain-specific talks and open discussion.

  • Some new results for streaming principal component analysis

    Special Seminar Series

    Abstract: While streaming PCA (also known as Oja’s algorithm) was proposed about four decades ago and has roots going back to 1949, theoretical resolution in terms of obtaining optimal convergence rates has been obtained only in the last decade. However, we are not aware of any available distributional guarantees, which can help provide confidence intervals on the quality of the solution. In this talk, I will present the problem of quantifying uncertainty for the estimation error of the leading eigenvector using Oja's algorithm for streaming PCA, where the data are generated IID from some unknown distribution. Combining classical tools from the U-statistics literature with recent results on high-dimensional central limit theorems for quadratic forms of random vectors and concentration of matrix products, we establish a distributional approximation result for the error between the population eigenvector and the output of Oja's algorithm. We also propose an online multiplier bootstrap algorithm and establish conditions under which the bootstrap distribution is close to the corresponding sampling distribution with high probability. While there are optimal rates for the streaming PCA problem, they typically apply to the IID setting, whereas in many applications like distributed optimization, the data is generated from a Markov chain and the goal is to infer parameters of the limiting stationary distribution. If time permits, I will also present our near-optimal finite sample guarantees which remove the logarithmic dependence on the sample size in previous work, where Markovian data is downsampled to get a nearly independent data stream.

  • Fireside Chat: Theory in the age of modern AI

    TILOS Fireside Chat on "Theory in the age of modern AI", which will be a conversation led by TILOS team members: Misha Belkin (UCSD), Arya Mazumdar (moderator, UCSD), Tara Javidi (UCSD), Visheeth […]

  • EGEMEN KOLEMEN | SEMINAR ON FUSION ENERGY AND AI/ML JOINT SEMINAR: CSE, HDSI, MAE, SDSC

    Recent advances in computing hardware (FPGAs, distributed parallel computing) and numerical methods (machine learning algorithms, automatic differentiation) create new possibilities for Fusion Power Plant optimization and control. In this talk, I will discuss some of the recent accomplishments of the Plasma Control Group at Princeton that take advantage of these new capabilities.

  • Steampunk Data Science

    Distinguished Lecturer Series

    How did scientists make sense of data before statistics and computing? This talk will explore this question by focusing on the discovery of vitamins, which occurred in the early 20th century just before the advent of modern statistical methodology. I will describe the varied practices in experimentation and reporting and highlight the sorts of insights required to uncover what "works." Through this discussion, I will draw connections to contemporary data science tools to illustrate their pros and cons in facilitating discovery.

  • In Silico: Simulators, Emulators and the Human Brain Project

    The European Human Brain Project, a flagship project of the European Union, recently ended after 10 years of research and development. The goals of the HBP were to (1) explore the complexity of the human brain in space and time; (2) to transfer the knowledge broadly; (3) to provide research infrastructure for neuro-science; and (4) to create a community of researchers. One of the major challenges is to model neural activity, from micro- to macro-scale, in a way that enables simulation of the human brain. This leads to so-called in silico experiments, which will be used “to validate models, and to perform investigations that are not possible in the laboratory”. I will present examples of such experiments and discuss how they relate to, and can benefit from, statistical research on the design and analysis of computer experiments. My students and colleagues and I have been working on the potential advantages of replacing a slow/expensive simulator with a much faster and cheaper statistical emulator. Emulators are empirical replicas trained on data generated with the simulator. We have often used Gaussian process regression for this purpose, but in some applications other methods (random forests, polynomial regression) proved more effective. Emulators can be especially useful when the simulator runs are matched to data in the context of statistical inference. I will discuss the modeling options and present examples, including simulation of neural basket cells, calcium induced neural reactions, and stochastic simulators like the Hodgkin-Huxley model.