Foundations of Optimization, Learning, and Data Science (FOLDS) Seminar

Logistics

About

What This seminar features leading experts in optimization, learning, and data science. Topics span algorithms, complexity, modeling, applications, and mathematical underpinnings.

Why Foundational advances in these fields are increasingly intertwined. This seminar serves as a university-wide hub to bring together the many communities across UPenn interested in optimization, learning, and data science: the Department of Statistics and Data Science, Electrical Engineering, Computer Science, Applied Mathematics, Economics, Wharton OID, and others. To help promote this internal interaction, several speakers will be from UPenn.

Funding We are grateful to the IDEAS Center for Innovation in Data Engineering and Science, Penn AI, and the Wharton Department of Statistics and Data Science.

Previous seminars

Fall 2026

Date Speaker
Sep 3 Shanyin Tong
Sep 10 Morgane Austern
Sep 17 Mateo Díaz
Sep 24 Pablo Gonzalez-Camara
Oct 1 Fall Break
Oct 8 Nicolas Loizou
Oct 15 Lachlan MacDonald
Oct 29 Meena Jagadeesan
Nov 5 Jonathan Niles-Weed
Nov 12 Cynthia Rush
Nov 19 Daniel Robinson

Spring 2027

Date Speaker
March 18 Carlos Fernandez-Granda
March 25 Eric Bradlow
April 1 Claire Donnat
April 8 Ying Jin
April 15 Enric Boix
April 22 Shipra Agarwal

Talk Abstracts

Shanyin Tong

Date: Sep 3, 2026

Title: Large Deviations for Rare-Event Estimation and Control

Abstract: Rare and extreme events, such as natural disasters, cascading failures, and accidents in autonomous systems, can have severe consequences but are inherently data-scarce: precisely because they occur infrequently, there are often too few observations to reliably learn their statistics directly from data. This creates a fundamental challenge for data-driven prediction and decision-making, particularly when the underlying systems involve high-dimensional uncertainty and expensive physical models. In this talk, I will present a computational framework based on large deviation theory (LDT) for rare-event estimation and control. LDT connects the probability of a rare event to a deterministic optimization problem over the uncertain parameters, whose solution identifies the most likely mechanism leading to the event. Building on this connection, I will discuss sampling-free approximations for rare-event probabilities, LDT-informed importance sampling algorithms, and optimization under rare chance constraints. These approaches exploit the geometry of the rare-event set and information from the associated optimization problem to substantially reduce the computational cost of both probability estimation and risk-aware control.

I will illustrate these methods through applications to physical and engineered systems, with a particular focus on traffic systems involving autonomous vehicles. These developments form part of RareDT, a broader effort to develop digital twins that go beyond predicting typical system behavior and instead incorporate the quantification and control of rare but consequential events.


Morgane Austern

Date: Sep 10, 2026

Title: Dependence and degeneracy create: Multiple descent in overparameterized models

Abstract: Over-parameterized linear regression has been widely studied over the last decade. However, most existing works assume that the covariates are independent and that their covariance matrices are non-degenerate. In this talk, we will relax both assumptions and derive deterministic equivalents for the prediction risk in a vanishing-ridge regime. We show that degeneracy of the covariance matrices and dependence can lead to multiple descent, and characterize where the corresponding peaks can occur. Our proofs use a novel graph representation of the variance profile. We show that maximum matchings and the Dulmage–Mendelsohn decomposition of the associated bipartite graph identify the configurations at which the variance becomes singular.


Mateo Díaz

Date: Sep 17, 2026

Title: Leveraging Structure for Faster Algorithms in Optimization and Diffusion

Abstract: Large-scale iterative methods drive modern AI, yet their theoretical foundations often lag behind their empirical success. We argue that bridging this gap requires identifying the inherent problem structure that enables these algorithms to perform well. This talk instantiates this principle across two domains: optimization and generative modeling.

First, we derive new theoretical guarantees for the Levenberg–Morrison-Marquardt method. Although this method is ubiquitous in settings that demand highly accurate solutions—for instance, when training physics-informed neural networks for scientific discovery—classical guarantees do not explain its strong empirical performance in modern overparameterized, ill-conditioned regimes. By reframing it through the lens of composite optimization, we uncover geometric conditions that ensure fast convergence even in these challenging modern regimes.

Second, we introduce Proximal Diffusion Models (PDM). While standard diffusion models rely on score-matching and forward discretization, we demonstrate that a backward discretization using proximal maps offers significant theoretical and practical advantages. Under mild conditions, we prove that PDM achieves ε-accuracy in KL-divergence within $\widetilde{O}(d/\sqrt{\varepsilon})$ steps and empirically demonstrate that it outperforms conventional methods using fewer sampling iterations.

Bio: Mateo Díaz is an Assistant Professor in the Department of Statistics and the Data Science Institute at the University of Chicago. His research lies at the intersection of continuous optimization, geometry, and statistics. He develops mathematical foundations and scalable algorithms for problems arising in data science, machine learning, and signal processing. Before joining the University of Chicago, he was an Assistant Professor of Applied Mathematics and Statistics at Johns Hopkins University. He previously spent two years as a postdoctoral scholar at Caltech. Mateo received his Ph.D. in Applied Mathematics from Cornell University in 2021 and earned bachelor’s degrees in Mathematics and in Systems and Computing Engineering, as well as a master’s degree in Mathematics, from Universidad de los Andes in Colombia. His work has been recognized with the Beale–Orchard-Hays Prize, an NSF CAREER Award, and a Sloan Research Fellowship in Mathematics.


Pablo Gonzalez-Camara

Date: Sep 24, 2026

Title: Metric Geometry and the Study of Biological Shape Across Scales

Abstract: The systematic study of biological shape has been central to major advances in biology, from the development of the neuron doctrine in the 19th century to the discovery of the molecular basis of sickle cell disease and the structures of DNA and hemoglobin. Across scales, biological shape encodes fundamental information about biological function and organization. Over the past two decades, advances in microscopy and structural biology have enabled increasingly rich and precise measurements of biological structures at the cellular, subcellular, and molecular scales. These developments create a need for rigorous quantitative methods for representing, comparing, and analyzing biological shapes, especially when the objects under study are highly heterogeneous. In this talk, I will discuss recent applications of ideas from metric geometry, and particularly Gromov-Wasserstein couplings, to these problems. We will focus on three settings: comparing cellular morphologies across heterogeneous and morphologically complex populations, such as neurons and glia; analyzing subcellular organization in ways that are robust to variation in cellular morphology; and comparing three-dimensional protein structures. We will also discuss how these methods can be naturally combined with metric-learning neural networks to achieve the scalability of deep learning while retaining interpretability and generalizability. Together, these results illustrate how metric geometry can provide a rigorous and flexible quantitative framework for studying biological shape across scales.

Bio Pablo G. Cámara is an Associate Professor of Genetics at the University of Pennsylvania and a faculty member of the Institute for Biomedical Informatics and the AI2D Center for AI and Data Science for Integrated Diagnostics. His research focuses on the cellular and molecular organization of the brain and brain tumors, utilizing mathematical principles to analyze high-dimensional omics and imaging data. He received a Ph.D. in Theoretical Physics from Universidad Autónoma de Madrid and continued his postdoctoral work at École Polytechnique (France), the European Organization for Nuclear Research (CERN, Switzerland), and the University of Barcelona. Fascinated by open fundamental questions in biomedicine, he shifted his focus to quantitative biology in 2014 as a postdoctoral fellow at Columbia University and the Institute for Advanced Study in Princeton. He joined the University of Pennsylvania as a faculty member in 2018, where he is developing geometry- and topology-based algorithms for integrating and analyzing single-cell omics, cytometry, and imaging data, and using them to elucidate the cellular ecosystem and oncogenic pathways of glioma, and to characterize the functional properties of CAR-T cell immunotherapies.


Nicolas Loizou

Date: Oct 8, 2026

Title:

Abstract:


Lachlan MacDonald

Date: Oct 15, 2026

Title:

Abstract:


Meena Jagadeesan

Date: Oct 29, 2026

Title:

Abstract:


Jonathan Niles-Weed

Date: Nov 5, 2026

Title:

Abstract:


Cynthia Rush

Date: Nov 12, 2026

Title:

Abstract:


Daniel Robinson

Date: Nov 19, 2026

Title:

Abstract:


Carlos Fernandez-Granda

Date: March 18, 2027

Title:

Abstract:


Eric Bradlow

Date: March 25, 2027

Title:

Abstract:


Claire Donnat

Date: April 1, 2027

Title:

Abstract:


Ying Jin

Date: April 8, 2027

Title:

Abstract:


Enric Boix

Date: April 15, 2027

Title:

Abstract:


Shipra Agarwal

Date: April 22, 2027

Title:

Abstract: