Machine Learning and AI Seminar Series

About

This seminar series invites experts from across the country to come to Columbia and present the latest cutting-edge research in the field of Machine Learning and Artificial Intelligence. Running the gamut between theory and empirics, the seminar provides a single, unified space to bring together the ML/AI community at Columbia. Topics of interest include, but are not limited to, Language Models, Optimization for Deep Learning, Reinforcement and Imitation Learning,  Learning Theory, Interpretability and AI Alignment, AI for science, Probabilistic ML, and Bayesian methods.

Hosts & Co-Sponsors: DSI Foundations of Data Science Center, Department of Statistics, Arts and Sciences, and Columbia Engineering


Fall 2026 Seminars

This event series will be held on Fridays from 11 a.m.- noon in the Columbia University Department of Statistics (Room 903), located in the School of Social Work Building. See speakers, dates, and registration links below.

The Columbia Morningside campus is open to the Columbia community. If you do not have an active CUID, registration is required. The deadline to register is at noon the day before the event. External guests will receive a QR code by email to enter campus.

Headshot of speaker: Yoav Artzi
Friday, Sept. 18

Yoav Artzi, Associate Professor in the Department of Computer Science, Cornell Tech at Cornell University

 

Headshot of speaker: Ben Eysenbach
Friday, Oct. 2

Benjamin Eysenbach, Assistant Professor of Computer Science at Princeton University

REGISTER

Headshot of speaker: Aviral Kumar
Friday, Oct. 16

Aviral Kumar, Assistant Professor of Computer Science and Machine Learning at CMU

REGISTER

Headshot of speaker: Stephen Tu
Friday, Oct. 30

Stephen Tu, Assistant Professor in the Department of Electrical and Computer Engineering at the University of Southern California

REGISTER

Headshot of speaker: Qi Lei
Friday, Nov. 13

Qi Lei, Assistant Professor of Mathematics, Data Science, and Computer Science at NYU

REGISTER 


Speaker Abstracts: 2026 - 2027 Academic Year

Event Date: September 18, 2026

Sparks of New Pre-Training

This talk covers new pre-training techniques. First, we introduce the hypothesis that state and prediction representations, which are entangled in transformers, are better separated. We design a simple architectural modification that effectively separates them, and provides 2.6x token efficiency during pre-training. The second technique is focused on externalizing knowledge by pre-training an LLM to rely on an external knowledge base, while inducing this KB from the pre-training data. We re-model the Limited Memory Language Model (LMLM) paradigm we introduced in prior work, with a new expressive continuous query mechanism. This dramatically increases the expressivity of the LMLM paradigm, and allows scaling to general web text. Our Co-LMLM model gives lower perplexity than a model trained on 40x the amount of tokens, and provides SimpleQA performance on par with Claude Sonnet 4.5. All that while presenting all the advantages of the LMLM class -- knowledge control, provenance, editing, and factuality. Together, these techniques demonstrate two of many possible avenues to bring about fundamental change in LLMs through new pre-training paradigms.

Event Date: October 2, 2026

Seeing like an Empowered Agent

Empowerment measures an agent's ability to actively control its environment. Appealing as an information theoretic quantity, empowerment is often juxtaposed with more geometric approaches to exploration, such as those that build a 3D model of the world. In this talk, I'll discuss recent results linking empowerment and geometry, which provides an answer to a longstanding open question on the connections between empowerment and centrality. These results suggest new approaches to representation learning, exploration, and information gathering, which we demonstrate on robotics benchmarks and open-ended games.

Event Date: October 16, 2026 

Event Date: October 30, 2026 

Multistep Backwards Latent Losses Learn Unstable Linear Dynamics

Model-based control synthesis relies on an accurate description of the system dynamics and the environment. When such models are unavailable or difficult to derive from first principles, they must be learned from system trajectories. Latent world models have recently emerged as a promising general-purpose approach, learning low-dimensional representations together with dynamics that evolve these representations to predict future behavior. However, it remains unclear which dynamical properties these models preserve, and whether they support control synthesis with guarantees of closed-loop stability and safety. Resolving these questions is essential to establishing latent world modeling as a reliable tool for control design. In this talk, we initiate a study of the dynamical properties captured by latent world model designs. For trajectory data generated by a linear dynamical system, we show that multi-step backward latent reconstruction losses recover the system’s unstable subspace, provided the latent dimension is large enough to represent the unstable modes. We then show that an H-infinity controller synthesized using the learned latent dynamics stabilizes the original system. Crucially, both multi-step prediction and reconstruction in the original space (i.e., backward reconstruction) are necessary for this guarantee: removing either can lead to an unstable closed-loop system. Although our analysis focuses on linear systems, experiments on high-dimensional nonlinear systems reveal behavior consistent with our theoretical findings.

Event Date: November 13, 2026