Skip to content

Details

Join our virtual meetup to hear talks from researchers at NYU on cutting-edge topics across AI, ML, and computer vision.

Date, Time and Location

Aug 25, 2026
9:00 AM - 11:00 AM PST
Online. Register for the Zoom!

Using Computer Vision to Advance the Sciences

I'll present some of our ongoing work on using computer vision to create impact in the sciences. These target a two areas, solar physics and evolutionary biology, that deal with objects of radically different sizes but are unified by a need for high quality, trustworthy data.

I'll show off our efforts, done in collaboration with domain experts, that aim to produce the best possible maps of the Sun's powerful magnetic field and have created some of the world's largest repositories of data about bird morphology.

About the Speaker

David Fouhey is an Associate Professor at New York University and a research scientist at Polymathic AI. Before joining NYU, he received a PhD in robotics from Carnegie Mellon, was a postdoc at UC Berkeley, and was a professor at University of Michigan.

Dream to Assembly: Rebuilding the Broken by Imagining the Whole

Give an intelligent system a pile of broken pieces. Can it recover the object they once formed? And what if some pieces are missing? In this talk, I will present our recent work, from GARF to CRAG, on learning to reconstruct fragmented objects. GARF focuses on generalizing reassembly from synthetic training data to complex real-world fractures through large-scale fracture-aware pretraining and SE(3) flow matching. CRAG goes a step further by asking whether a model can imagine the whole in order to better assemble the parts. By jointly reasoning about fragment poses and complete object geometry, it uses global shape context to resolve ambiguous arrangements and recover missing geometry. Together, these works move 3D reassembly from matching visible pieces toward reasoning about the complete object behind incomplete observations.

About the Speaker

Jing Zhang is an incoming Faculty Fellow at New York University and currently a postdoctoral researcher working at the intersection of embodied AI, 3D computer vision, robotics, and AI for science. Her research explores how intelligent systems can learn from real-world observations, reason about 3D space, imagine what is not directly observable, and use these capabilities to navigate, interact with, and reconstruct the physical world. Her work spans open-world robot navigation, spatial reasoning, generative world models, 3D reconstruction and reassembly, and computational methods for archaeological and paleoanthropological materials. She was selected as a 2025 Rising Star in EECS.

Solaris: Building a Multiplayer Video World Model in Minecraft

This talk will introduce Solaris: a multiplayer video world model in Minecraft. I will first present SolarisEngine, the software platform we built to simulate realistic multiplayer gameplay between bots at scale, enabling us to collect a large training dataset of aligned multiplayer actions and frames.

I will then discuss our staged training pipeline, starting with single-player pre-training before converting the model into a long-horizon multiplayer generator through bidirectional training, followed by causal training, and concluding with Self Forcing. I will also cover our memory-efficient implementation of Self Forcing, called Checkpointed Self Forcing.

Finally, I will showcase generated videos illustrating how Solaris maintains coherent long-horizon multiplayer interactions.

About the Speaker

Oscar Michel is a PhD student at NYU advised by Prof. Saining Xie. His research studies world models: generative models of agents interacting in an environment.

Closing the human to robot gap for dexterous hands

Collecting task-specific robot data for multi-fingered hands is challenging due to the many difficulties that arise in teleoperation. That is why recently there has been a major focus on learning robot policies directly from human demonstrations. However, human demonstrations are difficult to work with; there is a major morphological and visual gap between human and robot hands, as well as between the environments they operate in.

In this talk, I'd like to discuss my efforts on closing this gap.

About the Speaker

Irmak Guzey I'm Irmak (she/her), a rising 3rd year PhD student at New York University, currently advised by Lerrel Pinto. My research focuses on robot learning for dexterous manipulation. I have been awarded a Fulbright scholarship and NYU's Best Master's Thesis Award in the past.

Related topics

Artificial Intelligence
Computer Vision
Machine Intelligence
Machine Learning
Data Science

You may also like