Vision's LLM Moment: an evening with Robert Geirhos (Google DeepMind)
Details
Robert Geirhos is coming to Heidelberg to talk about one of the more interesting questions in AI right now: are video models quietly turning into general-purpose vision models, the same way LLMs became general-purpose language models?
Think about how much language models changed things. A few years back you needed a separate model for every job. One for translation, a different one for summarizing, another for answering questions. Then LLMs showed up and you could suddenly do all of it by just asking. And the recipe behind that shift was pretty plain once you saw it: take a big generative model, train it on a huge pile of web data.
Robert's argument is that the same thing might be starting to happen in vision, and video is where to look.
His team has been poking at Veo 3, and it keeps doing things nobody trained it to do. It'll pick objects out of a scene, find edges, edit images, reason about how physical stuff behaves, work out what you can actually do with an object, even solve mazes and symmetry puzzles. None of that was the training goal. It just emerged, the same way a lot of LLM abilities did.
If that pattern holds up, it's a big deal. It hints at a future where one vision model handles most perception tasks, instead of a new specialized model every single time.
Worth coming whether you build with this stuff, research it, or you're just curious where it's headed. You'll be hearing the case from one of the people actually making it.
About Robert
Robert is a Staff Research Scientist at Google DeepMind in Zürich, working on reaching visual intelligence through video models like Veo 3. He's one of the senior authors on the paper this talk is based on, "Video models are zero-shot learners and reasoners." If you've been around computer vision for a while, you probably already know some of his earlier work: the 2019 paper showing that image classifiers lean on texture far more than shape, and "Shortcut learning in deep neural networks" from 2020. Both get cited constantly. Along the way he's picked up the ELLIS PhD Award, a NeurIPS Outstanding Paper Award, and multiple orals across ICLR, NeurIPS, ICML, and VSS.
How the evening runs
Robert talks for about 60 minutes, then we open it up for questions. After that, stick around to meet people and chat over drinks and snacks.
