Vision, Language, Action: The Hard Parts of Building AI That Works
Details
This is an event repost, registration link and further details on each talk can be found under: https://voxel51.com/events/zurich-ai-ml-and-computer-vision-meetup-july-30-2026
Description
AI systems are rapidly moving beyond isolated models and simple prompts. Today’s most ambitious applications combine computer vision, large language models, multimodal generation, retrieval, orchestration, and autonomous agents into complex end-to-end workflows.
The promise is significant: systems that can analyze visual data, generate cinematic content, support high-stakes inspection tasks, and uncover connections that traditional search methods might miss.
But how well do these systems work in practice? And what does it take to move from an impressive prototype to a reliable, scalable AI application?
Modern computer vision still depends heavily on dataset quality, evaluation, and production infrastructure. AI video generation often remains a manual “prompt and pray” process. Visual inspection systems must perform reliably in critical domains such as healthcare and infrastructure. Meanwhile, LLM-powered retrieval systems can struggle when discovery requires indirect, unexpected, or non-obvious connections.
In this session, we’ll explore how researchers and engineering teams are addressing these challenges across vision, language, and multimodal AI. Through practical demonstrations and real-world case studies, the speakers will examine how open-source models, local infrastructure, foundation models, video APIs, and agentic architectures can be combined into more capable and dependable AI workflows.
💡 You’ll Learn:
- How modern computer vision systems are built, evaluated, and improved
- Why dataset curation and model analysis remain critical in the foundation-model era
- How agentic pipelines can transform static manuscripts into cinematic video trailers
- How foundation models are being applied to medical diagnostics and civil infrastructure monitoring
- Why traditional Agentic RAG can miss unexpected connections
- How AI systems can be designed for exploration, discovery, and deeper insight
- The practical challenges of maintaining consistency, reliability, and narrative integrity across multimodal pipelines
🧠 Who Should Attend?
- Engineers building computer vision, LLM, or multimodal AI applications
- Researchers working with foundation models, retrieval, and intelligent agents
- AI practitioners moving systems from prototypes into production
- Product and innovation leaders exploring new AI capabilities
- Teams working with RAG who want to go beyond conventional semantic search
- Anyone interested in how vision, language, and agentic systems are converging
The talks will combine technical insights, practical demos, and real-world lessons, with plenty of time for discussion and Q&A.
