Skip to content

About us

🖖 This group is for AI researchers, machine learning engineers, roboticists and open source enthusiasts.

Every week we bring you a diverse set of speakers working at the cutting edge of AI, machine learning, robotics and computer vision.

This Meetup is sponsored by Voxel51, the multimodal data platform for physical AI. Learn More.

Interested in speaking at a future event? Submit a talk!

By becoming a member of this group you agree to Voxel51's Terms of Service and Privacy Statement and agree to receive occasional emails about upcoming events.

Upcoming events

4

See all
  • Network event
    Sept 17 - ADAS, AV, and AI Meetup

    Sept 17 - ADAS, AV, and AI Meetup

    ·
    Online
    Online
    203 attendees from 55 groups

    Join our virtual meetup to hear talks from experts on cutting-edge topics across AI, ML, and computer vision.

    Time, Date and Location

    Sep 17, 2026
    9:00 AM - 11:00 AM PST
    Online.
    Register for the Zoom!

    AI for Autonomous Driving: From Data to Decisions

    Building reliable automated driving systems is as much a data and engineering challenge as a modeling one. In this talk, Tin will share perspectives from his work at Porsche AG on applying modern AI methods across the autonomous driving development process, from making sense of large-scale driving data to understanding and evaluating how AI-based systems behave on the road. He'll discuss lessons learned from real-world development, where today's approaches shine, and where hard problems remain for the ADAS and AV community.

    About the Speaker

    Tin Stribor Sohn is a PhD Student at Porsche AG and Karlsruhe Institute of Technology in the area of Foundation Models for Scenario Understanding and Decision Making in Autonomous Robotics, Tech Lead at Data Driven Engineering for Autonomous Driving, Prior: Master in CS at University of Tuebingen with focus on Computer Vision and Deep Learning and co-founder of a software company for smart EV charging

    Advancing ADAS and Autonomous Vehicle Development with Multimodal Data

    ADAS and autonomous vehicle systems rely on increasingly complex data from cameras, video, LiDAR, radar, and other sensor streams. In this session, Murilo will introduce Voxel51 and explore how the latest multimodal capabilities in FiftyOne help teams bring these data sources together to better understand their datasets and model behavior. He’ll discuss how unified workflows for visualization, search, curation, and evaluation can help ADAS and AV teams uncover challenging scenarios, investigate model failures, and build safer, more reliable autonomous systems.

    About the Speaker

    Murilo Gustineli is a Machine Learning Engineer at Voxel51 working at the intersection of representation learning and computer vision. He holds an M.S. in Computer Science from Georgia Tech, where he co-leads the DS@GT Applied Research & Competitions group, advancing machine learning research through competitive challenges and peer-reviewed publications.

    From Survey-Grade Maps to Physical AI: Scaling Real-World Data for Training and Simulation

    Physical AI systems are increasingly constrained not by model architectures, but by the availability of scalable, high-fidelity real-world data. This talk explores how Dynamic Map Platform transforms survey-grade road assets collected across 1.8 million km of roads worldwide into training- and simulation-ready datasets, including point clouds, imagery, HD maps, road surface models, and 3D Gaussian Splatting representations.

    We will discuss why geometric accuracy, semantic understanding, and real-world diversity are critical to building robust autonomous driving systems. Attendees will learn how real-world geospatial data can be structured and scaled for AI training and simulation workflows.

    About the Speaker

    Ryoto Miyake is a Software Engineer at Dynamic Map Platform, where he works on transforming large-scale geospatial data into AI-ready datasets for training, simulation, and validation, such as HD maps and 3D Gaussian Splatting. With a background in transportation engineering, he works closely with automotive manufacturers and industry partners to bridge large-scale real-world mapping data with next-generation AI and mobility systems.

    • Photo of the user
    • Photo of the user
    2 attendees from this group
  • Network event
    Oct 15 - AI, ML, and Computer Vision Meetup

    Oct 15 - AI, ML, and Computer Vision Meetup

    ·
    Online
    Online
    55 attendees from 55 groups

    Join our virtual meetup to hear talks from experts on cutting-edge topics across AI, ML, and computer vision.

    Time, Place and Location

    Oct 15, 2026
    9:00 AM - 11:00 AM PST
    Online.
    Register for the Zoom!

    Testing AI Systems in Production: Data Quality, Drift, and Model Evaluation

    AI systems can pass offline evaluation and still fail in production when real-world data changes, features become stale, labels or feedback signals are incomplete, or model behavior drifts away from expected outcomes. This talk shares practical patterns for testing and evaluating AI systems after deployment, including data quality checks, drift detection, online/offline metric comparison, model monitoring, and rollback analysis.

    Using personalization and recommendation systems as examples, we will examine how teams can build evaluation workflows that catch quality issues before users do. Attendees will leave with a practical checklist for making AI-backed systems easier to evaluate, debug, and operate as data changes over time.

    About the Speaker

    Jayakumar Ramalingam is a Staff Software Engineer and Cloud Architect at SiriusXM with over 16 years of experience building cloud-native platforms, real-time data pipelines, resilient APIs, and AI/ML-enabled applications at production scale.

    Where Should Your Model Live? A Framework for Tiering Computer Vision Deployments

    Where should a computer vision model actually run - on-device, near the edge, or in the cloud? It's a decision that looks simple until requirements like latency, cost, connectivity, and update cadence start pulling in different directions, often revealing themselves only after deployment.

    Drawing on hands-on experience developing and deploying CV models across Hailo, Nvidia, Qualcomm and AWS platforms, this talk introduces a practical framework for tiering computer vision deployments based on real project requirements and constraints. Discussion will include what changes at each tier - from development to deployment to monitoring and update strategy - with relevant industry examples.

    Attendees will leave with a set of questions or a framework they can use to place their own CV projects into the right tier.

    About the Speaker

    Ajaykumaar Sivacoumare is an AI Software Engineer specializing in computer vision and edge AI, with production experience developing and deploying CV models across Nvidia, Hailo, Qualcomm and AWS-based platforms.

    From 2D Slices to 3D Tumors: Lightweight Volumetric Detection Without Heavy 3D Networks

    Slice-wise 2D detectors are fast and scalable, but they struggle to produce reliable 3D bounding boxes from volumetric medical data. This talk presents YOLO-PVC, a lightweight post-processing framework that consolidates slice-wise YOLO detections into coherent 3D bounding boxes using percentile-based geometric aggregation and a minimal MLP calibration module.

    This talk demonstrates consistent improvements in volumetric IoU across three liver tumor categories i.e., HCC, CCA, and Mixed, without requiring dense 3D annotations or memory-intensive architectures. The talk covers the clinical motivation, the technical approach, and practical lessons from deploying computer vision on real hospital MRI data.

    About the Speaker

    Talha Waqas is a second-year PhD student at ESME Research Lab, Paris and LISSI, Université Paris-Est, working on computer vision applied to medical imaging, with a focus on tumor classification, detection, and segmentation in multi-phase liver MRI.

    The Two-Loop Architecture for Voice AI

    Building responsive voice AI requires balancing latency with intelligence. This talk introduces a practical architecture that separates real-time conversation from asynchronous reasoning, enabling richer interactions without slowing the user experience. The session covers reusable design patterns drawn from production-inspired conversational AI systems.

    About the Speaker

    Abhinav Tushar is an ML engineer and researcher specializing in Conversational AI and Speech Technology.

  • Network event
    Oct 28 - SF Physical AI Workshop and Meetup

    Oct 28 - SF Physical AI Workshop and Meetup

    Antler, 144 Townsend St, San Francisco, CA 94107, USA, San Francisco, CA, US
    10 attendees from 11 groups

    Join us at Antler VC on Oct 28 for the SF Physical AI Workshop and Meetup, co-presented by Nebius and Voxel51.

    Seats are limited, Pre-registration is mandatory.

    Date, Time and Location

    Oct 28, 2026
    5:30 PM - 8:30 PM
    Antler San Francisco
    144 Townsend St (3rd Floor)
    San Francisco, CA 94107

    Meetup Agenda

    5:30–6:30 PM

    • Networking, food, drinks, and lightning talks

    Workshop Agenda
    This hands-on session uses DROID, a real-world robotics dataset loaded into FiftyOne as a native multimodal MCAP recording, and YOLO11n, fine-tuned live during the session.

    6:30–7:30 PM

    • Welcome + framing: from raw robot logs to a trained detector
    • Explore a real DROID robotics recording in FiftyOne's native multimodal MCAP viewer: camera, proprioception, and language on one synced timeline, no ROS install required
    • Curate: extract and browse frames from the recording, filter and deduplicate
    • Compute embeddings on the curated frames; explore via similarity search and embeddings visualization (via Nebius Serverless AI Jobs)

    7:30–8:00 PM

    • Auto-label: open-vocabulary detection to generate bounding boxes for the robot gripper and target objects
    • Train: fine-tune a YOLO11n detector on the auto-labeled frames (via Nebius Serverless AI Jobs)
    • Evaluate results and close the loop: view predictions back on the original MCAP timeline

    8:00–8:30 PM

    • What else Nebius offers: Token Factory walkthrough — chat/vision models, fine-tuning, credits
    • Photo of the user
    • Photo of the user
    3 attendees from this group
  • Network event
    Nov 11 - AI, ML and Computer Vision Meetup

    Nov 11 - AI, ML and Computer Vision Meetup

    ·
    Online
    Online
    49 attendees from 55 groups

    Join our virtual meetup to hear talks from experts on cutting-edge topics across AI, ML, and computer vision.

    Time, Date and Location

    Nov 11, 2026
    9:00 AM - 11:00 AM PST
    Online.
    Register for the Zoom!

    Agentic RAG: Beyond Retrieve-and-Generate

    Retrieval-Augmented Generation (RAG) has become the default pattern for grounding large language models in external knowledge, but most implementations still follow a rigid retrieve-once, generate-once pipeline — one that struggles with multi-hop questions, ambiguous queries, and knowing when its own retrieved context is insufficient. This talk introduces Agentic RAG, where retrieval is treated as an action within an agent's reasoning loop rather than a fixed upstream step.

    We'll examine query decomposition and routing for breaking complex questions into targeted sub-queries, self-reflective and corrective retrieval loops that let an agent judge and re-query its own results, and tool-orchestration patterns (via MCP) that let retrieval sit alongside other agent actions like database lookups and API calls. Using a live architecture — evolving a standard RAG chatbot into an agentic, MCP-connected system — we'll walk through what changes in design, and where these systems introduce new failure modes: grounding drift, latency and cost from repeated retrieval loops, and cases where a simpler RAG pipeline still wins.

    Attendees will leave with a practical framework for deciding when the added complexity of agentic RAG is worth it, and a set of design patterns for building it correctly.

    About the Speaker

    Balaji Venkatasubramaniyar is a Technical Lead at Wisdom Infotech, leading a 15+ person engineering team delivering enterprise solutions. With 13+ years of experience in enterprise software and insurance technology, he specializes in agentic AI systems, RAG architectures, and vector databases.

    GeoAI for the Physical World: Earth Observation, Foundation Models, and Urban Digital Twins

    Earth observation provides a unique form of computer vision for understanding the physical world at city to continental scales. In this talk, I will show how satellite imagery, geospatial data, machine learning, and foundation-model representations can be combined to characterize urban environments and environmental conditions.

    I will present UrbanScope Open, an open GeoAI digital-twin prototype integrating Earth observation with 3D buildings, vegetation, land-surface temperature, air quality, noise, population, and other urban data. I will also share lessons from my research using geospatial foundation-model embeddings for environmental prediction across Europe.

    The talk will discuss how these approaches can contribute to increasingly multimodal AI systems capable of reasoning about real-world environments.

    About the Speaker

    Cesar Alvarez is a researcher at the University of Augsburg working at the intersection of GeoAI, Earth observation, remote sensing, and environmental intelligence. His research applies machine learning, computer vision, and geospatial foundation models to problems including urban environments, climate risk, air quality, and agriculture.

    Can agents get curious?

    Most AI agents are good at answering a question once we tell them exactly what to look for. The harder problem is building agents that can explore a complex dataset autonomously: generating hypotheses, deciding which analyses are worth running, allocating additional compute when evidence is ambiguous, and knowing when they have enough evidence to stop.

    In this talk, I’ll show an architecture for autonomous research agents that combines structured knowledge, iterative tool use, and explicit evidence tracking to turn open-ended questions into a sequence of testable investigations. I’ll discuss practical lessons from building and evaluating these systems, including why more test-time compute does not automatically produce better research and how provenance and evaluation can make long-running agents more reliable.

    I’ll close with a live example of an agent exploring a dataset, revising its hypotheses, and choosing what to investigate next.

    About the Speaker

    Srivatsa P is a member of technical staff at Sigma Computing, where he works on AI agents that reason over complex enterprise data. Previously, he worked on machine learning at Apple and conducted research at the MIT Media Lab; outside of traditional ML, he has also worked on mapping coral reefs through underwater imaging, which sparked an enduring interest in how intelligent systems make sense of messy real-world data.

Group links

Organizers

Super Organizer