Skip to content

About us

🖖 This virtual group is for data scientists, machine learning engineers, and open source enthusiasts.

Every month we’ll bring you diverse speakers working at the cutting edge of AI, machine learning, and computer vision.

  • Are you interested in speaking at a future Meetup?
  • Is your company interested in sponsoring a Meetup?

Send me a DM on Linkedin

This Meetup is sponsored by Voxel51, the lead maintainers of the open source FiftyOne computer vision toolset. To learn more, visit the FiftyOne project page on GitHub.

Upcoming events

12

See all
  • July 27 - London AI, ML, and Computer Vision Meetup

    July 27 - London AI, ML, and Computer Vision Meetup

    Skempton Building, Imperial College, SW7 2AZ, London, GB

    Join our in-person meetup to hear talks from experts on cutting-edge topics across AI, ML, and computer vision.

    Register to reserve your seat!

    Date, Time and Location

    Jul 27, 2026
    5:30 PM - 8:30 PM BST
    Imperial College London, Skempton Building (LT201), South Kensington, London SW7 2AZ

    UniLight: Unified Multi-Modal Lighting Representation

    Lighting has a strong influence on visual appearance, yet understanding and representing lighting in images remains notoriously difficult. UniLight introduces a joint latent space to unify previously incompatible lighting representation - environment maps, images, irradiance and text descriptions.

    Modality-specific encoders are trained contrastively to align their representations, with an auxiliary spherical-harmonics prediction task reinforcing directional understanding. Our joint lighting embedding enables applications such as retrieval, example-based light control during image generation, and environment map generation from various modalities.

    About the Speaker

    Zitian Zhang - is a PhD candidate in Computer Science at Université Laval, and a research scientist intern in Adobe Research London. He focuses on image understanding, generation, and lighting representations through foundation models.

    LoST: Level of Semantics Tokenization for 3D Shapes

    Tokenization is fundamental to generative modeling and especially important for autoregressive 3D generation. However, current 3D shape tokenizers rely on geometric level-of-detail hierarchies that are token-inefficient and poorly aligned with semantic structure.

    We propose Level-of-Semantics Tokenization (LoST), which orders tokens by semantic salience so early tokens produce complete, plausible shapes and later tokens refine detailed geometry and semantics. LoST is trained with Relational Inter-Distance Alignment (RIDA), a semantic alignment loss that matches relationships in 3D shape latent space to those in DINO feature space.

    Experiments show that LoST achieves state-of-the-art reconstruction and efficient high-quality AR 3D generation while using only 0.1%–10% of the tokens required by prior methods.

    About the Speaker

    Niladri Dutt - is an ELLIS PhD student at University College London (UCL), sponsored by Adobe Research. He is advised by Prof Niloy Mitra (UCL) and Duygu Ceylan (Adobe).

    Material selection in 2D and beyond - methods, tricks and applications

    In this talk, we'll explore reasoning about images from a material-centric perspective, namely through the lens of material understanding. Materials distinguish themselves by their response to light, which is governed and modelled through physical properties like roughness or gloss - however, understanding such properties is a non-trivial task for current algorithms and models.

    We'll see how we can select materials similar to a given query material, significantly improve selection fidelity and eventually even venture beyond 2D, to enable selection in the 3D domain.

    About the Speaker

    Michael Fischer - is a research scientist at Adobe research London. He obtained his PhD from University College London (UCL), advised by Niloy Mitra and Tobias Ritschel. Michael has authored several top-tier publications (CVPR, ICCV, SIGGRAPH, ...) and is a recipient of both the Meta PhD scholarship and the Rabin Ezra scholarship as well as the Eurographics PhD Thesis award 2026.

    Lessons from the Trenches of Agentic Engineering

    A candid lessons-learned from running an agentic engineering consultancy with clients ranging from federal governments to early-stage AI startups. I'll cover what's held up under real production pressure, what I tried and abandoned, and the approaches that are quietly dead but still being sold. Expect specifics, opinions, and a few uncomfortable conclusions.

    About the Speaker

    John Adeojo - runs Brainqub3 an agentic engineering consultancy serving clients from federal governments to early-stage AI startups, and recently served briefly as CTO of a pre-seed AI startup. He previously led the data science function at RBS International and held senior IC roles at HSBC, NatWest Group, and Shawbrook Bank.

    Building Real-World Computer Vision Systems

    This talk will explore practical workflows for building, evaluating, and improving modern computer vision systems. We’ll dive into real-world approaches to dataset curation, model analysis, multimodal AI workflows, and production-ready vision pipelines using open-source technologies.

    The session is designed for engineers, researchers, and AI practitioners looking to better understand how teams are developing and scaling computer vision applications today. Expect practical demos, technical insights, and discussions around the evolving AI tooling ecosystem.

    About the Speaker

    Harpreet Sahota - is a hacker-in-residence and machine learning engineer with a passion for deep learning and generative AI. He’s got a deep interest in RAG, Agents, and Multimodal AI.

    • Photo of the user
    • Photo of the user
    81 attendees
  • Network event
    July 29 - MCP, Agents and Skills Meetup

    July 29 - MCP, Agents and Skills Meetup

    ·
    Online
    Online
    410 attendees from 48 groups

    Join our virtual meetup to hear talks from experts on cutting-edge topics across MCP, Agents and Skills.

    Date, Time and Location

    Jul 29, 2026
    9:00 AM - 11:00 AM PST
    Online.
    Register for the Zoom!

    The Agent Control Plane: Turning Coding Agents into Reliable Engineering Workflows

    AI coding agents are powerful but often unreliable — they hallucinate, lose context, and produce inconsistent results across runs. In this talk, Alex introduces Atomic, an open-source control plane that adds persistent memory, deterministic workflow phases (Research → Specify → Implement → Ship), and human-in-the-loop gates around coding agents like Claude Code and GitHub Copilot. The result: repeatable, auditable engineering workflows that teams can actually trust in production.

    About the Speaker

    Alex Lavaee is an Applied AI engineer at Microsoft Research and the creator of Atomic, an open-source SDK that wraps deterministic, research-to-execution workflows around AI coding agents. He previously conducted AI research at Harvard Medical School and Boston University, and has worked as an MLE and data scientist at companies including Boeing and Themis AI, an MIT CSAIL spinoff.

    UISurf: Toward Universal UI Automation with Cross-Environment Agents

    In this talk, we introduce UISurf, an open-source multimodal agentic UI automation platform in which agents can perceive, reason, and collaborate across browser and desktop environments to complete end-to-end tasks that require interaction with multiple user interfaces.

    UISurf comprises three main components: uisurf-agent, the runtime for UI automation agents; uisurf-admin, the session orchestration and management service; and uisurf-app, the full-stack user application. Its multi-agent architecture includes a planning_agent that transforms natural-language requests into structured execution plans, specialized Browser and Desktop Agents for environment-specific interaction, an automation_agent that coordinates execution and inter-agent handoff through Agent-to-Agent (A2A) communication, and a summarization_agent that produces the final task summary for the user. UISurf supports both autonomous execution and human-in-the-loop supervision, offering a practical and extensible framework for studying and deploying cross-environment UI automation.

    About the Speaker

    Dr. Henry Ruiz is a Research Scientist at Texas A&M University @ AgriLife Research, specializing in Artificial Intelligence (AI) and Remote Sensing. His work focuses on the development of advanced software systems and computational algorithms for analyzing multi-source remote sensing data, including satellite imagery, UAVs (Unmanned Aerial Vehicles), LiDAR (Light Detection and Ranging), and Ground Penetrating Radar (GPR).

    From Manual Workflows to AI-Assisted Skills: Building Reliable Internal Automation

    In this session, I will discuss how teams can turn repetitive manual workflows into reliable AI-assisted and automation-driven “skills.” I will share practical lessons from building internal tools for CAD and engineering workflows, including how automation can reduce manual effort, improve consistency, and support better process control. The talk will also cover why many AI/agent experiments fail when they are not connected to real team workflows, standards, and validation steps. Attendees will walk away with a practical framework for identifying repeatable workflows, designing useful internal tools, and adopting AI assistance without losing accuracy or trust.

    About the Speaker

    Janvi Vijaykumar Saddi - Janvi Saddi is a Computer Science graduate and CAD/Data Automation professional with experience in data center design workflows, AutoCAD automation, process improvement, and data analytics. She currently works as a CAD Tech 2 at Astreya, supporting Google data center design workflows by building internal tools that reduce manual effort, improve accuracy, and streamline engineering processes. Her background also includes SQL, Power BI, market research analytics, and AI-assisted development.

    Building Safe Agent Sandboxes: Let Agents Act Without Breaking Production

    AI agents become truly useful when they can take action, not just generate text. But giving agents access to code, data, and systems raises an important question: how do you let them explore, execute, fail, and improve without putting production at risk?

    In this talk, we'll explore the sandbox pattern for agent systems and how to equip agents with tools to read, write, execute, and iterate within controlled environments while using permissions, human approval, and safety guardrails to keep them reliable. We'll cover practical architectures and lessons learned for building agents that can safely evolve from experimentation to production..

    About the Speaker

    Adonai Vera - Adonai Vera - Machine Learning Engineer & DevRel at Voxel51. With over 7 years of experience building computer vision and machine learning models using TensorFlow, Docker, and OpenCV.

    • Photo of the user
    • Photo of the user
    • Photo of the user
    17 attendees from this group
  • Network event
    Aug 4 - Visual AI in Manufacturing

    Aug 4 - Visual AI in Manufacturing

    ·
    Online
    Online
    111 attendees from 52 groups

    Join our virtual meetup to hear talks from experts on cutting-edge topics at the intersection of manufacturing, AI, ML, and computer vision.

    Date, Time and Location

    Aug 04, 2026
    9:00 AM - 11:00 AM PST
    Online.
    Register for the Zoom!

    Enabling Multimodal Agents on the Edge

    The next generation of AI agents is moving beyond cloud-based text-only models and will interact with the physical multimodal world in real-time. For example in the vision domain, AI agents rely on Vision-Language Models (VLMs) in their backbone. However, deploying massive VLMs with billions of parameters on the edge devices remains a significant engineering hurdle.

    Drawing on our recent ICML and CVPR research papers, this session explores advancements in agentic model optimizations, specifically how distillation and pruning transform 'heavyweight' models into lean, edge-ready engines. Lastly, I present our UI agent running on the actual phone that is being developed by our lab's team.

    About the Speaker

    Denis Gudovskiy is a Distinguished AI Engineer at Panasonic North America where he conducts R&D activities of various core AI methods, including multimodal and hardware-efficient agents, supervised and RL training pipelines, and robustness to out-of-distribution scenarios.

    When the Camera Can’t Be Trusted: Health-Aware Visual AI for Reliable Near-Miss Detection

    Near-miss detection systems are often evaluated as though every camera frame is equally trustworthy, even though blur, poor exposure, occlusion, contamination, and changing lighting can silently degrade the visual evidence used to make safety decisions. This talk presents an online camera-health framework that estimates visual reliability before downstream perception performance significantly deteriorates.

    I will discuss how camera-health signals can support condition-aware evaluation, prioritize human review, reduce unreliable alerts, and trigger appropriate fallback behavior. Drawing from research in safety-critical visual perception, the talk will demonstrate how these principles can be adapted to industrial video systems operating across different cameras, shifts, layouts, and environmental conditions.

    The presentation will also connect camera-health monitoring with rare-event discovery and failure-driven dataset improvement for more trustworthy near-miss detection.

    About the Speaker

    Shiva Aher is a computer vision researcher with a graduate background in computer science from the Georgia Institute of Technology, specializing in artificial intelligence.

    Agentic VLM applications in manufacturing

    Vision Language Models (VLMs) introduce net-new functionality to vision workloads in manufacturing that traditional computer vision models simply do not offer (e.g., open-vocabulary detection, in-context-learning). Even so, fine-tuned models like YOLO offer a level of precision and recall that today's VLMs struggle to match out-of-the-box.

    Through agentic harnesses that coordinate calls to VLMs, we can start to deliver similar reliability on manufacturing-relevant tasks (e.g., many-class, many-instance detection), while also supporting the net new functionalities (e.g., multimodal search) that make VLMs distinct. In this talk, we walk through the design of these harnesses, how you serve them efficiently, and how they deliver value in manufacturing.

    About the Speaker

    Subraiz Ahmed is a member of the Technical Staff at Perceptron AI. He builds the infrastructure to serve frontier vision models. He previously founded a series of startups.

    • Photo of the user
    • Photo of the user
    • Photo of the user
    7 attendees from this group
  • Network event
    Aug 6 - Audio and AI Meetup

    Aug 6 - Audio and AI Meetup

    ·
    Online
    Online
    170 attendees from 51 groups

    Join our virtual meetup to hear talks from experts on cutting-edge topics across AI, ML, and computer vision.

    Date, Time and Location

    Aug 06, 2026
    9:00 AM - 11:00 AM PST
    Online. Register for the Zoom!

    Do Speech Models Actually Understand Speech? Evaluating Speech LLMs Under Realistic Spoken Instruction Conditions

    Speech Large Language Models (SLLMs) are increasingly capable; but are we evaluating them the right way? Most benchmarks rely on text prompts, yet real users interact with these systems through speech, a modality that introduces noise, disfluencies, and stylistic variation that text simply doesn't capture.
    In this talk, we present findings from a systematic study across 11 tasks, 12 languages, and five prompt styles, examining how prompt modality, language, and task type shape SLLM performance.

    About the Speaker

    Maike Züfle is a PhD student at the Karlsruhe Institute of Technology (KIT), working in Prof. Jan Niehues's group on interactive speech systems for more natural human–machine communication. Her research focuses on instruction-following speech models with speech as both input and output, with a recent emphasis on full-duplex systems. Beyond her research, she co-organises the instruction-following and speech translation metrics shared tasks at IWSLT. She is a 2026 Apple Scholar in AI/ML.

    AI based Audio Forensics

    In this presentation, attendees will discover several modules developed by Gradiant for the detection and analysis of synthetically generated or manipulated audio. The session will be delivered by one of the developers involved in the design and implementation of these technologies, providing first-hand insight into their capabilities and underlying methodology.

    The presentation will cover the traceability module, which helps identify the origin of AI-generated content. It will also cover the segment detection tool, designed to locate manipulated regions within an audio recording, as well as the complete audio detection tool, which assesses whether an entire recording has been synthetically generated.

    About the Speaker

    Daniel Paniagua Ares is a research engineer at Gradiant. Graduated in computer engineering from the FIC and with a master's degree in AI from the VIU.

    Curating, Searching, and Evaluating Audio Datasets in FiftyOne

    In this talk, we'll start with the ESC-50 environmental-sound dataset to show how FiftyOne represents audio: browsing clips in the tabular view, rendering spectrograms directly in the sample grid with a custom renderer, and turning sounds into searchable vectors with CLAP embeddings. Then we'll demo a similarity-search panel that lets you query an entire audio collection by example clip or a natural-language prompt to quickly find matching sounds.

    We'll conclude with a live research problem: Audio Moment Retrieval from the DCASE 2026 Challenge, where the goal is to localize the exact moment in a long recording that matches a text query. We'll frame this as temporal detection, evaluate predictions, and visualize ground-truth vs. predicted moments on an interactive timeline to intuitively expose model failure modes.

    Attendees will leave with a concrete blueprint and open code for applying visual data-centric AI practices to their own audio and multimodal datasets.

    About the Speaker

    John Duncan is a Machine Learning Engineer, Customer Success at Voxel51. His research interests include vision, LiDAR, and audio perception for robots and intelligent systems.

    • Photo of the user
    • Photo of the user
    • Photo of the user
    6 attendees from this group

Group links

Organizers

Super Organizer

Members

4,118
See all