
About us
đź–– This group is for AI researchers, machine learning engineers, roboticists and open source enthusiasts.
Every week we bring you a diverse set of speakers working at the cutting edge of AI, machine learning, robotics and computer vision.
This Meetup is sponsored by Voxel51, the multimodal data platform for physical AI. Learn More.
Interested in speaking at a future event? Submit a talk!
By becoming a member of this group you agree to Voxel51's Terms of Service and Privacy Statement and agree to receive occasional emails about upcoming events.
Upcoming events
9
- Network event

Oct 1 - APAC AI, ML and Computer Vision Meetup
·OnlineOnline162 attendees from 55 groupsJoin our APAC time-zone friendly virtual meetup to hear talks from experts on cutting-edge topics across AI, ML, and computer vision.
Time, Date and Location
Oct 1, 2026
6:00 PM - 8:00 PM PDT
Online. Register for the Zoom!Beyond Exact Matches: Detecting Modified 3D Assets at Marketplace Scale
How can a marketplace identify copied 3D assets when their orientation, geometry, or composition has changed? Drawing on my work in 3D content understanding at Roblox, this talk will explore multi-view and rotation-invariant representations for similarity and duplicate detection, including the challenges posed by deformed and fragmented copies.
It will examine how geometric and semantic signals can complement one another, and discuss practical trade-offs in evaluating detection quality and deploying these methods at scale. The presentation will draw on published patent applications and publicly shareable examples to offer practical lessons for engineers building visual search, content-understanding, and marketplace-safety systems.About the Speaker
Phani Harish Wajjala is a Principal Machine Learning Engineer at Roblox specializing in 3D computer vision, multimodal AI, and large-scale content understanding.
Sign Language: Towards Sign Understanding for Robot Autonomy
Navigational signs are common aids for human wayfinding and scene understanding, but are underutilized by robots. We argue that they benefit robot navigation and scene understanding, by directly encoding privileged information on actions, spatial regions, and relations.
Interpreting signs in open-world settings remains a challenge owing to the complexity of scenes and signs, but recent advances in vision-language models (VLMs) make this feasible. To advance progress in this area, we introduce the task of visual sign grounding, which parses locations and associated directions from signs, and maps them to region in the sign’s local environment.
Additionally, we present a baseline approach using VLMs, and demonstrate their promise on the task. We also outline different applications, such as localization and navigation, which benefit from the spatial-symbolic information encoded by navigational signs.
About the Speaker
Nicky Zimmerman I am a postdoctoral researcher in NUS, working on open world navigation. My PhD thesis focused on human-inspired strategies for semantic localization and mapping. Previously, I worked as computer vision algorithm developer in General Motors and Intel.
MeMo: Memory as a Model
Large language models (LLMs) achieve strong performance across a wide range of tasks, but remain frozen after pretraining until subsequent updates. Many real-world applications require timely, domain-specific information, motivating the need for efficient mechanisms to incorporate new knowledge.
In this paper, we introduce MeMo (Memory as a Model), a modular framework that encodes new knowledge into a dedicated Memory model while keeping the LLM unchanged. Compared to existing methods, MeMo offers several advantages: (a) it captures complex cross-document relationships, (b) it is robust to retrieval noise, (c) it avoids catastrophic forgetting in the LLM, (d) it does not require access to the LLM’s weights or output logits that enabling plug-and-play integration with both open and proprietary LLMs, and (e) its retrieval cost is independent of corpus size at inference time.
Our experiments on three benchmarks, BrowseComp-Plus, NarrativeQA, and MuSiQue, show that MeMo achieves strong performance compared to existing methods across diverse settings.
About the Speaker
Arun Verma is a Postdoctoral Associate at the Singapore-MIT Alliance for Research and Technology Centre, where he works with Daniela Rus, Armando Solar-Lezama, and Bryan Low.
4 attendees from this group - Network event

Oct 14 - Advances in AI at SDSU
·OnlineOnline123 attendees from 55 groupsJoin our virtual meetup to hear talks from AI researchers at San Diego State University.
Date, Time Location
Oct 14, 2026
9:00 AM - 11:00 AM PST
Online. Register for the Zoom!Agent as Policy for Robotic Manipulation
This talk will introduces how a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training. We introduce Agent as Policy (AGP), which places task planning and execution under the agent’s control.
About the Speaker
Xiaobai Liu is a Professor of Computer Science at San Diego State University (SDSU), where he directs the Machine Vision and Perception Lab. Prior to joining SDSU in 2015, he conducted research and taught at UCLA.
Grounding Multimodal Open-World Learning for Physical AI
Physical AI systems such as robots, unmanned vehicles, and embodied agents must perceive and reason about a world that is dynamic, unstructured, and rarely matches their training distribution. Yet most multimodal models remain brittle when confronted with novel objects, unseen conditions, and a messy unstructured environment.
This talk develops how grounding perception across modalities and environments can make open-world learning more robust for physically situated systems. Drawing on our recent work, I will highlight the core challenges of multimodal open-world learning including out-of-distribution detection, distribution shift generalization, and cascaded semantic grounding and point toward multimodal systems that stay reliable when deployed in the unpredictable open world.
About the Speakers
Salimeh Sekeh is an Associate Professor of Computer Science at San Diego State University (SDSU), where she directs the Sekeh Laboratory. Her recognition includes an NSF CAREER Award and a Cisco research gift (both 2022), and the Maine College of Engineering and Computing Early Career Research Award (2023), along with multiple industry and federal research awards.
Mary Wisell is a second-year PhD student in the Sekeh Lab and leads the lab's work on environment-aware OOD detection and cascaded failure analysis for multimodal intelligence with several publications in top-tier Machine Learning and computer vision conferences.
Model-Free, Position-Free Signal Source Seeking Using Unmanned Maritime Systems
Long-range signal source detection in open-world environments presents a significant challenge, mainly due to the detrimental effects of the environment on signals propagation paths and presence of extraneous signal sources. While distributed static sensing systems are often employed, achieving scalability and comprehensive coverage across expansive areas is cost-prohibitive.
One affordable solution involves using low-cost autonomous unmanned vehicles (AUVs) that can leverage their mobility and actively explore the environment using extremum seeking control (ESC) algorithms. This talk presents novel ESC approaches to steer AUVs to sources of interest in an a priori unknown, highly non-convex map.
About the Speaker
Zahra Nili Ahmadabadi is Associate Professor with the Mechanical Engineering Department at San Diego State University (SDSU). She is a recipient of the ASME rising star award and ARO Early career award.
3 attendees from this group - Network event

Oct 15 - AI, ML, and Computer Vision Meetup
·OnlineOnline94 attendees from 55 groupsJoin our virtual meetup to hear talks from experts on cutting-edge topics across AI, ML, and computer vision.
Time, Place and Location
Oct 15, 2026
9:00 AM - 11:00 AM PST
Online. Register for the Zoom!Testing AI Systems in Production: Data Quality, Drift, and Model Evaluation
AI systems can pass offline evaluation and still fail in production when real-world data changes, features become stale, labels or feedback signals are incomplete, or model behavior drifts away from expected outcomes. This talk shares practical patterns for testing and evaluating AI systems after deployment, including data quality checks, drift detection, online/offline metric comparison, model monitoring, and rollback analysis.
Using personalization and recommendation systems as examples, we will examine how teams can build evaluation workflows that catch quality issues before users do. Attendees will leave with a practical checklist for making AI-backed systems easier to evaluate, debug, and operate as data changes over time.
About the Speaker
Jayakumar Ramalingam is a Staff Software Engineer and Cloud Architect at SiriusXM with over 16 years of experience building cloud-native platforms, real-time data pipelines, resilient APIs, and AI/ML-enabled applications at production scale.
Where Should Your Model Live? A Framework for Tiering Computer Vision Deployments
Where should a computer vision model actually run - on-device, near the edge, or in the cloud? It's a decision that looks simple until requirements like latency, cost, connectivity, and update cadence start pulling in different directions, often revealing themselves only after deployment.
Drawing on hands-on experience developing and deploying CV models across Hailo, Nvidia, Qualcomm and AWS platforms, this talk introduces a practical framework for tiering computer vision deployments based on real project requirements and constraints. Discussion will include what changes at each tier - from development to deployment to monitoring and update strategy - with relevant industry examples.
Attendees will leave with a set of questions or a framework they can use to place their own CV projects into the right tier.
About the Speaker
Ajaykumaar Sivacoumare is an AI Software Engineer specializing in computer vision and edge AI, with production experience developing and deploying CV models across Nvidia, Hailo, Qualcomm and AWS-based platforms.
From 2D Slices to 3D Tumors: Lightweight Volumetric Detection Without Heavy 3D Networks
Slice-wise 2D detectors are fast and scalable, but they struggle to produce reliable 3D bounding boxes from volumetric medical data. This talk presents YOLO-PVC, a lightweight post-processing framework that consolidates slice-wise YOLO detections into coherent 3D bounding boxes using percentile-based geometric aggregation and a minimal MLP calibration module.
This talk demonstrates consistent improvements in volumetric IoU across three liver tumor categories i.e., HCC, CCA, and Mixed, without requiring dense 3D annotations or memory-intensive architectures. The talk covers the clinical motivation, the technical approach, and practical lessons from deploying computer vision on real hospital MRI data.
About the Speaker
Talha Waqas is a second-year PhD student at ESME Research Lab, Paris and LISSI, Université Paris-Est, working on computer vision applied to medical imaging, with a focus on tumor classification, detection, and segmentation in multi-phase liver MRI.
The Two-Loop Architecture for Voice AI
Building responsive voice AI requires balancing latency with intelligence. This talk introduces a practical architecture that separates real-time conversation from asynchronous reasoning, enabling richer interactions without slowing the user experience. The session covers reusable design patterns drawn from production-inspired conversational AI systems.
About the Speaker
Abhinav Tushar is an ML engineer and researcher specializing in Conversational AI and Speech Technology.
5 attendees from this group - Network event

Oct 22 - Advances in AI at Virginia Tech
·OnlineOnline52 attendees from 55 groupsJoin our virtual meetup to hear talks from AI researchers at Virginia Tech!
Date, Time and Location
Oct 22, 2026
9:00 AM - 11:00 AM PST
Online. Register for the Zoom!Multi-Agent Communication: A framework, diagnostic and mechanistic perspective
Multi-agent LLM systems are increasingly used for collaborative reasoning, debate, and consensus, yet their communication dynamics remain poorly understood. This talk presents a framework for studying multi-agent communication through diagnostic and mechanistic perspectives.
I will discuss CONSENSAGENT, which improves consensus by mitigating sycophancy, alongside our diagnostic work on communication patterns and failure modes in real-world multi-agent debates. I will then present ongoing work that moves toward a mechanistic understanding of how these interaction patterns arise internally, with the broader goal of making multi-agent systems more interpretable, reliable, and controllable.
About the Speaker
Priya Pitre I am an Ph.D student in the Computer Science Department at Virginia Tech (VT), co-advised by Dr. Xuan Wang and Dr. Naren Ramakrishnan.
Exposing and Improving Fine-Grained Visual Grounding Abilities of Lightweight Multimodal LLMs
Lightweight multimodal LLMs can localize whole objects effectively, yet often struggle when a query targets a small object part or fine-grained visual detail. This talk presents a reasoning-guided framework that teaches compact models to ground parts through an explicit coarse-to-fine process: first locating the parent object, then identifying the requested part.
A part-aware reinforcement-learning objective provides stage-wise rewards for object accuracy, part containment, and the consistency of the model’s self-critique. Using these techniques, a compact 4B-parameter model achieves state-of-the-art zero-shot part grounding while preserving its object-level performance.
These advances can be used to enable lightweight MLLMs to support detail-oriented tasks in biology and robotics.
About the Speaker
Kazi Mehrab I a CS PhD candidate at Virginia Tech, where I currently focus on multimodal LLMs and computer vision tasks, including visual perception, reasoning and grounding.
Understanding Visual Generative Models for Precise Control
Despite remarkable progress in image and video generation, translating user intent into precise and consistent visual outputs remains a challenge. This talk explores how understanding the representations within generative models can enable finer control over what they create.
It connects semantic image editing with compositional generation, examining how visual concepts can be isolated, manipulated, and combined while preserving their identity and surrounding content. Building on these insights, structured visual inputs provide a way to express complex intent through subject references, poses, and spatial layouts.
The discussion then extends from images to video, where representations must evolve to preserve scene continuity while accommodating motion and change. Together, these directions establish a unified perspective on how visual representations can support controllable editing, composition, and coherent generation across space and time.
About the Speaker
Yusuf Dalva is a Ph.D. candidate at Virginia Tech, advised by Pinar Yanardag and affiliated with the Sanghani Center for Artificial Intelligence and Data Analytics.
2 attendees from this group
Past events
191

