BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Meetup//Meetup Calendar 1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
NAME:Rome AI, Machine Learning and Computer Vision Meetup
X-WR-CALNAME:Rome AI, Machine Learning and Computer Vision Meetup
BEGIN:VTIMEZONE
TZID:Europe/Rome
TZURL:http://tzurl.org/zoneinfo-outlook/Europe/Rome
X-LIC-LOCATION:Europe/Rome
BEGIN:DAYLIGHT
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
TZNAME:CEST
DTSTART:19700329T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=-1SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
TZNAME:CET
DTSTART:19701025T030000
RRULE:FREQ=YEARLY;BYMONTH=10;BYDAY=-1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
UID:event_315682798@meetup.com
SEQUENCE:1
DTSTAMP:20260729T190735Z
DTSTART;TZID=Europe/Rome:20260804T180000
DTEND;TZID=Europe/Rome:20260804T200000
SUMMARY:Aug 4 - Visual AI in Manufacturing
DESCRIPTION:Rome AI\, Machine Learning and Computer Vision Meetup\nJoin ou
 r virtual meetup to hear talks from experts on cutting-edge topics at the 
 intersection of manufacturing\, AI\, ML\, and computer vision.\n\n**Date\,
  Time and Location**\n\nAug 04\, 2026\n9:00 AM - 11:00 AM PST\nOnline. **[
 Register for the Zoom!](https://voxel51.com/events/visual-ai-in-manufactur
 ing-meetup-august-4-2026)**\n\n**Enabling Multimodal Agents on the Edge**\
 n\nThe next generation of AI agents is moving beyond cloud-based text-only
  models and will interact with the physical multimodal world in real-time.
  For example in the vision domain\, AI agents rely on Vision-Language Mode
 ls (VLMs) in their backbone. However\, deploying massive VLMs with billion
 s of parameters on the edge devices remains a significant engineering hurd
 le.\n\nDrawing on our recent ICML and CVPR research papers\, this session 
 explores advancements in agentic model optimizations\, specifically how di
 stillation and pruning transform 'heavyweight' models into lean\, edge-rea
 dy engines. Lastly\, I present our UI agent running on the actual phone th
 at is being developed by our lab's team.\n\n*About the Speaker*\n\n[Denis 
 Gudovskiy](https://www.linkedin.com/in/gudovskiy/) is a Distinguished AI E
 ngineer at Panasonic North America where he conducts R&D activities of var
 ious core AI methods\, including multimodal and hardware-efficient agents\
 , supervised and RL training pipelines\, and robustness to out-of-distribu
 tion scenarios.\n\n**When the Camera Can’t Be Trusted: Health-Aware Visu
 al AI for Reliable Near-Miss Detection**\n\nNear-miss detection systems ar
 e often evaluated as though every camera frame is equally trustworthy\, ev
 en though blur\, poor exposure\, occlusion\, contamination\, and changing 
 lighting can silently degrade the visual evidence used to make safety deci
 sions. This talk presents an online camera-health framework that estimates
  visual reliability before downstream perception performance significantly
  deteriorates.\n\nI will discuss how camera-health signals can support con
 dition-aware evaluation\, prioritize human review\, reduce unreliable aler
 ts\, and trigger appropriate fallback behavior. Drawing from research in s
 afety-critical visual perception\, the talk will demonstrate how these pri
 nciples can be adapted to industrial video systems operating across differ
 ent cameras\, shifts\, layouts\, and environmental conditions.\n\nThe pres
 entation will also connect camera-health monitoring with rare-event discov
 ery and failure-driven dataset improvement for more trustworthy near-miss 
 detection.\n\n*About the Speaker*\n\n[Shiva Aher](https://www.linkedin.com
 /in/shivaaher/) is a computer vision researcher with a graduate background
  in computer science from the Georgia Institute of Technology\, specializi
 ng in artificial intelligence.\n\n**Agentic VLM applications in manufactur
 ing**\n\nVision Language Models (VLMs) introduce net-new functionality to 
 vision workloads in manufacturing that traditional computer vision models 
 simply do not offer (e.g.\, open-vocabulary detection\, in-context-learnin
 g). Even so\, fine-tuned models like YOLO offer a level of precision and r
 ecall that today's VLMs struggle to match out-of-the-box.\n\nThrough agent
 ic harnesses that coordinate calls to VLMs\, we can start to deliver simil
 ar reliability on manufacturing-relevant tasks (e.g.\, many-class\, many-i
 nstance detection)\, while also supporting the net new functionalities (e.
 g.\, multimodal search) that make VLMs distinct. In this talk\, we walk th
 rough the design of these harnesses\, how you serve them efficiently\, and
  how they deliver value in manufacturing.\n\n*About the Speaker*\n\n[Subra
 iz Ahmed](https://www.linkedin.com/in/subraiz/) is a member of the Technic
 al Staff at Perceptron AI. He builds the infrastructure to serve frontier 
 vision models. He previously founded a series of startups.
URL;VALUE=URI:https://www.meetup.com/rome-ai-machine-learning-and-computer
 -vision-meetup/events/315682798/
STATUS:CONFIRMED
CREATED:20260714T180942Z
LAST-MODIFIED:20260714T180942Z
CLASS:PUBLIC
END:VEVENT
BEGIN:VEVENT
UID:event_315389154@meetup.com
SEQUENCE:1
DTSTAMP:20260729T190735Z
DTSTART;TZID=Europe/Rome:20260806T180000
DTEND;TZID=Europe/Rome:20260806T200000
SUMMARY:Aug 6 - Audio and AI Meetup
DESCRIPTION:Rome AI\, Machine Learning and Computer Vision Meetup\nJoin ou
 r virtual meetup to hear talks from experts on cutting-edge topics across 
 AI\, ML\, and computer vision.\n\n**Date\, Time and Location**\n\nAug 06\,
  2026\n9:00 AM - 11:00 AM PST\n**[Online. Register for the Zoom!](https://
 voxel51.com/events/audio-and-ai-meetup-august-6-2026)**\n\n**Do Speech Mod
 els Actually Understand Speech? Evaluating Speech LLMs Under Realistic Spo
 ken Instruction Conditions**\n\nSpeech Large Language Models (SLLMs) are i
 ncreasingly capable\; but are we evaluating them the right way? Most bench
 marks rely on text prompts\, yet real users interact with these systems th
 rough speech\, a modality that introduces noise\, disfluencies\, and styli
 stic variation that text simply doesn't capture.\nIn this talk\, we presen
 t findings from a systematic study across 11 tasks\, 12 languages\, and fi
 ve prompt styles\, examining how prompt modality\, language\, and task typ
 e shape SLLM performance.\n\n*About the Speaker*\n\n[Maike Züfle ](https:
 //voxel51.com/events/www.linkedin.com/in/maike-z%C3%BCfle)is a PhD student
  at the Karlsruhe Institute of Technology (KIT)\, working in Prof. Jan Nie
 hues's group on interactive speech systems for more natural human–machin
 e communication. Her research focuses on instruction-following speech mode
 ls with speech as both input and output\, with a recent emphasis on full-d
 uplex systems. Beyond her research\, she co-organises the instruction-foll
 owing and speech translation metrics shared tasks at IWSLT. She is a 2026 
 Apple Scholar in AI/ML.\n\n**AI based Audio Forensics**\n\nIn this present
 ation\, attendees will discover several modules developed by Gradiant for 
 the detection and analysis of synthetically generated or manipulated audio
 . The session will be delivered by one of the developers involved in the d
 esign and implementation of these technologies\, providing first-hand insi
 ght into their capabilities and underlying methodology.\n\nThe presentatio
 n will cover the traceability module\, which helps identify the origin of 
 AI-generated content. It will also cover the segment detection tool\, desi
 gned to locate manipulated regions within an audio recording\, as well as 
 the complete audio detection tool\, which assesses whether an entire recor
 ding has been synthetically generated.\n\n*About the Speaker*\n\n[Daniel P
 aniagua Ares ](https://voxel51.com/events/www.linkedin.com/in/daniel-pania
 gua-ares-96252918a)is a research engineer at Gradiant. Graduated in comput
 er engineering from the FIC and with a master's degree in AI from the VIU.
 \n\n**Curating\, Searching\, and Evaluating Audio Datasets in FiftyOne**\n
 \nIn this talk\, we'll start with the ESC-50 environmental-sound dataset t
 o show how FiftyOne represents audio: browsing clips in the tabular view\,
  rendering spectrograms directly in the sample grid with a custom renderer
 \, and turning sounds into searchable vectors with CLAP embeddings. Then w
 e'll demo a similarity-search panel that lets you query an entire audio co
 llection by example clip or a natural-language prompt to quickly find matc
 hing sounds.\n\nWe'll conclude with a live research problem: Audio Moment 
 Retrieval from the DCASE 2026 Challenge\, where the goal is to localize th
 e exact moment in a long recording that matches a text query. We'll frame 
 this as temporal detection\, evaluate predictions\, and visualize ground-t
 ruth vs. predicted moments on an interactive timeline to intuitively expos
 e model failure modes.\n\nAttendees will leave with a concrete blueprint a
 nd open code for applying visual data-centric AI practices to their own au
 dio and multimodal datasets.\n\n*About the Speaker*\n\n[John Duncan](https
 ://www.linkedin.com/in/john-a-duncan/) is a Machine Learning Engineer\, Cu
 stomer Success at Voxel51. His research interests include vision\, LiDAR\,
  and audio perception for robots and intelligent systems.\n\n**Real-Time A
 SR at 4x on Consumer Hardware: The Meetily Architecture**\n\nThis talk cov
 ers the engineering behind Meetily\, an open-source meeting assistant that
  runs Whisper and NVIDIA Parakeet transcription entirely on-device. We'll 
 walk through how we got Parakeet to roughly 4x real-time on consumer hardw
 are\, and the specific points where it still falls over.\n\nWe'll also get
  into the honest trade-offs between local and cloud inference: latency\, a
 ccuracy\, cost\, and what you actually give up by choosing one over the ot
 her. Wrapping ML inference in a Rust/Tauri desktop app came with its own c
 osts\, which we'll unpack as well.\n\nFinally\, we'll look at what "fully 
 local" really means at an architecture level\, where that boundary sits\, 
 and how easily it leaks once you add model downloads\, integrations\, or a
  pluggable LLM backend.\n\n*About the Speaker*\n\n[Sandeep Zachariah](http
 s://www.linkedin.com/in/sandeepzachariah/) is the Founder and CEO of Zackr
 iya Solutions and the leads the team behind Meetily\, an open-source\, pri
 vacy-first meeting assistant that runs Whisper and NVIDIA Parakeet transcr
 iption entirely on-device. He brings a rare full-stack perspective on audi
 o AI — from low-level embedded systems and hardware acceleration up thro
 ugh real-time ASR and local LLM summarization — with deep experience dep
 loying speech and ML models across servers\, GPUs and consumer hardware.
URL;VALUE=URI:https://www.meetup.com/rome-ai-machine-learning-and-computer
 -vision-meetup/events/315389154/
STATUS:CONFIRMED
CREATED:20260623T223549Z
LAST-MODIFIED:20260623T223549Z
CLASS:PUBLIC
END:VEVENT
BEGIN:VEVENT
UID:event_315700954@meetup.com
SEQUENCE:1
DTSTAMP:20260729T190735Z
DTSTART;TZID=Europe/Rome:20260806T180000
DTEND;TZID=Europe/Rome:20260806T200000
SUMMARY:Aug 6 - Audio and AI Meetup
DESCRIPTION:Rome AI\, Machine Learning and Computer Vision Meetup\nJoin us
  on Aug 6 for a special edition of the AI\, ML\, and Computer Vision Meetu
 p focused on audio use cases!\n\n**Date\, Time and Location**\n\nAug 06\, 
 2026\n9:00 AM - 11:00 AM PST\nOnline. **[Register for the Zoom](https://vo
 xel51.com/events/audio-and-ai-meetup-august-6-2026)**\n\n**Do Speech Model
 s Actually Understand Speech? Evaluating Speech LLMs Under Realistic Spoke
 n Instruction Conditions**\n\nSpeech Large Language Models (SLLMs) are inc
 reasingly capable\; but are we evaluating them the right way? Most benchma
 rks rely on text prompts\, yet real users interact with these systems thro
 ugh speech\, a modality that introduces noise\, disfluencies\, and stylist
 ic variation that text simply doesn't capture.\nIn this talk\, we present 
 findings from a systematic study across 11 tasks\, 12 languages\, and five
  prompt styles\, examining how prompt modality\, language\, and task type 
 shape SLLM performance.\n\n*About the Speaker*\n\n[Maike Züfle](https://v
 oxel51.com/events/www.linkedin.com/in/maike-z%C3%BCfle) is a PhD student a
 t the Karlsruhe Institute of Technology (KIT)\, working in Prof. Jan Niehu
 es's group on interactive speech systems for more natural human–machine 
 communication.\n\n**AI based Audio Forensics**\n\nIn this presentation\, a
 ttendees will discover several modules developed by Gradiant for the detec
 tion and analysis of synthetically generated or manipulated audio. The ses
 sion will be delivered by one of the developers involved in the design and
  implementation of these technologies\, providing first-hand insight into 
 their capabilities and underlying methodology.\n\nThe presentation will co
 ver the traceability module\, which helps identify the origin of AI-genera
 ted content. It will also cover the segment detection tool\, designed to l
 ocate manipulated regions within an audio recording\, as well as the compl
 ete audio detection tool\, which assesses whether an entire recording has 
 been synthetically generated.\n\n*About the Speaker*\n\n[Daniel Paniagua A
 res](https://voxel51.com/events/www.linkedin.com/in/daniel-paniagua-ares-9
 6252918a) is a research engineer at Gradiant. Graduated in computer engine
 ering from the FIC and with a master's degree in AI from the VIU.\n\n**Cur
 ating\, Searching\, and Evaluating Audio Datasets in FiftyOne**\n\nIn this
  talk\, we'll start with the ESC-50 environmental-sound dataset to show ho
 w FiftyOne represents audio: browsing clips in the tabular view\, renderin
 g spectrograms directly in the sample grid with a custom renderer\, and tu
 rning sounds into searchable vectors with CLAP embeddings. Then we'll demo
  a similarity-search panel that lets you query an entire audio collection 
 by example clip or a natural-language prompt to quickly find matching soun
 ds.\n\nWe'll conclude with a live research problem: Audio Moment Retrieval
  from the DCASE 2026 Challenge\, where the goal is to localize the exact m
 oment in a long recording that matches a text query. We'll frame this as t
 emporal detection\, evaluate predictions\, and visualize ground-truth vs. 
 predicted moments on an interactive timeline to intuitively expose model f
 ailure modes.\n\nAttendees will leave with a concrete blueprint and open c
 ode for applying visual data-centric AI practices to their own audio and m
 ultimodal datasets.\n\n*About the Speaker*\n\n[John Duncan](https://www.li
 nkedin.com/in/john-a-duncan/) is a Machine Learning Engineer\, Customer Su
 ccess at Voxel51. His research interests include vision\, LiDAR\, and audi
 o perception for robots and intelligent systems.
URL;VALUE=URI:https://www.meetup.com/rome-ai-machine-learning-and-computer
 -vision-meetup/events/315700954/
STATUS:CONFIRMED
CREATED:20260715T231958Z
LAST-MODIFIED:20260715T231958Z
CLASS:PUBLIC
END:VEVENT
BEGIN:VEVENT
UID:event_315683203@meetup.com
SEQUENCE:1
DTSTAMP:20260729T190735Z
DTSTART;TZID=Europe/Rome:20260811T180000
DTEND;TZID=Europe/Rome:20260811T200000
SUMMARY:Aug 11 - Debugging Physical AI Models at Scale with Multimodal Dat
 a Workshop
DESCRIPTION:Rome AI\, Machine Learning and Computer Vision Meetup\nJoin Vo
 xel51 for a live workshop on how multimodal data workflows in FiftyOne hel
 p teams inspect\, search\, and debug complex Physical AI datasets and expl
 ain black-box model behavior at scale. We’ll show how teams can work wit
 h synchronized video and sensor data\, query for similar scenarios across 
 their datasets\, and uncover patterns behind model failures faster than pl
 ayback-only visualization tools allow.\n\n**Date\, Time and Location**\n\n
 Aug 11\, 2026\n9:00 AM - 10:00 AM PST\nOnline. **[Register for the Zoom!](
 https://voxel51.com/events/debugging-physical-ai-models-at-scale-with-mult
 imodal-data-august-11-2026)**\n\nAs robotics and autonomous vehicle teams 
 move from traditional perception models to end-to-end Physical AI systems\
 , understanding model behavior is becoming harder than ever. These models 
 ingest synchronized inputs from cameras\, sensors\, and other data streams
 \, but their decisions can be difficult to explain\, reproduce\, and impro
 ve.\n\nYou’ll learn how to use multimodal data to investigate questions 
 like: when did the model swerve\, miss an object\, misinterpret a scene\, 
 or behave unexpectedly — and how can you find every similar moment acros
 s your dataset?\n\nDesigned for robotics\, AV\, and machine learning teams
 \, this session will show how FiftyOne helps turn multimodal data into a s
 calable workflow for model evaluation\, debugging\, and improvement.
URL;VALUE=URI:https://www.meetup.com/rome-ai-machine-learning-and-computer
 -vision-meetup/events/315683203/
STATUS:CONFIRMED
CREATED:20260714T183337Z
LAST-MODIFIED:20260714T183337Z
CLASS:PUBLIC
END:VEVENT
BEGIN:VEVENT
UID:event_315684124@meetup.com
SEQUENCE:1
DTSTAMP:20260729T190735Z
DTSTART;TZID=Europe/Rome:20260813T180000
DTEND;TZID=Europe/Rome:20260813T190000
SUMMARY:Aug 13 - How to Build Vision Data Agents with Tools\, Skills\, and
  MCP
DESCRIPTION:Rome AI\, Machine Learning and Computer Vision Meetup\nIn this
  session\, you’ll learn how to build production-ready AI agents that can
  reason over your data\, automate complex tasks\, and integrate seamlessly
  into your existing stack using tools\, skills\, and the Model Context Pro
 tocol (MCP).\n\n**Date\, Time and Location**\n\nAug 13\, 2026\n9:00 AM - 1
 0:00 AM PST\nOnline. **[Register for the Zoom!](https://voxel51.com/events
 /how-to-build-vision-data-agents-with-tools-skills-and-mcp-august-13-2026)
 **\n\nWe’ll walk through how modern agentic systems move beyond simple p
 rompts—leveraging structured tools like dataset operations\, embeddings\
 , evaluation pipelines\, and model execution to take real action. You’ll
  see how these agents can tag data\, run inference\, evaluate performance\
 , and surface insights automatically\, all within a unified workflow.\n\nB
 y combining natural language interfaces with programmable building blocks\
 , teams can dramatically reduce manual effort\, accelerate experimentation
 \, and unlock faster decision-making across the ML lifecycle.\n\nWhether y
 ou're building data-centric AI systems\, managing large-scale vision datas
 ets\, or exploring agentic workflows for the first time\, this session wil
 l give you a practical blueprint for getting started.\n\n*About the Speake
 r*\n\n[Adonai Vera](https://www.linkedin.com/in/adonai-vera/) \\- Machine 
 Learning Engineer & DevRel at Voxel51\\. With over 7 years of experience b
 uilding computer vision and machine learning models using TensorFlow\\\, D
 ocker\\\, and OpenCV\\. I started as a software developer\\\, moved into A
 I\\\, led teams\\\, and served as CTO\\. Today\\\, I connect code and comm
 unity to build open\\\, production\\-ready AI\\\, making technology simple
 \\\, accessible\\\, and reliable\\.
URL;VALUE=URI:https://www.meetup.com/rome-ai-machine-learning-and-computer
 -vision-meetup/events/315684124/
STATUS:CONFIRMED
CREATED:20260714T192455Z
LAST-MODIFIED:20260714T192455Z
CLASS:PUBLIC
END:VEVENT
BEGIN:VEVENT
UID:event_314739436@meetup.com
SEQUENCE:1
DTSTAMP:20260729T190735Z
DTSTART;TZID=Europe/Rome:20260820T180000
DTEND;TZID=Europe/Rome:20260820T200000
SUMMARY:Aug 20 - Cold Pool to Hot Queue: Annotation Curation with FiftyOne
DESCRIPTION:Rome AI\, Machine Learning and Computer Vision Meetup\nIn this
  hands-on workshop\, you'll use [FiftyOne ](https://docs.voxel51.com/)to r
 un the full rare-class mining loop end-to-end on a large unlabeled image p
 ool: compress the pool with near-duplicate detection\, embed images with a
  modern vision backbone\, mine candidate positives via seeded similarity f
 rom a tiny labeled set\, confirm them through targeted human review\, and 
 prioritize the survivors for annotation using representativeness and uniqu
 eness scores.\n\n**Time\, Date and Location**\n\nAug 20\, 2026\n9:00 AM - 
 11:00 AM PST\, 2026\nOnline. **[Register for the Zoom!](https://voxel51.co
 m/events/from-cold-pool-to-hot-queue-annotation-curation-with-fiftyone-aug
 ust-20-2026)**\n\n**What You'll Walk Away With**\n\n* A working FiftyOne p
 ipeline for finding rare classes in any visual dataset you own\n* A repeat
 able four-stage funnel — compress\, mine\, confirm\, prioritize — with
  a clear objective at each stage\n* A fine-tuned detector that demonstrabl
 y outperforms one trained on the same number of randomly sampled images\n*
  The mental model that data curation — not architecture or hyperparamete
 rs — is the highest-leverage thing you can do to improve a rare-class de
 tector
URL;VALUE=URI:https://www.meetup.com/rome-ai-machine-learning-and-computer
 -vision-meetup/events/314739436/
STATUS:CONFIRMED
CREATED:20260511T191209Z
LAST-MODIFIED:20260511T191209Z
CLASS:PUBLIC
END:VEVENT
BEGIN:VEVENT
UID:event_315684972@meetup.com
SEQUENCE:1
DTSTAMP:20260729T190735Z
DTSTART;TZID=Europe/Rome:20260825T180000
DTEND;TZID=Europe/Rome:20260825T200000
SUMMARY:Aug 25 - Advances in AI at NYU
DESCRIPTION:Rome AI\, Machine Learning and Computer Vision Meetup\nJoin ou
 r virtual meetup to hear talks from researchers at NYU on cutting-edge top
 ics across AI\, ML\, and computer vision.\n\n**Date\, Time and Location**\
 n\nAug 25\, 2026\n9:00 AM - 11:00 AM PST\nOnline. **[Register for the Zoom
 !](https://voxel51.com/events/advances-in-ai-at-nyu-august-25-2026)**\n\n*
 *Using Computer Vision to Advance the Sciences**\n\nI'll present some of o
 ur ongoing work on using computer vision to create impact in the sciences.
  These target a two areas\, solar physics and evolutionary biology\, that 
 deal with objects of radically different sizes but are unified by a need f
 or high quality\, trustworthy data.\n\nI'll show off our efforts\, done in
  collaboration with domain experts\, that aim to produce the best possible
  maps of the Sun's powerful magnetic field and have created some of the wo
 rld's largest repositories of data about bird morphology.\n\n*About the Sp
 eaker*\n\n[David Fouhey](https://www.linkedin.com/in/david-fouhey-64212037
 9/) is an Associate Professor at New York University and a research scient
 ist at Polymathic AI. Before joining NYU\, he received a PhD in robotics f
 rom Carnegie Mellon\, was a postdoc at UC Berkeley\, and was a professor a
 t University of Michigan.\n\n**Solaris: Building a Multiplayer Video World
  Model in Minecraft**\n\nThis talk will introduce Solaris: a multiplayer v
 ideo world model in Minecraft. I will first present SolarisEngine\, the so
 ftware platform we built to simulate realistic multiplayer gameplay betwee
 n bots at scale\, enabling us to collect a large training dataset of align
 ed multiplayer actions and frames.\n\nI will then discuss our staged train
 ing pipeline\, starting with single-player pre-training before converting 
 the model into a long-horizon multiplayer generator through bidirectional 
 training\, followed by causal training\, and concluding with Self Forcing.
  I will also cover our memory-efficient implementation of Self Forcing\, c
 alled Checkpointed Self Forcing.\n\nFinally\, I will showcase generated vi
 deos illustrating how Solaris maintains coherent long-horizon multiplayer 
 interactions.\n\n*About the Speaker*\n\n[Oscar Michel ](https://www.linked
 in.com/in/oscar-michel-82a162145/)is a PhD student at NYU advised by Prof.
  Saining Xie. His research studies world models: generative models of agen
 ts interacting in an environment.\n\n**Closing the human to robot gap for 
 dexterous hands**\n\nCollecting task-specific robot data for multi-fingere
 d hands is challenging due to the many difficulties that arise in teleoper
 ation. That is why recently there has been a major focus on learning robot
  policies directly from human demonstrations. However\, human demonstratio
 ns are difficult to work with\; there is a major morphological and visual 
 gap between human and robot hands\, as well as between the environments th
 ey operate in.\n\nIn this talk\, I'd like to discuss my efforts on closing
  this gap.\n\n*About the Speaker*\n\n[Irmak Guzey](https://www.linkedin.co
 m/in/irmak-guzey-6a9010175) I'm Irmak (she/her)\, a rising 3rd year PhD st
 udent at New York University\, currently advised by Lerrel Pinto. My resea
 rch focuses on robot learning for dexterous manipulation. I have been awar
 ded a Fulbright scholarship and NYU's Best Master's Thesis Award in the pa
 st.
URL;VALUE=URI:https://www.meetup.com/rome-ai-machine-learning-and-computer
 -vision-meetup/events/315684972/
STATUS:CONFIRMED
CREATED:20260714T204416Z
LAST-MODIFIED:20260714T204416Z
CLASS:PUBLIC
END:VEVENT
BEGIN:VEVENT
UID:event_314999424@meetup.com
SEQUENCE:1
DTSTAMP:20260729T190735Z
DTSTART;TZID=Europe/Rome:20260827T180000
DTEND;TZID=Europe/Rome:20260827T200000
SUMMARY:Aug 27 - AI\, ML\, and Computer Vision Meetup
DESCRIPTION:Rome AI\, Machine Learning and Computer Vision Meetup\nJoin ou
 r virtual meetup to hear talks from experts on cutting-edge topics across 
 AI\, ML\, and computer vision.\n\nDate\, Time\, and Location\n\nAug 27\, 2
 026\n9:00 AM - 11:00 AM PST\nOnline. **[Register for the Zoom!](https://vo
 xel51.com/events/ai-ml-and-computer-vision-meetup-august-27-2026)**\n\n**R
 obust Concept Protection against Diffusion-Based Image Editing and Persona
 lization**\n\nDiffusion-based image editing and personalization models hav
 e made it increasingly easy to manipulate and replicate visual concepts fr
 om only a few reference images. However\, existing protection methods ofte
 n overfit to a single attack model and fail to generalize across diverse e
 diting pipelines.\n\nIn this presentation\, I will discuss recent advances
  in concept protection for generative AI systems\, focusing on targeted pe
 rturbation strategies and style-sensitive diffusion representations. I wil
 l also present experimental findings across multiple editing and fine-tuni
 ng scenarios\, highlighting the challenges of robustness\, transferability
 \, and imperceptibility in practical protection settings. Finally\, I will
  discuss open problems and future directions toward trustworthy generative
  content ownership.\n\n*About the Speaker*\n\n[Qiuyu Tang](https://www.lin
 kedin.com/in/qiuyutang/) is a Ph.D. student in Computer Science and Engine
 ering at Lehigh University. Her research focuses on trustworthy AI\, media
  forensics\, and robust protection methods against diffusion-based image e
 diting and personalization systems. Her recent work explores concept prote
 ction\, style safeguarding\, semantic image manipulation\, and generative 
 AI robustness. She has contributed to multiple publications in computer vi
 sion and AI safety\, including research on diffusion model protection and 
 manipulation detection\, and has also served as a conference workshop orga
 nizer.\n\n**From Pixels to the Planet: Building Scalable and Grounded AI f
 or Science**\n\nAI has demonstrated a lot of new possibilities\, from draf
 ting emails to image editing and generation. The efficacy of AI models is 
 largely built upon a standard machine learning pipeline\, where data is fe
 d into models to get representations and predictions\, and the performance
  is evaluated with controlled benchmarks and metrics. However\, the mismat
 ch arises when we try to transit this pipeline to the interaction with the
  real world and use AI for scientific discovery. Beyond close-set decision
 s\, scientists want to discover new categories and propose new hypotheses.
  In this talk\, I will share how I address the challenges of AI for scienc
 e from the perspectives of data-centric methods and interpretability appro
 aches.\n\n*About the Speaker*\n\n[Jianyang Gu](https://www.linkedin.com/in
 /jianyang-gu-9235151b9/) is a postdoctoral scholar at The Ohio State Unive
 rsity. His research focuses on using data-centric methods to build scalabl
 e and interpretable foundation models for science.\n\n**Beyond the Barn: N
 on-Invasive Acidosis Detection in Dairy Cattle Through Multimodal Gas Emis
 sion Intelligence**\n\nRumen acidosis silently costs the global dairy indu
 stry billions annually and compromises animal welfare\, yet current detect
 ion methods remain invasive\, delayed\, and impractical at scale. Our lab 
 has pioneered a fundamentally new approach: capturing and analyzing exhale
 d CO₂ and CH₄ gas emission patterns through synchronized RGB-thermal i
 maging\, turning every breath into a diagnostic signal. We developed DualG
 asNet\, a dual-stream deep learning architecture with cross-attention fusi
 on that detects acidosis non-invasively and in real time\, achieving state
 -of-the-art accuracy on a first-of-its-kind livestock gas emission dataset
  we constructed from scratch.\n\nTo push toward explainable\, farm-ready A
 I\, we integrate vision-language models — CLIP and LLaVA-1.5 — enablin
 g zero-shot diagnostic reasoning that bridges the gap between deep learnin
 g predictions and actionable veterinary insight. This talk will walk throu
 gh the full pipeline from custom dataset creation to multimodal fusion to 
 VLM-powered interpretation\, offering the audience a compelling case study
  in how computer vision can solve high-impact\, real-world problems outsid
 e traditional benchmarks.\n\n*About the Speaker*\n\n[Taminul Islam](https:
 //www.linkedin.com/in/taminul-islam/) is a Doctoral Research Fellow and Ph
 D candidate at Southern Illinois University Carbondale with 40+ publicatio
 ns\, 740+ citations\, and an h-index of 16 — with publications in CVPR 2
 026\, WACV 2026 (Oral)\, ICCV 2025\, and Nature Scientific Reports\, inclu
 ding a Highly Cited Paper for 2024–25.\n\n**Building Real-World Computer
  Vision Systems with Voxel51**\n\nThis talk will explore practical workflo
 ws for building\, evaluating\, and improving modern computer vision system
 s. We’ll dive into real-world approaches to dataset curation\, model ana
 lysis\, multimodal AI workflows\, and production-ready vision pipelines us
 ing open-source technologies.\n\nThe session is designed for engineers\, r
 esearchers\, and AI practitioners looking to better understand how teams a
 re developing and scaling computer vision applications today. Expect pract
 ical demos\, technical insights\, and discussions around the evolving AI t
 ooling ecosystem.\n\n*About the Speaker*\n\n[Daniel Gural](https://www.lin
 kedin.com/in/daniel-gural/) is an expert in Physical AI and has been worki
 ng in the field for over 8 years. Working across healthcare he has experie
 nce in both operating use case as well as using Visual AI as an aid in psy
 chology applications as well.
URL;VALUE=URI:https://www.meetup.com/rome-ai-machine-learning-and-computer
 -vision-meetup/events/314999424/
STATUS:CONFIRMED
CREATED:20260528T185051Z
LAST-MODIFIED:20260528T185051Z
CLASS:PUBLIC
END:VEVENT
BEGIN:VEVENT
UID:event_314739577@meetup.com
SEQUENCE:1
DTSTAMP:20260729T190735Z
DTSTART;TZID=Europe/Rome:20260902T180000
DTEND;TZID=Europe/Rome:20260902T200000
SUMMARY:Sept 2 - Document Visual AI Workshop
DESCRIPTION:Rome AI\, Machine Learning and Computer Vision Meetup\nIn this
  hands-on workshop\, you'll use [FiftyOne](https://docs.voxel51.com/) and 
 the High Quality Invoice Images for OCR dataset to run the full data-centr
 ic loop end-to-end: embed invoices with a modern visual document model\, c
 luster them by structure\, run LightOnOCR as your base model\, and use per
 -sample evaluation scores layered onto embedding space to find \\*where\\*
  and \\*why\\* it fails.\n\n**Time\, Date and Location**\n\nSep 02\, 2026\
 n9:00 AM - 11:00 AM PST\nOnline. **[Register for the Zoom!](https://voxel5
 1.com/events/document-visual-ai-workshop-september-2-2026)**\n\n**What You
 'll Walk Away With**\n\n* A working FiftyOne pipeline for any document col
 lection you own\n* A repeatable curation query that combines evaluation + 
 embedding signals\n* A fine-tuned LightOnOCR checkpoint that demonstrably 
 outperforms the base model on your invoices\n* The mental model that data 
 curation — not architecture or hyperparameters — is the highest-levera
 ge thing you can do to improve a document AI system
URL;VALUE=URI:https://www.meetup.com/rome-ai-machine-learning-and-computer
 -vision-meetup/events/314739577/
STATUS:CONFIRMED
CREATED:20260511T191801Z
LAST-MODIFIED:20260511T191801Z
CLASS:PUBLIC
END:VEVENT
BEGIN:VEVENT
UID:event_315388191@meetup.com
SEQUENCE:1
DTSTAMP:20260729T190735Z
DTSTART;TZID=Europe/Rome:20260924T180000
DTEND;TZID=Europe/Rome:20260924T200000
SUMMARY:Sept 24 - AI\, ML and Computer Vision Meetup
DESCRIPTION:Rome AI\, Machine Learning and Computer Vision Meetup\nJoin ou
 r virtual meetup on September 24 to hear talks from experts on cutting-edg
 e topics across AI\, ML\, and computer vision.\n\n**Date\, Time and Locati
 on**\n\nSep 24\, 2026\n9:00 AM - 11:00 AM PST\nOnline. **[Register for the
  Zoom!](https://voxel51.com/events/ai-ml-and-computer-vision-meetup-septem
 ber-24-2026)**\n\n**How Do Mercedes-Benz AI Principles Drive our Innovatio
 n?**\n\nAt Mercedes-Benz\, our AI Principles guide every step of innovatio
 n\, emphasizing responsible use\, safety and reliability\, explainability\
 , and the protection of privacy. These principles go beyond statements and
  actively shape how we design\, test\, and deploy AI systems in real-world
  automotive and enterprise settings. In this talk\, I will present how the
 se principles inspired our recent research on when reusing LoRA (Low-Rank 
 Adaptation) is effective. By combining theoretical analysis with synthetic
  data as a proxy for enterprise scenarios\, we uncovered the strengths and
  limitations of modular AI components under constrained data access. Our f
 indings provide practical guidance on when reused LoRAs could deliver high
 -quality results.\n\n*About the Speaker*\n\n[Mei-Yen Chen](https://www.lin
 kedin.com/in/mei-yen-chen-22937787) is a Senior Data Scientist at Mercedes
 -Benz Tech Innovation GmbH in Germany with 10 years of industry experience
  in AI and data solutions. She leads early-stage AI projects across busine
 ss functions and collaborates with research institutions on machine learni
 ng and responsible AI.\n\n**Region Tokens as the Visual Primitive: From Re
 cognition to World Modeling**\n\nPatch-based tokenization has become the d
 efault interface between vision encoders and downstream models\, yet patch
 es carry no semantic structure and scale poorly with resolution and tempor
 al extent. This talk presents a research program centered on replacing pat
 ch tokens with region-level representations — semantically dense tokens 
 grounded in visual entities rather than arbitrary grid crops.\n\nI will de
 scribe RELOCATE\, REN\, and T-REN\, a progression of methods that produce 
 region tokens via pooling\, train them with region-level objectives\, and 
 extend them to video with temporal coherence. I will then present ongoing 
 work integrating region tokens into VLMs to directly expand visual context
  capacity\, and preliminary results on future region trajectory prediction
  as a foundation for world modeling.\n\nThe broader thesis is that region-
 level tokens are a more natural unit of visual computation than patches\, 
 and their advantage compounds as task complexity\, resolution\, and tempor
 al horizon increase.\n\n*About the Speaker*\n\n[Savya Khosla](https://www.
 linkedin.com/in/savyakhosla/) is a second-year Ph.D. student at the Univer
 sity of Illinois Urbana-Champaign\, advised by Prof. Derek Hoiem and Prof.
  Alex Schwing.\n\n**Leveraging Text-To-Image Diffusion Models for Consiste
 nt Set-to-Set Generation**\n\nImage collections are humans' primary way of
  capturing the world\, yet advances in generative editing remain largely i
 napplicable to this modality. We address this gap by introducing Match-and
 -Fuse - a zero-shot\, training-free method for consistent set-to-set gener
 ation from image collections that share a common visual element but differ
  in viewpoint\, capture time\, and surrounding content.\nOur key idea is a
  unified graph-based framework that combines dense correspondences with an
  emergent prior in text-to-image diffusion models to generate coherent can
 vases. We achieve state-of-the-art consistency and visual quality\, and un
 lock new creative capabilities for content generation.\n\n*About the Speak
 er*\n\n[Kate Feingold](https://www.linkedin.com/in/katefeingold/) is a PhD
  student in Computer Vision at the Weizmann Institute of Science. Her rese
 arch sits at the intersection of generative models\, 3D/4D perception\, an
 d multimodal learning\, focusing on problems where vision meets other moda
 lities or paradigms in creative tasks.\n\n**Yield Estimation of a Coffee i
 n a dense environment**\n\nThis presentation provides a detailed workflow 
 related to coffee yield estimation in a dense environment. With photos of 
 pre-harvest coffee plants from a couple of coffee estates\, details relate
 d to pre-processing\, annotation to detect regions of interest (ROI)\, obj
 ect detection training and inferencing results with various Yolo models an
 d finally segmentation with SAM2 and Yolo\\*-seg with training and inferen
 ce results to determine the count of raw\, pre-mature\, mature and over-ma
 ture coffee berries and finally the yield of the entire estate. All this i
 s based on real world data captured on iPhone and android phones.\n\n*Abou
 t the Speaker*\n\n[Raghu M. Rao](http://www.linkedin.com/in/raghumrao) is 
 a consultant working on applications of computer vision AI models. He was 
 previously with AMD and Xilinx. He has a Ph.D. in Wireless Communications 
 from UCLA and is a Senior Member\, IEEE. His current interests are in appl
 ications of AI for agriculture\, health care and wireless communications.
URL;VALUE=URI:https://www.meetup.com/rome-ai-machine-learning-and-computer
 -vision-meetup/events/315388191/
STATUS:CONFIRMED
CREATED:20260623T210742Z
LAST-MODIFIED:20260623T210742Z
CLASS:PUBLIC
END:VEVENT
X-ORIGINAL-URL:https://www.meetup.com/rome-ai-machine-learning-and-compute
 r-vision-meetup/events/ical/
X-WR-CALNAME:Rome AI\, Machine Learning and Computer Vision Meetup
END:VCALENDAR