Aug 22 - Visual Agent Workshop Part 2: From Pixels to Predictions


Details
Welcome to the three part Visual Agents Workshop virtual series...your hands on opportunity to learn about visual agents - how they work, how to develop them and how to fine-tune them.
Date and Time
Aug 22, 2025 at 9 AM Pacific
Part 2: From Pixels to Predictions - Building Your GUI Dataset
Hands-On Dataset Creation and Curation with FiftyOne
The best GUI models are only as good as their training data, and the best datasets are built by understanding what makes GUI interactions fundamentally different from natural images. In this practical session, you'll build a complete GUI dataset from scratch, learning to capture the precise annotations that GUI agents need.
Using FiftyOne as your data management backbone, you'll import diverse GUI screenshots, explore annotation strategies that go beyond bounding boxes, and implement efficient labeling workflows. We'll tackle the real challenges: handling platform differences, managing annotation quality, and creating datasets that transfer to new domains. You'll also learn advanced techniques like synthetic data generation and automated prelabeling to scale your annotation efforts.
Walk away with a production-ready dataset and the skills to build more—because in GUI agents, data quality determines everything.
By the end, you'll have both a dataset and the methodology to build the next generation of GUI training data.
About the Instructor
Harpreet Sahota is a hacker-in-residence and machine learning engineer with a passion for deep learning and generative AI. He’s got a deep interest in RAG, Agents, and Multimodal AI.

Aug 22 - Visual Agent Workshop Part 2: From Pixels to Predictions