How do you know your AI actually works? Evals for real projects
Details
You built something with AI. How do you know it works? We're doing evals — the unglamorous part that separates a demo from something you'd ship. All levels.
New home: Schuler Books on 28th Street, in the Chapbook Cafe. Coffee, real tables, and quiet enough to hear each other.
WHAT WE'RE DIGGING INTO
How to tell whether a prompt or an agent got better or just different. Writing a small eval set you'll actually run. Using a model to grade output, and where that quietly misleads you. Catching regressions before your users do.
WHO THIS IS FOR
Anyone shipping something with an LLM in it, and anyone who's noticed their AI feature works great in the demo and badly in real life. No prior eval experience assumed — that's the point.
HOW THE EVENING GOES
6:30 — coffee, introductions, what everyone's building
6:50 — I walk through the eval set I use on Digi-Bot, including what it failed to catch
7:20 — open floor: bring a project and we'll sketch an eval for it
8:30 — whoever's still talking keeps talking
Free to attend, no cover. Coffee and books are on you. Schuler Books, 2660 28th St SE — big parking lot, easy to find.
