Skip to content

Details

Details
Benchmarks Are Lying to You: How to Tell Whether Models and Agents Actually Work

Description
AI benchmarks are everywhere, but what do they actually measure, and how much should we trust them?

In this talk, we’ll break down the major benchmarks used to evaluate language models, coding systems, multimodal models, and autonomous agents, from MMLU and GPQA to SWE bench, GAIA, and OSWorld.

We’ll explore common pitfalls such as benchmark contamination, judge bias, agent scaffolding, hidden test-time compute, and misleading leaderboard comparisons, then discuss how to design practical evaluations for real world AI products.

Attendees will leave with a clearer framework for interpreting benchmark results, comparing models fairly, and measuring whether an AI system actually works in production.

Speaker
Rakshak Talwar

Info
Austin Deep Learning Journal Club is group for committed machine learning practitioners and researchers alike. The group typically meets every first Tuesday of each month to discuss research publications. The publications are usually the ones that laid foundation to ML/DL or explore novel promising ideas and are selected by a vote. Participants are expected to read the publications to be able to contribute to discussion and learn from others. This is also a great opportunity to showcase your implementations to get feedback from other experts.

Sponsors:
STATION Austin is the center of gravity for entrepreneurs in Texas. Day and night, in-person and online, we gather the best founders, programmers, and designers outside of Silicon Valley and introduce them to investors, employees, and customers who help their ideas launch. STATION Austin is powered by Capital Factory, whose investments and leadership have helped power the Texas startup ecosystem for more than a decade. To sign up for a STATION Austin membership, click here.

Related topics

You may also like