Zum Inhalt springen

Details

Calibrating AI Judges and Building Reliable Production Metrics - Tejas Kumar

​A green dashboard does not always mean a healthy system. Engineering teams must build accurate evaluation frameworks when using language models to grade other models. A properly calibrated AI judge will prioritize factual answers over confident but incorrect responses. Resolving these biases ensures that rising metrics actually reflect an improving customer experience.

​In our upcoming episode featuring Tejas Kumar we will explore why standard AI evaluations fail and how engineering teams can fix them. The conversation will detail how to calibrate an AI judge against human labels and shift from relying on a single fuzzy score to using targeted deterministic probes. We will also discuss how to evaluate retrieval augmented generation systems stage by stage to pinpoint exact failure points in production.

​He’ll cover:

  • ​Overcoming length and position bias in AI judges
  • ​Calibrating AI judges using human labels and strict rubrics
  • ​Using instant deterministic checks instead of expensive model evaluations
  • ​How to evaluate RAG systems stage by stage
  • ​Pinpointing data retrieval versus text generation failures
  • ​The link between agent harnesses and successful evaluations

​About the Speaker:
​Tejas Kumar is an AI Engineer at IBM, based in Berlin, Germany. He is the best selling author of Fluent React (O’Reilly), the host of the ConTejas Code podcast, and has given 81 recorded talks at 65 events since 2012. His talk “Harnesses in AI: A Deep Dive” at AI Engineer Europe 2026 has over 275,000 views, and his talk on LLM evaluation, “Your Evals Are Lying to You”, was accepted at AI Engineer World’s Fair 2026. He is also an angel investor and advisor to early stage startups, 3 of the 6 he has backed since acquired, by CoreWeave, the Linux Foundation and Supabase.

**Join our Slack: https://datatalks.club/slack.html**

Das könnte dir auch gefallen