Hands-on AI Evaluation: From Eval Design to Production
Details
An agent may produce the right final answer while choosing the wrong tool, taking 20 unnecessary steps, or making unauthorized calls. If we only look at the final output, the run may appear perfectโbut in terms of reliability, efficiency, and permission boundaries, serious problems may still exist.
As AI applications move into production, teams need to understand:
- ๐ When is an application truly reliable?
- ๐งช Where is it likely to fail?
- ๐ Does each update genuinely improve quality?
- ๐ ๏ธ How should failures be diagnosed?
- โ๏ธ Can we trust an LLM to evaluate another AI system?
In this practical session, Jessie, a former Senior AI Engineer at Deloitte, will use a real agent execution trace to explore evaluation design, failure diagnosis, LLM-as-a-judge calibration, and production implementation.
You will learn how to build an AI evaluation system that is measurable, diagnosable, and continuously improvable.
## Event Details
- ๐ Date: Monday, 21 September 2026
- ๐ Time: 7:00โ8:00 PM Sydney time
- ๐ป Format: Online livestream
- ๐ค Speaker: Jessie โ Former Senior AI Engineer at Deloitte
- ๐ฌ Q&A: 10-minute interactive Q&A
- ๐๏ธ Cost: Free
## Why Does Evaluation Matter? ๐
Generating a reasonable-looking answer is often only the beginning. In real business workflows, we also need to ask:
- ๐ Can the application handle different versions of the same task?
- ๐งฐ Were the tools and parameters used correctly?
- ๐งฉ Did a component failure affect the overall result?
- ๐ง Can we trust an AI judgeโs evaluation?
- ๐ Do pre-launch tests reflect real user needs?
Evaluation turns these questions into clear test cases, quality criteria, and feedback signalsโhelping teams compare versions, identify problems, and decide what to improve next.
## What Weโll Cover ๐งฉ
### 1. A โCorrect but Failedโ Agent Run
Explore a real execution trace where the agent gives the correct answer but still demonstrates tool misuse, unnecessary steps, and unauthorized calls.
### 2. The Core Components of Eval
Understand Dataset, Scorer, and Target, as well as the roles of offline and online evaluation.
### 3. Three Levels of Agent Evaluation
Evaluate agents through:
- โ Final results
- ๐ค๏ธ Execution trajectories
- โ๏ธ Individual components
### 4. LLM-as-a-Judge
Learn how to identify judge bias, build calibration datasets, and validate AI evaluators through human review.
### 5. Build Your First Evaluation Workflow
Start with 10โ20 critical cases, integrate evaluation into CI, and continuously add real production failures.
### 6. Who Defines โGoodโ? ๐ฏ
Understand how domain experts and engineering teams can work together to define standards and maintain evaluation quality.
## What Youโll Take Away ๐ก
- ๐ A practical framework for measurable AI quality
- ๐ฉบ A structured approach to diagnosing failures
- โ๏ธ Strategies for calibrating LLM judges
- ๐ง A practical path to integrating evaluation into CI and production
- ๐ค Clearer collaboration between domain experts and engineering teams
## Who Should Attend? ๐ฅ
This session is ideal for:
- ๐จโ๐ป Engineers and technical leads building AI applications, RAG systems, or agents
- ๐ฆ Product managers responsible for AI product quality and launch readiness
- ๐งโ๐ฌ Domain experts converting business knowledge into evaluation criteria
- ๐งโ๐ป Developers who want to move from an AI demo to a systematic evaluation process
## Q&A and Discussion ๐ฌ
Bring your questions about test-case selection, tool-call evaluation, inconsistent judge scores, and how to implement evaluation for an existing AI application.
๐ Join us online and take the first step toward building more reliable AI applications.
Note: The original event page lists both 7:00โ8:00 PM and 9:00โ10:00 PM Sydney time. Please confirm the correct time before publishing on Luma.
