Skip to content

Details

An agent may produce the right final answer while choosing the wrong tool, taking 20 unnecessary steps, or making unauthorized calls. If we only look at the final output, the run may appear perfectโ€”but in terms of reliability, efficiency, and permission boundaries, serious problems may still exist.
As AI applications move into production, teams need to understand:

  • ๐Ÿ” When is an application truly reliable?
  • ๐Ÿงช Where is it likely to fail?
  • ๐Ÿ“ˆ Does each update genuinely improve quality?
  • ๐Ÿ› ๏ธ How should failures be diagnosed?
  • โš–๏ธ Can we trust an LLM to evaluate another AI system?

In this practical session, Jessie, a former Senior AI Engineer at Deloitte, will use a real agent execution trace to explore evaluation design, failure diagnosis, LLM-as-a-judge calibration, and production implementation.
You will learn how to build an AI evaluation system that is measurable, diagnosable, and continuously improvable.

## Event Details

  • ๐Ÿ“… Date: Monday, 21 September 2026
  • ๐Ÿ•’ Time: 7:00โ€“8:00 PM Sydney time
  • ๐Ÿ’ป Format: Online livestream
  • ๐ŸŽค Speaker: Jessie โ€” Former Senior AI Engineer at Deloitte
  • ๐Ÿ’ฌ Q&A: 10-minute interactive Q&A
  • ๐ŸŽŸ๏ธ Cost: Free

## Why Does Evaluation Matter? ๐Ÿ”

Generating a reasonable-looking answer is often only the beginning. In real business workflows, we also need to ask:

  • ๐Ÿ”„ Can the application handle different versions of the same task?
  • ๐Ÿงฐ Were the tools and parameters used correctly?
  • ๐Ÿงฉ Did a component failure affect the overall result?
  • ๐Ÿง  Can we trust an AI judgeโ€™s evaluation?
  • ๐Ÿš€ Do pre-launch tests reflect real user needs?

Evaluation turns these questions into clear test cases, quality criteria, and feedback signalsโ€”helping teams compare versions, identify problems, and decide what to improve next.

## What Weโ€™ll Cover ๐Ÿงฉ

### 1. A โ€œCorrect but Failedโ€ Agent Run

Explore a real execution trace where the agent gives the correct answer but still demonstrates tool misuse, unnecessary steps, and unauthorized calls.

### 2. The Core Components of Eval

Understand Dataset, Scorer, and Target, as well as the roles of offline and online evaluation.

### 3. Three Levels of Agent Evaluation

Evaluate agents through:

  • โœ… Final results
  • ๐Ÿ›ค๏ธ Execution trajectories
  • โš™๏ธ Individual components

### 4. LLM-as-a-Judge

Learn how to identify judge bias, build calibration datasets, and validate AI evaluators through human review.

### 5. Build Your First Evaluation Workflow

Start with 10โ€“20 critical cases, integrate evaluation into CI, and continuously add real production failures.

### 6. Who Defines โ€œGoodโ€? ๐ŸŽฏ

Understand how domain experts and engineering teams can work together to define standards and maintain evaluation quality.

## What Youโ€™ll Take Away ๐Ÿ’ก

  • ๐Ÿ“ A practical framework for measurable AI quality
  • ๐Ÿฉบ A structured approach to diagnosing failures
  • โš–๏ธ Strategies for calibrating LLM judges
  • ๐Ÿ”ง A practical path to integrating evaluation into CI and production
  • ๐Ÿค Clearer collaboration between domain experts and engineering teams

## Who Should Attend? ๐Ÿ‘ฅ

This session is ideal for:

  • ๐Ÿ‘จโ€๐Ÿ’ป Engineers and technical leads building AI applications, RAG systems, or agents
  • ๐Ÿ“ฆ Product managers responsible for AI product quality and launch readiness
  • ๐Ÿง‘โ€๐Ÿ”ฌ Domain experts converting business knowledge into evaluation criteria
  • ๐Ÿง‘โ€๐Ÿ’ป Developers who want to move from an AI demo to a systematic evaluation process

## Q&A and Discussion ๐Ÿ’ฌ

Bring your questions about test-case selection, tool-call evaluation, inconsistent judge scores, and how to implement evaluation for an existing AI application.
๐ŸŒŸ Join us online and take the first step toward building more reliable AI applications.
Note: The original event page lists both 7:00โ€“8:00 PM and 9:00โ€“10:00 PM Sydney time. Please confirm the correct time before publishing on Luma.

Related topics

Career Coaching
Career Network
Job Search
Professional Development
Education & Technology

You may also like