Skip to content

Details

!We are starting one hour later!

Miton AI Times and evals.cz have joined forces to prepare the next event.

This time it will be about evaluation, evals.cz is a Prague meetup for people who build AI products and need to measure them.

Model benchmarks are everywhere, but what you care about is different: does it work for you, on your data, for your users?

Evelina Gabasova: The Weird and Wonderful World of Cybersecurity Evals
AI cybersecurity evals have suddenly become a mainstream concern, and the OpenAI/Hugging Face incident shows benchmarks are now adversarial environments. In this talk, we take a practical view of evaluating models for vulnerability discovery and remediation. Using examples from AISLE’s work, we’ll discuss the problems with public benchmarks, challenges in building our own, and failure modes when the eval itself is the attack surface.
Antonín Hoskovec: Evaluating Open Models on Your Task and in Your Language
Every new open-weight release raises the same questions: is it better on my task, is it better in my language, and could a smaller model do the job? I'll show how I answer them, using Czech as the example. For instance I use calibrated LLM judges on task-specific data to find the smallest model that is still sufficient, and I evaluate models on BenCzechMark and a Czech localisation of τ²-bench for agentic tool use. The whole pipeline can run in GitHub Actions, so a model swap, prompt change, or inference-stack update is tested like any other code change and regressions are caught before deployment.
Michal Spiegel: Generating Synthetic Evaluation Data for Search Agents
Evaluating search agents is hard, manual annotation is expensive and often infeasible. Synthetic generation scales, but it is easy to produce questions that no real user would ask. This talk describes how we generate synthetic evaluation data for search agents that answer questions over a large set of case files. Humans stay in the loop throughout: they steer the generation and review the results manually. The talk covers how we generate deep research QA data, as well as how we measure and control their quality and diversity.

-------------------------------------------------------------

⌚️ Start: 17:00 and end at 21:00, both online and offline in Truhlárna Karlín (Šaldova 388/5) - Onsite attendance is limited to 70 people.
🎙️ Speakers: Evelina Gabasova (AISLE), Tonda Hoskovec (Miton, anex.sh), Michal Spiegel (Filevine).
🍻 Networking after the seminar – great food and cold beer waiting for you!
🎥 Recording: After the event we will publish a recording and post a link to it in the comments.
🚪Doors open at 16:45, and the event officially starts at 17:00.
Your expertise is about to take off. Can't wait to have you on board!
-------------------------------------------------------------
Who is hosting the event
Miton is a Czech VC with portfolio companies like Rossum, Equilibre, or Rohlik. Apart from supporting Miton AI Times, Miton also issues the bi-weekly AI Newsletter.
Read more about Miton and AI.
**Evals.cz**
A Prague meetup and learning platform for people building AI-powered products.

Related topics

Events in Prague, CZ
Artificial Intelligence
Artificial Intelligence Applications
Big Data
Startup Businesses

You may also like