Designing Test Strategies for AI-From Prompt Variability to Safety Regression
Details
NOTE: the following description was created with the assistance of AI
AI products don’t behave like traditional software. One prompt can produce ten different answers, small wording changes can shift outcomes, and “correct” may depend on context, tone, and user intent. In this session, software testers will learn how to design practical, repeatable test strategies that keep AI features reliable—and safe—over time.
We’ll cover how to think about prompt variability (and how to turn it into testable scenarios), how to build test suites that catch failures beyond simple correctness, and how to measure stability across versions, models, and system changes. Then we’ll move into safety regression: creating guardrails, validating policy-aligned behavior, and preventing reintroductions of known issues when models or prompts evolve.
Come for real testing patterns you can apply immediately, leave with a clearer approach to coverage, risk, and regression testing for AI systems—so your next release doesn’t just work on paper, it behaves responsibly in the wild.
