Webinar
60 Mins
Testing & Evaluation in the Agentic Age
Build Trust in AI Agents with Golden Datasets, Silent Evaluation, and Governance
Overview
As organizations move from AI-assisted experiences to increasingly autonomous agentic systems, one challenge is becoming clear: building an AI agent is only half the problem. Proving it can be trusted in production is where the real work begins.
Traditional pass/fail testing was designed for deterministic applications. Agentic systems operate differently. The same input can produce multiple plausible outcomes, making quality assurance less about validating a single result and more about continuously evaluating performance, gathering evidence, and monitoring behavior over time.
Join Florian Lauck-Wunderlich and Glen Aronson for a practical AI Expert Circle session exploring how organizations can design, validate, monitor, and govern GenAI and agentic solutions with confidence.
Event Time: 09:00 AM ET | 2:00 PM UK (BST) | 3:00 PM CET | 6:30 PM IST | 09:00 PM SGT
Through real-world examples and demonstrations, we'll explore how to create golden reference datasets, compare prompt and agent variants, apply multi-signal evaluation techniques, and use silent evaluation to gather production evidence without operational risk. We'll also discuss how human feedback, SME annotations, technical metrics, process outcomes, and business signals can be combined to support rollout decisions and governance.
During this session, we'll cover:
- Why testing GenAI and agentic systems requires a new paradigm
- Building golden reference datasets and evaluation baselines
- Comparing prompts, contexts, and agent variants using repeatable evaluation approaches
- Applying multi-signal evaluation frameworks that combine human, technical, process, and business outcomes
- Using silent evaluation and SME annotations to collect evidence before expanding autonomy
- Designing Human-in-the-Loop patterns for a safe progression from assisted to autonomous operation
- Governance and monitoring practices for production-ready AI agents
Whether you're designing agentic solutions, leading delivery teams, or defining AI governance standards, this session will provide practical techniques and reusable patterns for building trust in AI systems before and after go-live.
Reserve your place today and learn how leading teams are moving from testing prompts to governing agentic behavior with evidence, confidence, and control.
Presenters