Automated Agent Testing with Synthetic Datasets | Galileo Experiments
Learn how to systematically test AI agents using Galileo's datasets and experiments platform. This demo shows how to generate synthetic test data, run controlled experiments, and catch regressions before productionβwith built-in CI/CD integration for automated quality gates. In this demo, our Field Engineer, Al Chen, will walk you through: - Creating synthetic datasets with 6 user behavior profiles - Auto-generating 50 test cases from few-shot examples - Running systematic experiments with out-of-the-box metrics - Filtering and analyzing results to identify failure patterns - Integrating experiments into CI/CD pipelines - Comparing experiments across different datasets and configurations Built for field service agents helping customers troubleshoot home appliances. Demonstrates how teams can stress-test applications against toxic inputs, off-topic queries, and edge cases before production deployment. π Try Galileo: https://app.galileo.ai/sign-up?utm_medium=organic&utm_source=youtube π Docs: https://v2docs.galileo.ai/ Perfect for AI engineering teams building production agents that need comprehensive testing coverage without manual test case creation. 0:00 - Systematic Experiments for Agent Testing 0:30 - Spotting Errors & Regressions Before Production 1:00 - CI/CD Pipeline Integration with Quality Gates 1:30 - Use Case: Field Service Technician Agent 2:00 - Creating Synthetic Datasets in Console 2:30 - Auto-Generate Test Dataset Feature 3:00 - Configuring User Behavior Profiles 3:30 - Generating 50 Varied Test Cases 4:00 - Agent System Prompt Configuration 4:30 - Running First Experiment with Built-In Metrics 5:00 - Experiment Results in Console 5:30 - Input Toxicity Metric Analysis 6:00 - Comparing Multiple Experiments