Evals in your SDLC. Eval Engineering for AI Developers , lesson 5 - learn how evals fit in your SDLC
Learn Eval Engineering in this free, 5-part, hands-on course presented by @jimbobbennett 90% of AI agents don't make it successfully to production. The biggest reason is the AI engineers building these apps don't have a clear way of evaluating that these agents are doing what they should do, and using the results of this evaluation to fix them. In this course, you will learn all about evals for AI applications. You'll start with some out-of-the-box metrics and learn about evals, then move onto understanding observability for AI apps, analyzing failure states, defining custom metrics, then finally using these across your whole SDLC. This will be hands on, so be prepared to write some code, create some metrics, and do some homework! In this final lesson, you will - Learn how evals fit into the SDLC - Build unit tests using evals that can be run in your CI/CD pipeline - Learn about using evals as guardrails at runtime - Add observability and alerts to detect when your application is failing Prerequisites: - A basic knowledge of Python - Access to an OpenAI API key - A free Galileo account (we will be using Galileo as the evals platform). Sign up at https://galileo.ai/sign-up. - Course materials from https://github.com/rungalileo/eval-engineering Catch the rest of the lessons here: https://youtube.com/playlist?list=PLS7keRo8770OODJn9HAEN8JKbIzTHME60 0:00:00 - Introduction & Welcome 0:04:05 - The Software Development Life Cycle (SDLC) Overview 0:14:43 - Eval Engineering in the Requirements Phase 0:22:31 - Eval Engineering in the Design Phase 0:32:06 - Eval Engineering in the Implementation Phase 0:35:18 - Demo: Running Experiments in Galileo 0:47:16 - Eval Engineering in the Testing Phase 0:50:52 - Continuous Iteration: Test, Fix, Improve 0:54:50 - Eval Engineering in Deployment & Production 1:00:50 - Guardrails & Governance 1:13:00 - Q&A: Handling Multimodal Evals (Voice/Video) 1:15:50 - Wrap Up & Final Thoughts