AI Evals For Engineers & PMs course homework 3 using Galileo - LLM-as-Judge for RecipeBot Evaluation
This video by @jimbobbennett has a walkthrough for the third homework from the AI Evals For Engineers & PMs course from Hamel Husain and Shreya Shankar. This walks you through: - Splitting pass and fail datasets for train, dev, and test - Creating a custom metric prompt, and setting up the metric in Galileo - Running experiments to calculate the TPR and TNR - Finding errors in datasets The Jupyter notebook for this is in the Galileo SDK examples repo - https://github.com/rungalileo/sdk-examples/tree/main/ai-evals-course/hw3 To get started with Galileo, and to integrate the Recipe Chatbot, check out our setup video - https://youtu.be/Lb_uxu8C5wc?si=FquikejiPHqKf6a_ You can find a full playlist with all the homework walkthroughs and Galileo setup here - https://youtube.com/playlist?list=PLS7keRo8770Ps8_gGRBzhkCEiiqCBH1P6&si=VXACui-rElE5HfMX Learn more about Galileo - https://galileo.ai/ Galileo docs - https://v2docs.galileo.ai/ Join the Galileo community - https://community.galileo.ai/