Mastering GenAI Development with MLflow: AI Observability, Evaluation, Debugging, and Deployment

MLflow
1,905 views February 25, 2026

In this series, Jules Damji provides a technical walkthrough of the AI agent development lifecycle using MLflow. We move beyond basic prompting to cover the rigorous engineering required for production-grade agents, including instrumentation for observability, automated evaluation patterns like LLM-as-a-Judge, and systemic prompt optimization. These tutorials provide a roadmap from initial environment configuration to the deployment of a fully instrumented Retrieval-Augmented Generation (RAG) system. Follow the series chronologically or jump to specific modules to address architectural and knowledge gaps. Tutorial Roadmap 🔹 Module 1 & 2: Infrastructure & Tracking – Environment setup, credential management, and hierarchical run comparison (Parent/Child). 🔹 Module 3 & 4: AI Observability & Tracing – Implementing auto-tracing and manual decorators to monitor tool calls and latencies, and using MLflow Assistant for debugging and analysing root cause analysis 🔹 Module 5: Prompt Engineering – Versioning prompts via the Prompt Registry and optimization using the GAPA algorithm. 🔹 Module 6 & 7: Agent Evaluation & Integrated Frameworks – Scaling evaluation with the "LLM-as-a-Judge" pattern and integrating with LangChain/LlamaIndex. Final Project: Building and deploying an end-to-end instrumented RAG application.

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close