Lesson 2: LLM Evaluation & MLflow Tracking
🧠 The Concept (Explain Like I'm 5)
If you change the instructions (Prompt) you give to an AI, how do you mathematically prove the AI got smarter or dumber? MLflow is a giant spreadsheet that tracks every prompt you tested, what model you used, and the score the AI achieved, so you can pick the best combination for production.
🏢 The Enterprise Context
- Prompt Engineering Lifecycle: Prompts are treated as code. They must be versioned, tested against datasets, and tracked.
- LLM-as-a-Judge: Since AI answers are text, not math, enterprises use a larger AI (like GPT-4) to grade the answers of a smaller AI (like Llama-3) automatically during the CI/CD pipeline.
🗺️ Visual Architecture: Prompt CI/CD
flowchart LR
Dev["Developer<br/>Updates Prompt"] --> Test["Automated Test Suite"]
Test -->|Sends Query| LLM["LLM (App)"]
LLM -->|Returns Answer| Judge["GPT-4 Judge"]
Judge -->|Scores Answer| MLflow["MLflow (Experiment Tracker)"]
MLflow -->|If Score > 90%| Deploy["Deploy Prompt to Prod"]