Skip to content

Lesson 2: LLM Evaluation & MLflow Tracking

🧠 The Concept (Explain Like I'm 5)

If you change the instructions (Prompt) you give to an AI, how do you mathematically prove the AI got smarter or dumber? MLflow is a giant spreadsheet that tracks every prompt you tested, what model you used, and the score the AI achieved, so you can pick the best combination for production.


🏢 The Enterprise Context

  • Prompt Engineering Lifecycle: Prompts are treated as code. They must be versioned, tested against datasets, and tracked.
  • LLM-as-a-Judge: Since AI answers are text, not math, enterprises use a larger AI (like GPT-4) to grade the answers of a smaller AI (like Llama-3) automatically during the CI/CD pipeline.

🗺️ Visual Architecture: Prompt CI/CD

flowchart LR
    Dev["Developer<br/>Updates Prompt"] --> Test["Automated Test Suite"]
    Test -->|Sends Query| LLM["LLM (App)"]
    LLM -->|Returns Answer| Judge["GPT-4 Judge"]
    Judge -->|Scores Answer| MLflow["MLflow (Experiment Tracker)"]

    MLflow -->|If Score > 90%| Deploy["Deploy Prompt to Prod"]