Skip to content

Phase 8: Hands-On Job-Essential Lab Solutions

This module contains complete production manifests, Python evaluation scripts, and step-by-step verification commands for the Phase 8 LLMOps Job-Essential Exercises.


πŸ› οΈ Lab 1: Deploying a Distributed Vector Database (Milvus on Kubernetes)

Objective

Deploy a distributed, production-grade Milvus vector database cluster on Kubernetes using Helm, configured with MinIO object storage, etcd coordination, and NVMe-backed persistent volume storage.

1. Production Milvus Helm Values (milvus-values.yaml)

cluster:
  enabled: true

# Coordinators (Root, Data, Query, Index)
rootCoord:
  replicas: 1
dataCoord:
  replicas: 1
queryCoord:
  replicas: 1
indexCoord:
  replicas: 1

# Stateless Data & Query Nodes (Scale horizontally)
queryNode:
  replicas: 2
  resources:
    limits:
      cpu: "4000m"
      memory: "16Gi"
    requests:
      cpu: "2000m"
      memory: "8Gi"

dataNode:
  replicas: 2

# External High-Availability Storage Dependencies
minio:
  enabled: true
  mode: distributed
  replicas: 4
  persistence:
    size: 100Gi
    storageClass: gp3 # AWS EBS gp3

etcd:
  replicaCount: 3
  persistence:
    size: 20Gi

2. Deployment & Cluster Health Verification

# Add Milvus Helm repo
helm repo add milvus https://zilliztech.github.io/milvus-helm/
helm repo update

# Install Milvus distributed cluster
helm install my-milvus milvus/milvus -n llmops --create-namespace -f milvus-values.yaml

# Verify all coordinator, query, and data pods are running
kubectl get pods -n llmops -l app.kubernetes.io/instance=my-milvus

πŸ› οΈ Lab 2: Enterprise AI Gateway (LiteLLM with Multi-Model Fallback)

Objective

Deploy LiteLLM proxy as a centralized AI Gateway inside Kubernetes. Configure a primary route to OpenAI (gpt-4o) with an automated fallback to AWS Bedrock (anthropic.claude-3-5-sonnet) when OpenAI returns a 429 Too Many Requests or 5xx Server Error.

1. LiteLLM Kubernetes ConfigMap & Deployment

apiVersion: v1
kind: ConfigMap
metadata:
  name: litellm-config
  namespace: llmops
data:
  config.yaml: |
    model_list:
      - model_name: enterprise-chat
        litellm_params:
          model: openai/gpt-4o
          api_key: os.environ/OPENAI_API_KEY
          rpm: 500

      - model_name: enterprise-chat-fallback
        litellm_params:
          model: bedrock/anthropic.claude-3-5-sonnet-20240620-v1:0
          aws_region_name: us-east-1

    router_settings:
      routing_strategy: latency-based-routing
      fallbacks:
        - enterprise-chat: ["enterprise-chat-fallback"]
      timeout: 5
      num_retries: 2

2. Verify Fallback Execution via curl

# Query LiteLLM Gateway
curl -X POST http://litellm.llmops.svc:4000/v1/chat/completions \
  -H "Authorization: Bearer sk-corp-virtual-team-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "enterprise-chat",
    "messages": [{"role": "user", "content": "What is enterprise platform engineering?"}]
  }'
(LiteLLM automatically handles token quota tracking and seamless model failover).


πŸ› οΈ Lab 3: Automated Prompt Evaluation Pipeline (MLflow)

Objective

Build a Python CI pipeline using MLflow that evaluates a RAG prompt against a golden test dataset. Compute faithfulness and relevance scores using an LLM-as-a-Judge, and fail the CI pipeline if accuracy drops below 90%.

Production Evaluation Script (eval_prompt.py)

import sys
import mlflow
import pandas as pd
from mlflow.metrics.genai import faithfulness, answer_relevance

# 1. Connect to MLflow Tracking Server
mlflow.set_tracking_uri("http://mlflow.corp.internal:5000")
mlflow.set_experiment("rag-prompt-ci-evaluations")

# 2. Golden QA Benchmark Dataset
test_df = pd.DataFrame({
    "inputs": [
        "What is the company password rotation policy?",
        "How do I request an AWS sandbox account?"
    ],
    "context": [
        "Passwords must be at least 16 characters and rotated every 90 days via Okta.",
        "Engineers can request an ephemeral AWS sandbox account via Backstage in 60 seconds."
    ],
    "ground_truth": [
        "Passwords must be at least 16 characters and rotated every 90 days.",
        "Request an AWS sandbox via Backstage."
    ]
})

with mlflow.start_run(run_name="pr-402-strict-eval"):
    # Run automated evaluation
    eval_results = mlflow.evaluate(
        data=test_df,
        targets="ground_truth",
        model_type="question-answering",
        extra_metrics=[
            faithfulness(model="openai:/gpt-4o"),
            answer_relevance(model="openai:/gpt-4o")
        ]
    )

    faith_score = eval_results.metrics.get("faithfulness/v1/score", 0.0)
    relevance_score = eval_results.metrics.get("answer_relevance/v1/score", 0.0)

    print(f"Evaluation Metrics: Faithfulness={faith_score:.2f}, Relevance={relevance_score:.2f}")

    # Enforce CI Quality Gate
    if faith_score < 0.90 or relevance_score < 0.90:
        print("❌ CI Quality Gate FAILED: Scores below 0.90 threshold!")
        sys.exit(1)
    else:
        print("βœ… CI Quality Gate PASSED: Prompt ready for production deployment!")
        sys.exit(0)