Phase 8: Hands-On Job-Essential Lab Solutions
This module contains complete production manifests, Python evaluation scripts, and step-by-step verification commands for the Phase 8 LLMOps Job-Essential Exercises.
π οΈ Lab 1: Deploying a Distributed Vector Database (Milvus on Kubernetes)
Objective
Deploy a distributed, production-grade Milvus vector database cluster on Kubernetes using Helm, configured with MinIO object storage, etcd coordination, and NVMe-backed persistent volume storage.
1. Production Milvus Helm Values (milvus-values.yaml)
cluster:
enabled: true
# Coordinators (Root, Data, Query, Index)
rootCoord:
replicas: 1
dataCoord:
replicas: 1
queryCoord:
replicas: 1
indexCoord:
replicas: 1
# Stateless Data & Query Nodes (Scale horizontally)
queryNode:
replicas: 2
resources:
limits:
cpu: "4000m"
memory: "16Gi"
requests:
cpu: "2000m"
memory: "8Gi"
dataNode:
replicas: 2
# External High-Availability Storage Dependencies
minio:
enabled: true
mode: distributed
replicas: 4
persistence:
size: 100Gi
storageClass: gp3 # AWS EBS gp3
etcd:
replicaCount: 3
persistence:
size: 20Gi
2. Deployment & Cluster Health Verification
# Add Milvus Helm repo
helm repo add milvus https://zilliztech.github.io/milvus-helm/
helm repo update
# Install Milvus distributed cluster
helm install my-milvus milvus/milvus -n llmops --create-namespace -f milvus-values.yaml
# Verify all coordinator, query, and data pods are running
kubectl get pods -n llmops -l app.kubernetes.io/instance=my-milvus
π οΈ Lab 2: Enterprise AI Gateway (LiteLLM with Multi-Model Fallback)
Objective
Deploy LiteLLM proxy as a centralized AI Gateway inside Kubernetes. Configure a primary route to OpenAI (gpt-4o) with an automated fallback to AWS Bedrock (anthropic.claude-3-5-sonnet) when OpenAI returns a 429 Too Many Requests or 5xx Server Error.
1. LiteLLM Kubernetes ConfigMap & Deployment
apiVersion: v1
kind: ConfigMap
metadata:
name: litellm-config
namespace: llmops
data:
config.yaml: |
model_list:
- model_name: enterprise-chat
litellm_params:
model: openai/gpt-4o
api_key: os.environ/OPENAI_API_KEY
rpm: 500
- model_name: enterprise-chat-fallback
litellm_params:
model: bedrock/anthropic.claude-3-5-sonnet-20240620-v1:0
aws_region_name: us-east-1
router_settings:
routing_strategy: latency-based-routing
fallbacks:
- enterprise-chat: ["enterprise-chat-fallback"]
timeout: 5
num_retries: 2
2. Verify Fallback Execution via curl
# Query LiteLLM Gateway
curl -X POST http://litellm.llmops.svc:4000/v1/chat/completions \
-H "Authorization: Bearer sk-corp-virtual-team-key" \
-H "Content-Type: application/json" \
-d '{
"model": "enterprise-chat",
"messages": [{"role": "user", "content": "What is enterprise platform engineering?"}]
}'
π οΈ Lab 3: Automated Prompt Evaluation Pipeline (MLflow)
Objective
Build a Python CI pipeline using MLflow that evaluates a RAG prompt against a golden test dataset. Compute faithfulness and relevance scores using an LLM-as-a-Judge, and fail the CI pipeline if accuracy drops below 90%.
Production Evaluation Script (eval_prompt.py)
import sys
import mlflow
import pandas as pd
from mlflow.metrics.genai import faithfulness, answer_relevance
# 1. Connect to MLflow Tracking Server
mlflow.set_tracking_uri("http://mlflow.corp.internal:5000")
mlflow.set_experiment("rag-prompt-ci-evaluations")
# 2. Golden QA Benchmark Dataset
test_df = pd.DataFrame({
"inputs": [
"What is the company password rotation policy?",
"How do I request an AWS sandbox account?"
],
"context": [
"Passwords must be at least 16 characters and rotated every 90 days via Okta.",
"Engineers can request an ephemeral AWS sandbox account via Backstage in 60 seconds."
],
"ground_truth": [
"Passwords must be at least 16 characters and rotated every 90 days.",
"Request an AWS sandbox via Backstage."
]
})
with mlflow.start_run(run_name="pr-402-strict-eval"):
# Run automated evaluation
eval_results = mlflow.evaluate(
data=test_df,
targets="ground_truth",
model_type="question-answering",
extra_metrics=[
faithfulness(model="openai:/gpt-4o"),
answer_relevance(model="openai:/gpt-4o")
]
)
faith_score = eval_results.metrics.get("faithfulness/v1/score", 0.0)
relevance_score = eval_results.metrics.get("answer_relevance/v1/score", 0.0)
print(f"Evaluation Metrics: Faithfulness={faith_score:.2f}, Relevance={relevance_score:.2f}")
# Enforce CI Quality Gate
if faith_score < 0.90 or relevance_score < 0.90:
print("β CI Quality Gate FAILED: Scores below 0.90 threshold!")
sys.exit(1)
else:
print("β
CI Quality Gate PASSED: Prompt ready for production deployment!")
sys.exit(0)