Operate LLM applications in production
LLMOps, RAG & Agents
Run LLM systems for real: the LLM application lifecycle, production RAG pipelines, reliable AI agents with guardrails and human approval, and end-to-end LLM observability.
What you'll learn
Module 1Module 0: Prerequisites and Environment Readiness
- Who This Course Is For
- Python, APIs, and Structured Data
- Transformer and Generation Fundamentals
- Embeddings and Retrieval Basics
- Kubernetes and Production-Service Basics
- Evaluation, Observability, and Security Baseline
- Readiness Exercise
- Knowledge Check
Module 2Module 1: LLMOps Foundations
- From MLOps to LLMOps
- The LLM Application Lifecycle
- Prompt and Configuration Management
- Evaluation-Driven Development
- Release Strategies and Model Routing
- Knowledge Check
Module 3Module 2: Production Retrieval-Augmented Generation
- RAG Architecture and When to Use It
- Ingestion, Parsing, and Chunking
- Embeddings, Search, and Reranking
- Context Assembly, Citations, and Security
- Evaluating and Operating RAG
- Knowledge Check
Module 4Module 3: Building Reliable AI Agents
- Agent Architecture and the Control Loop
- Tool Design and Structured Interfaces
- Memory and State
- Agent Safety and Human Approval
- Evaluating and Observing Agents
- Knowledge Check
Module 5Module 4: Operating the Complete System
- End-to-End Observability
- Reliability, Cost, and Fallbacks
- Security Testing and Red Teaming
- Capstone: The Support Resolution Agent
- Production Readiness and Continuous Improvement
- Knowledge Check
Preview the first lesson
This course is for platform engineers, MLOps engineers, application engineers, data scientists, and operators moving LLM applications from prototypes to dependable services. You do not need to train a foundation model, but you should be comfortable reading Python, using a terminal, and reasoning about web services.
Before continuing, review Git, containers, HTTP APIs, JSON, environment variables, and development versus production configuration. You should recognize Kubernetes pods, deployments, services, namespaces, configuration, secrets, resource requests, and logs. The practical goal is to follow a request across code and infrastructure and explain how to reproduce or roll back a change.
AI tutor on every page
Ask questions in context — the tutor reads the exact lesson you're on.
Hands-on terminal labs
Realistic scenario labs with a simulated cluster — no setup required.
Shareable certificate
Pass the Knowledge Checks to earn a verifiable certificate of completion.