LLM Evaluation & Safety
Master evaluation fundamentals, LLM-as-judge, RAGAS, regression testing, red teaming, and production observability. Pulls from the Evaluation & Safety phase, the Testing AI Code course, and the Agent Evaluation phase.
The route
- Evaluation & Safety (full phase) — ML & AI Engineering (phase 08). 7 lessons: eval fundamentals, golden datasets, LLM-as-judge, red teaming & safety, OWASP Top 10 for LLMs, production observability, responsible AI, and model cards.
- Testing AI-Written Code (full course) — Testing AI-Written Code. A dedicated course on testing AI-generated code: test design, coverage strategies, mutation testing, and building CI/CD gates for AI code quality.
- Agent Evaluation & Observability — Agentic AI Engineering (phase c2-09). 5 lessons on agent eval frameworks, LLM-as-judge with self-consistency, benchmark design, CI/CD for agents, and structured red team campaigns.
What you build
A complete evaluation suite with golden datasets, LLM-as-judge scoring, regression tests, a red team report, and CI/CD integration.
Every learning path
Create a free account — the opening phases of 24 of 30 courses are free, no credit card. Or see Pro pricing.
All courses · Pricing · About · FAQ · Glossary