LLM Evaluation & Safety

Master evaluation fundamentals, LLM-as-judge, RAGAS, regression testing, red teaming, and production observability. Pulls from the Evaluation & Safety phase, the Testing AI Code course, and the Agent Evaluation phase.

The route

  1. Evaluation & Safety (full phase)ML & AI Engineering (phase 08). 7 lessons: eval fundamentals, golden datasets, LLM-as-judge, red teaming & safety, OWASP Top 10 for LLMs, production observability, responsible AI, and model cards.
  2. Testing AI-Written Code (full course)Testing AI-Written Code. A dedicated course on testing AI-generated code: test design, coverage strategies, mutation testing, and building CI/CD gates for AI code quality.
  3. Agent Evaluation & ObservabilityAgentic AI Engineering (phase c2-09). 5 lessons on agent eval frameworks, LLM-as-judge with self-consistency, benchmark design, CI/CD for agents, and structured red team campaigns.

What you build

A complete evaluation suite with golden datasets, LLM-as-judge scoring, regression tests, a red team report, and CI/CD integration.

Every learning path

Create a free account — the opening phases of 24 of 30 courses are free, no credit card. Or see Pro pricing.

All courses · Pricing · About · FAQ · Glossary