LLM Alignment & RLHF

Master the full LLM alignment pipeline: from RL fundamentals to PPO to RLHF to DPO. Understand exactly how ChatGPT and Claude were trained — and build the same pipeline yourself. For engineers who want to work on LLM training, not just use LLMs.

The route

  1. Pre-training & Fine-tuningTransformers & LLMs from Scratch (phase tf-03). Understand how base LLMs are pre-trained and fine-tuned — the foundation you need before understanding RLHF.
  2. Reinforcement Learning (full course)Reinforcement Learning. 4 phases, 15 lessons: RL fundamentals (MDPs, bandits), model-free RL (Q-learning, DQN, PPO), model-based RL (MuZero, AlphaGo), and RLHF for LLMs (reward modeling, PPO alignment, DPO, Constitutional AI). Build an RL agent and an RLHF pipeline.

What you build

A trained RL agent that learns to play a game from scratch, plus a complete RLHF pipeline that aligns an LLM with human preferences — the same techniques used to train ChatGPT and Claude.

Every learning path

Create a free account — the opening phases of 24 of 30 courses are free, no credit card. Or see Pro pricing.

All courses · Pricing · About · FAQ · Glossary