Reinforcement Learning
From bandits to RLHF — build agents that learn by doing
4 phases. 15 lessons. 15 labs. 1 capstone. Reinforcement learning from foundations to modern RLHF — RL fundamentals (MDPs, bandits, value functions, policy gradients), model-free RL (Q-learning, DQN, policy gradient methods, Actor-Critic, PPO), model-based RL & advanced methods (world models, MuZero, AlphaGo-style search, offline RL), and RLHF for LLMs (reward modeling, PPO for alignment, DPO, Constitutional AI). You build an RL agent that learns to play a game and an RLHF pipeline that aligns a
- Lessons: —
- Labs: —
- Projects: —
- Level: Beginner
Curriculum
- RL Foundations — MDPs, value functions, and Q-learning from scratch.
- Policy Gradient Methods — Policy gradients, REINFORCE, and PPO from scratch.
- Advanced RL — Actor-critic, DDPG, SAC, and RLHF.
Skills You Will Learn
- Markov Decision Processes (MDPs)
- Multi-Armed Bandits & Exploration
- Q-Learning & Deep Q-Networks (DQN)
- Policy Gradients (REINFORCE)
- Actor-Critic & Advantage Estimation
- PPO (Proximal Policy Optimization)
- Model-Based RL & World Models
- Offline RL
- RLHF — Reward Modeling & PPO for LLMs
- DPO & Constitutional AI
Related Courses
Browse all courses · View pricing · DeVenture Academy