Glossary › Training
Reinforcement Learning from Human Feedback. Process: (1) collect human preference comparisons, (2) train reward model, (3) optimize LLM policy via PPO/DPO. Aligns model outputs with human preferences.
Backpropagation · Batch Size · Fine-Tuning · Gradient Descent · LoRA (Low-Rank Adaptation) · DPO (Direct Preference Optimization) · QLoRA · Learning Rate
DeVenture Academy teaches AI engineering through projects — every lesson is paired with a hands-on lab you run on your own machine with real tools.
Create a free account — the opening phases of 24 of 30 courses are free, no credit card. Or see Pro pricing.
All courses · Pricing · About · FAQ · Glossary