About this Course
The two dominant approaches to making LLMs follow instructions: both have sharp tradeoffs. This comprehensive curriculum is designed to give you hands-on experience with RLHF & DPO: Aligning Language Models. This AI Alignment & ML Engineers curriculum is designed to give you hands-on experience and deep conceptual understanding.
Across 4 intensive modules, you'll tackle real-world challenges and build practical projects that reinforce your learning. By the end of this journey, you'll have the skills and proof of work to demonstrate your expertise.
What you'll learn
Master the core concepts of supervised fine-tuning (sft) & data formatting.
Gain hands-on experience with reward model training for rlhf.
Understand the architecture behind proximal policy optimization (ppo) for llms.
Implement production-grade direct preference optimization (dpo) & modern variants.
W1
Supervised Fine-Tuning (SFT) & Data Formatting
Master the core concepts of supervised fine-tuning (sft) & data formatting.
3 videos•129m
3 readings
3 topics
1 homework
W2
Reward Model Training for RLHF
Gain hands-on experience with reward model training for rlhf.
3 videos•66m
3 readings
3 topics
1 homework
W3
Proximal Policy Optimization (PPO) for LLMs
Understand the architecture behind proximal policy optimization (ppo) for llms.
3 videos•112m
2 readings
3 topics
1 homework
W4
Direct Preference Optimization (DPO) & Modern Variants
Implement production-grade direct preference optimization (dpo) & modern variants.
3 videos•63m
3 readings
3 topics
1 homework
01
Learn
Watch curated videos and read study resources
02
Practice
Practice what you learned
03
Build Projects
Build projects using your new gained knowledge
04
Submit & Verify
Submit your project and get verified by our system
References
Rate this course
Help the community find verified technical paths.
Community Insights
0Join the discussion
Sign in to share your thoughts and technical insights.
Loading insights...