EulerFold
AI Alignment & ML Engineers

RLHF & DPO: Aligning Language Models

4 weeks
0 Learners
Jul 23

The two dominant approaches to making LLMs follow instructions: both have sharp tradeoffs. This comprehensive curriculum is designed to give you hands-on experience with RLHF & DPO: Aligning Language Models.

Share:

About this Course

The two dominant approaches to making LLMs follow instructions: both have sharp tradeoffs. This comprehensive curriculum is designed to give you hands-on experience with RLHF & DPO: Aligning Language Models. This AI Alignment & ML Engineers curriculum is designed to give you hands-on experience and deep conceptual understanding. Across 4 intensive modules, you'll tackle real-world challenges and build practical projects that reinforce your learning. By the end of this journey, you'll have the skills and proof of work to demonstrate your expertise.

What you'll learn

Master the core concepts of supervised fine-tuning (sft) & data formatting.
Gain hands-on experience with reward model training for rlhf.
Understand the architecture behind proximal policy optimization (ppo) for llms.
Implement production-grade direct preference optimization (dpo) & modern variants.

Prerequisites

intermediate Level

Requires basic familiarity with the tech stack.

  • Data structures (Trees, Graphs)

Ideal for

AI Alignment & ML Engineers

AI Alignment & ML Engineers Professionals
Tech Enthusiasts
W1

Supervised Fine-Tuning (SFT) & Data Formatting

Master the core concepts of supervised fine-tuning (sft) & data formatting.

3 videos129m
3 readings
3 topics
1 homework
Learn

Topics

1.1
SFT Dataset Formatting & Chat Templates
41 minutes
1.2
Parameter-Efficient Fine-Tuning (LoRA & QLoRA)
28 minutes
1.3
Cross-Entropy Loss & Training Dynamics
60 minutes
W2

Reward Model Training for RLHF

Gain hands-on experience with reward model training for rlhf.

3 videos66m
3 readings
3 topics
1 homework
Learn
W3

Proximal Policy Optimization (PPO) for LLMs

Understand the architecture behind proximal policy optimization (ppo) for llms.

3 videos112m
2 readings
3 topics
1 homework
Learn
W4

Direct Preference Optimization (DPO) & Modern Variants

Implement production-grade direct preference optimization (dpo) & modern variants.

3 videos63m
3 readings
3 topics
1 homework
Learn
01

Learn

Watch curated videos and read study resources

02

Practice

Practice what you learned

03

Build Projects

Build projects using your new gained knowledge

04

Submit & Verify

Submit your project and get verified by our system

Rate this course

0.0
0 reviews

Help the community find verified technical paths.

Community Insights

0

Join the discussion

Sign in to share your thoughts and technical insights.

Loading insights...