Training and post-training
SFT / RLHF / DPO
A family of post-training methods for teaching instruction following and preference-aligned behavior. SFT learns from demonstrations; RLHF and DPO use preference comparisons through different optimization procedures.
Key terms
- supervised fine-tuning
- preference data
- reward model
- alignment
Connected across the project
Builds on
Related technology
Historical context
Primary sources
This is a curated technical map, not a claim of comprehensive coverage.