English
← Back to AI Technology

Training and post-training

SFT / RLHF / DPO

A family of post-training methods for teaching instruction following and preference-aligned behavior. SFT learns from demonstrations; RLHF and DPO use preference comparisons through different optimization procedures.

Reference period
2017 / 2022 / 2023
Last reviewed

Key terms

  • supervised fine-tuning
  • preference data
  • reward model
  • alignment

Connected across the project

Primary sources

  1. Deep reinforcement learning from human preferences
  2. Training language models to follow instructions with human feedback
  3. Direct Preference Optimization

This is a curated technical map, not a claim of comprehensive coverage.