Skip to content
All apps

RLHF trainer — you are the human

About this app

You compare answers: a reward model learns from them (σ(r_A − r_B)), the policy shifts under a KL leash; with a loose leash the flattering answer wins (reward hacking); plus the DPO route.

Subject: Neural networks

More from Neural networks

A VisuApp by heyprof: interactive, free in your browser, no sign-up. What is a VisuApp? · All apps