Mastery card
I can explain RLHF
Teaching a model using human preference ratings.
Reinforcement learning from human feedback (RLHF) optimises a model using preference data so outputs better match what people rate as helpful and safe.
Sticky trick
RLHF = humans thumbs-up the better answer, model learns.
Take into the room
In plain English, what is RLHF?
Or open the share card 👍
Or start your own 5-word trail from home.