Hacker Newsnew | past | comments | ask | show | jobs | submit | fromlogin
RLHF and Post-Training Course by Nathan Lambert (rlhfbook.com)
2 points by ankitg12 56 days ago | past | 1 comment
Reinforcement Learning (I.e. Policy Gradient Algorithms) (rlhfbook.com)
2 points by vinhnx 6 months ago | past
Reinforcement Learning from Human Feedback (rlhfbook.com)
133 points by onurkanbkrc 7 months ago | past | 5 comments
RLHF Book (rlhfbook.com)
479 points by jxmorris12 on Feb 1, 2025 | past | 37 comments

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: