Skip to content

Details

Novembers's book is "Reinforcement Learning from Human Feedback"!

​This is a casual-style event. Not a structured presentation on topics. Sometimes, the discussion even drifts away from the chapters, but feel free to grab the mic to help steer it back.

​Feel free to join the discussion even if you have not read the book chapters! :)
​Want to discuss the contents during the reading week? Join the Flyte MLOps Slack group.
​-------------------------------------------------
About the book:

  • ​Title: Reinforcement Learning from Human Feedback
  • ​Authors: Nathan Lambert
  • ​Published: August 2026

​Manning (Promo code: AIBookClub should give you 45% off: https://www.manning.com/books/reinforcement-learning-from-human-feedback
​O'rielly platform: https://learning.oreilly.com/library/view/reinforcement-learning-from/9781633434301/
​
​Chapters:

  • ​1 Introduction
  • ​2 A tiny history of RLHF
  • ​3 Training overview
  • ​4 Instruction fine-tuning
  • ​5 Reward modeling
  • ​6 Reinforcement learning
  • ​7 Reasoning and inference-time scaling
  • ​8 Direct-alignment algorithms
  • ​9 Rejection sampling
  • ​10 The nature of preferences
  • ​11 Preference data
  • ​12 Synthetic data
  • ​13 Tool use and function calling
  • ​14 Over-optimization
  • ​15 Regularization
  • ​16 Evaluation
  • ​17 Crafting model character and products

​Book Description
​Reinforcement Learning from Human Feedback: LLM alignment and post-training helps you understand how modern AI models can be adapted to better match the needs and expectations of their users. Rather than surveying the vast field of reinforcement learning, elite AI researcher Nathan Lambert concentrates exclusively on RLHF and its immediate importance to post-training generative AI models.

This compact book gets right to the point. Early chapters establish the training overview, explain instruction fine-tuning, and build reliable reward models. The middle chapters transition into the heart of alignment, exploring core policy gradient algorithms, Direct Preference Optimization (DPO), and inference-time scaling. Later chapters tackle the messy reality of data, guiding you through preference data collection, synthetic data generation, and the nuances of function calling.

Related topics

Artificial Intelligence
Artificial Intelligence Applications
Artificial Intelligence Machine Learning Robotics
Deep Learning
Machine Learning with Python

You may also like