Optimizing Continuous RL @ Fireworks: Rollout, Reward, Update, Repeat!
Details
Zoom link: https://us02web.zoom.us/j/82308186562
Talk #0: Introductions and Meetup Updates
by Chris Fregly and Antje Barth
Talk #1: Rollout, Reward, Update, Repeat: How Fireworks Optimizes Continuous Post-Training
by Sinan Ozdemir @ Fireworks.ai
The RL loop is three simple steps: rollout, reward, weight update. What makes it work continuously in production is all in the data plumbing and in how you design your learning environment.
We'll walk through a post-training RL session on Kimi K3 using Fireworks' serverless API, and talk about the choices of algorithms, loss functions, LoRA rank, hyperparameters, reward design, and how they all interact with how a model learns (or doesn't).
Zoom link: https://us02web.zoom.us/j/82308186562
Related Links
Github Repo: https://github.com/cfregly/ai-performance-engineering/
O'Reilly Book: https://www.amazon.com/Systems-Performance-Engineering-Optimizing-Algorithms/dp/B0F47689K8/
YouTube: https://www.youtube.com/@AIPerformanceEngineering
DeepLearning.ai: https://bit.ly/gllm
