Supervised Fine-Tuning to Reinforcement Learning: Train Models to Think Better
Details
Note: It is mandatory to both RSVP and fill out this Google Form https://forms.gle/pJamac2Ejq3LYkZY6 if you miss either of these, the security will not allow you to attend the event. Please ensure that you do not miss filling it out.
This is an in-person meeting only.
Note: Certificates will be provided for those who attend the event and fill out the Google Form that will be shared at the end of the event.
About the event
How do you train a model to get better at a task? What examples do you give it? How do you decide which answers deserve a reward? And once you’ve trained it, how do you know it actually improved?
We’ll work through these questions together, starting with a pretrained model and using Supervised Fine-Tuning (SFT) with LoRA to train it on examples. Then we’ll explore DPO for learning from preferences and GRPO for reinforcement learning with rewards. Along the way, we’ll look at what changes inside the model, where things get tricky and what you need to try this yourself.
This is a hands-on session, so we’ll spend time with code, compare results and make sense of the numbers. We’ll also tackle a question that’s easy to overlook when training seems to be going well and that is, is the model getting better at the task or just getting better at earning the reward?
If you’re curious about how models learn after pre-training, come along. Some familiarity with Python and machine learning will help, but you don’t need to have trained a model before.
Looking forward to seeing you at the venue and happy learning.
