Building a Transformer LM
Details
Agenda:
Join us as we work through the second part of Assignment 1 of Stanford's CS336: Language Modeling from Scratch. We'll focus on the transformer architecture, overviewing each block. We'll cover: einsum notation, Linear and Embedding modules, RMSNorm, the SwiGLU feed-forward network, RoPE, softmax and scaled dot-product attention, causal multi-head self-attention, and assembling the pre-norm transformer block into the full LM. We'll finish by training the model on TinyStories and generating some sample output.
https://github.com/stanford-cs336/assignment1-basics/blob/main/cs336_assignment1_basics.pdf
Venue:
Boston Public Library (Kirstein Business Library & Innovation Center (KBLIC): Alcove 4)
700 Boylston St, Boston, MA 02116
Time:
6:00 - 8:00 PM EST
