vLLM: LLM inference and serving
Details
This session introduces vLLM, an open-source library for LLM inference and serving. We will build an intuitive systems-level understanding of how vLLM approaches these production challenges, from a token’s journey through prefill and decode to the scheduling and memory-management techniques that make high-throughput serving practical.
Start exploring at vllm.
Slides for past meetups posted: Github
Recordings posted at: YanAITalk
Feel free to reach out if you want to present at upcoming meetups!
Note: You must have a Zoom account to login (free account is sufficient). Zoom Link will be posted to the event page one day before the meetup.
Related topics
Artificial Intelligence
Deep Learning
Machine Learning
GPU
Data Science
