Skip to content

Details

Important time note: Please plan on arriving between 5:30 and 6:00 as the elevators lock after 6 and you'll need to message us and we'll need to come get you.
The building address is 4450 Bridge Park
The entrance is 6620 Mooney St, Suite 400

You will need to scan your ID at the door to get a visitor badge.

Abstract
Tensors, Tokens & Too Much VRAM: A Developer's Guide to Running LLMs Locally

What actually happens when you send a prompt to an LLM? What are tokens, tensors, attention, quantization, and all those mysterious gigabytes of VRAM doing? And why does your GPU suddenly become the most important piece of hardware in the room?

This talk takes a developer-friendly tour through the fundamentals of modern LLMs. We’ll try to demystify concepts like context windows, model quantization, and inference performance without requiring a PhD in linear algebra.

Then we'll take those concepts out of the textbook and into the real world, looking at tools such as Ollama, LM Studio, and vLLM and what it takes to run useful models on your own hardware.

By the end, you'll have a mental model for what an LLM is actually doing, how local inference works, and why sometimes the answer to "How much VRAM do I need?" is apparently "all of it."

YouTube Link
TBD

Related topics

Artificial Intelligence
.NET
Computer Programming
Software Development
Software Engineering

Sponsors

Central Insurance

Central Insurance

Meeting space and food

You may also like