LLM Inference at Scale: The Hidden Cost Nobody Talks About
Details
Most teams obsess over model selection and fine-tuning — then get blindsided when the inference bill hits. A single poorly configured deployment can burn through hundreds of thousands of dollars before anyone notices. This session breaks down exactly how to prevent that.
We'll go from first principles — how LLM inference actually works under the hood — to the decisions that make or break your production costs: GPU architecture selection, batching strategies, KV cache management, quantization tradeoffs, and throughput optimization for handling thousands of concurrent requests. You'll leave with a mental model for auditing your own inference stack and a practical checklist for cutting costs without sacrificing latency.
About the Speaker
Saniya Jaswani is an Applied AI Data Scientist with over 7 years of experience building production-grade AI systems. She holds an MTech in AI from IIT Jodhpur and currently works on multi-agent LLM systems, LLM fine-tuning, and inference infrastructure optimization. She has hands-on experience across the full inference stack — from model quantization and mixed precision training to deploying agentic pipelines at scale.
