75th Deep Learning Meetup: AI Manipulation / Scaling LLM Infrastructure
Details
Hi Deep Learners,
We are happy to announce our first Vienna Deep Learning Meetup after the summer break: On September 23
we'll be hosted by 42 Vienna (in Heiligenstadt) and will feature two exciting topics: AI manipulation and Scaling AI Infrastructure on premise.
***
Agenda:
- 18:15 Arrival
- 18:30 Welcome by the meetup organizers
- Introduction by the host: 42 Vienna
- 18:45 Talk 1: (tbd) by Jason Hoelscher-Obermaier (Director of Research at Apart Research)
- 19:30 Announcements
- Networking Break
- 20:00 Talk 2: AI Inference Engines and Instances: Strategies for Scaling LLM Infrastructure by Jonas Aaron Vander (CTO at Xinity)
- 20:30 Networking
- ~22:00 Wrap up & End
***
Talk Details:
--------------
Talk 1:
(coming soon)
About the speaker:
Jason Hoelscher-Obermaier
Talk 2: AI Inference Engines and Instances: Strategies for Scaling LLM Infrastructure
LLM inference is more than just deploying a model. With Ollama, vLLM, SGLang, there are a variety of specialized engines, each with its own strengths in latency, throughput, ease of use, and configuration. In this talk, we take a practical look at modern LLM infrastructure: Which engine to use when? How do we scale beyond single nodes? And why is context-aware orchestration the key to efficiency when connecting different engines in a cluster?
By the end, you’ll have a roadmap for choosing the right engine for your use case and managing it in a scaled production environment.
Section topics (tentative):
- Modern LLM infrastructure is multi engine
- Ollama excels at local and edge efficiency
- vLLM dominates high-throughput production serving tasks
- SGLang is ideal for structured outputs
- TGI suits HuggingFace ecosystem integration
- Generic routers suffer from cache blindness
- Context-aware orchestration is the future
About the Speaker
Jonas Aaron Vander is CTO and co-founder of Xinity, a Vienna-based sovereign AI infrastructure company, and the architect behind Xinity Runtime, an open-source, OpenAI-compatible inference platform that lets regulated European enterprises run LLMs entirely on their own hardware, currently serving production workloads like Mediengruppe Wiener Zeitung. His background is in AI solution architecture and MLOps.
We are looking forward to welcoming you at this meetup!
Your VDLM organizer team
