Model Optimization & CPU Based Inferencing
Details
All right!!!
Meetup #39 will take place September 24 at AI Sweden and focus on "Model Optimization & CPU Based Inferencing". We're super excited to team up with Spanish Quantum AI wizards Multiverse Computing and tech giants Intel and HPE. The program is currently being worked out, so stay tuned for updates!
There is a tremendous amount happening in the area of model optimization and inferencing. Novel approaches to both seem to be announced daily, such as making it possible to run the recently released 2.78 trillion parameter model Kimi K3 on a single CPU with 8GB of memory...!!! or Multiverse Computing's July 23 announcement that all their compressed models now run on Intel Xeon 6 Processors.
So buckle up and brace for impact!
Event Program
- Doors open at 17:00 CET
- Talks begin at 17:45 CET
- There will be pizzas & drinks
- There may be a moderated Q&A session....
Speaker Line-Up TBD...
- Franco Serra, Solutions Architect at Multiverse Computing will give a talk titled "Ultra Efficient Models to Scale your GenAI Datacenter & Fit for purpose on the Edge deployment". Abstract: As organizations race to deploy Generative AI, two challenges dominate: how to scale inference economically in the datacenter, and how to bring intelligence on the edge where connectivity, power and footprint are constrained. Multiverse Computing makes ultra-efficient, compressed AI models, including LLMs, VLMs, speech-to-text, and computer vision models — engineered to deliver the same accuracy with a fraction of the memory requirements. By integrating with Intel® hardware, enterprises can deploy larger AI models on existing infrastructure rather than expanding it. In the datacenter, leaner models translate directly into higher throughput, lower energy consumption, and improved ROI by enabling more users and workloads to run on the same infrastructure. At the edge, the same compression breakthroughs make it possible to deploy GenAI in constrained environments — bringing secure, low-latency AI to tactical environments.
- Thomas Melzer, Senior Solutions Architect at Intel will give a talk on the resurgence of the relevance of CPUs for inferencing....
- HPE (on the dynamics of storage, streaming and memory)
Looking forward to seeing you there,
/Patrick & the Stockholm MLOps Team
