

About us
What's happening Stockholm. We are firing up a local MLOps chapter for this amazing city!
The MLOps Community fills the need to share real-world Machine Learning Operations best practices from engineers in the field. While MLOps shares a lot of ground with DevOps, the differences are as big as the similarities. We needed a community laser-focused on solving the unique challenges we deal with every day building production ML pipelines.
We’re in this together. Come learn with us in a community open to everyone. Share knowledge. Ask questions. Get answers.
Upcoming events
1

#39 - Model Optimization & CPU Based Inferencing
AI Sweden, Folkungagatan 44, Stockholm, SE***Do note that you will have to sign-up at the Luma event page prior to attending. Do so here: https://luma.com/4si349kz***
All right!!!
Meetup #39 will take place September 24 at AI Sweden and focus on "Model Optimization & CPU Based Inferencing". We're super excited to team up with Spanish Quantum AI wizards Multiverse Computing and tech giants Intel and HPE. The program is currently being worked out, so stay tuned for updates!
There is a tremendous amount happening in the area of model optimization and inferencing. Novel approaches to both seem to be announced daily, such as making it possible to run the recently released 2.78 trillion parameter model Kimi K3 on a single CPU with 8GB of memory...!!! or Multiverse Computing's July 23 announcement that all their compressed models now run on Intel Xeon 6 Processors.
So buckle up and brace for impact!
Event Program
- Doors open at 17:00 CET
- Talks begin at 17:45 CET
- There will be pizzas & drinks
- There may be a moderated Q&A session....Speaker Line-Up TBD...
- Franco Serra, Solutions Architect at Multiverse Computing will give a talk titled "Ultra Efficient Models to Scale your GenAI Datacenter & Fit for purpose on the Edge deployment". Abstract: As organizations race to deploy Generative AI, two challenges dominate: how to scale inference economically in the datacenter, and how to bring intelligence on the edge where connectivity, power and footprint are constrained. Multiverse Computing makes ultra-efficient, compressed AI models, including LLMs, VLMs, speech-to-text, and computer vision models — engineered to deliver the same accuracy with a fraction of the memory requirements. By integrating with Intel® hardware, enterprises can deploy larger AI models on existing infrastructure rather than expanding it. In the datacenter, leaner models translate directly into higher throughput, lower energy consumption, and improved ROI by enabling more users and workloads to run on the same infrastructure. At the edge, the same compression breakthroughs make it possible to deploy GenAI in constrained environments — bringing secure, low-latency AI to tactical environments.
- Jonas Svennebring, Principal Engineer at Intel and Theo Charitidis, Machine Learning Engineer also at Intel will give a talk titled "Architecting the future AI/ML CPUs". Abstract: The future of AI/ML capable CPUs will depend on architectures that balance computational performance, memory efficiency, and adaptability. To that end, it is essential that high multiply-accumulate (MAC) throughput enabled, supported by a well-balanced memory subsystem. Another key architectural guidance is figuring out which workloads should realisticly run on general-purpose cores and which should be delegated to specialized accelerators. Finally, supporting the correct dynamic range and numerical precision will be critical for striking the right balance between reliable application results and minimizing hardware costs."
- Johan Fondin, Solution Architect at HPE will give a talk titled "Out of the box KV-cache acceleration". Abstract: How do you increase the number of concurrent sessions on a GPU by a hundredfold, and 20x times faster as well? Learn how HPE solves this issue with their new KV-cache acceleration technology for AI platforms .
Looking forward to seeing you there,
/Patrick & the Stockholm MLOps Team
136 attendees
Past events
45

