Ultimate AI - Running inference directly in Go apps without CGO. Kronkai.com
Details
Running open-source models is about much more than loading weights and sending a prompt. Every model server must make the same kinds of decisions: which model and quantization fit the available hardware, how requests are admitted and batched, how prompts are rendered, how context is cached, how tokens are sampled, and how concurrent generations share compute without sharing state. In this talk I will show the foundations of open-source model serving through Kronk, a model SDK and server written for Go. Kronk gives us a concrete system to examine, but the concepts apply broadly to model servers built on inference engines such as llama.cpp. We begin with the architecture of an inference stack and small working programs. Then I will provide a high-level view of what it takes to run inference.
