Kubernetes and nano-vllm Working Group
Details
This is a working group. There is no formal lecture or teaching format. Do your own demos like it is work and build a set of demos.
The company you work for blew their token budget on OpenAI. Your job is to replace the internal demand for llms with open source models running on gpu rentals from a cloud vendor of your choice. Present a POC or usable system consisting of a dashboard, a query router and cost caps per user.
LLM Deployment and Kubernetes:
https://github.com/kserve/kserve
https://github.com/vllm-project/llm-compressor
https://github.com/llm-d/llm-d
https://github.com/GeeeekExplorer/nano-vllm
