Skip to content

Details

Twenty models. Five teams calling them. One invoice nobody can explain — and a runaway agent that burned through budget before anyone noticed.

Enterprises scaling GenAI have no mechanism to answer basic operational questions: who's calling which model, what's it costing per team, and what happens when someone exceeds their budget? We solved API governance a decade ago — usage plans, rate limits, per-consumer throttling. But for LLM calls? We're back to zero.

In this talk, we will introduce "token sovereignty" — the principle that every token flowing through your org must be attributed, budgeted, and governable — and show a practical AWS-native architecture that enforces it.
You'll learn how to: (1) attribute cost to teams, projects, and users at token granularity, (2) enforce budget ceilings and (3) rate-limit AI consumers with usage plans the way you'd rate-limit API consumers, Using AWS services, no third-party layers — the control plane your AI workloads have been missing.

This event is hybrid - join us at the CGI - Franklin office or through Zoom.

In person Location:
6640 Carothers Pkwy
#400 CGI Office
Franklin, TN 37067
Via Zoom Meeting:
https://us06web.zoom.us/j/87116082506?pwd=xMtuS7hsZU4mCxBIrHamPoQg1qRHqP.1
Meeting ID: 871 1608 2506
Passcode: AWS

Related topics

Events in Franklin, TN
Artificial Intelligence
Amazon Web Services

You may also like