Building an AI Platform: SDK Design, Model Routing, Vector Search & Evals
Details
Beyond the Model API Call - Suhas Suresha
A successful AI prototype often leads to a slow and expensive production reality. The model API call makes up only 2% of a functioning AI system. The remaining 98% involves the complex infrastructure that engineering teams frequently overlook.
In the first part featuring the authors of Designing AI Systems, Suhas Suresha will break down the reality of bringing LLM applications to scalable production. We will explore how decentralized development creates technology sprawl and missed deadlines. He will reveal the architectural shifts needed to build a unified AI platform where models, memory, guardrails, and evaluations operate as independent services behind a seamless single SDK experience.
He’ll cover:
- SDK and API design: how to make a distributed platform feel like a few lines of Python for the developer
- Model service: covering provider abstraction, routing, fallbacks, and cost tracking
- Data service: covering ingestion, vector search, and hybrid retrieval
- Experimentation service: treating evals and A/B testing of AI prompts as a platform capability rather than a one-off script.
About the Speaker:
Suhas Suresha is a Senior Machine Learning Engineer at Adobe, where he builds large-scale generative AI platforms across the full machine learning lifecycle. He previously co-founded QALY, where he helped deploy real-time ECG analysis models to more than 100,000 users. He holds a master's degree in computational and applied mathematics from Stanford University.
**Join our Slack: https://datatalks.club/slack.html**
