Skip to content

Details

Your prompt worked in testing. Then the model got updated, or a user phrased something differently, and it quietly stopped working. Nobody noticed until a customer did. Most teams treat evals as something they will get to later. This session is the argument for doing it first, and a look at the tool that makes "first" practical.
Serj Smorodinsky (DSPy contributor) and Brett Kennedy, authors of the Manning book "Building LLM Applications with DSPy", will show how to stop hand-tuning prompts and start measuring them: build a baseline, define what a correct answer looks like, and let an optimizer beat you.

What you'll learn:

  • Why hand-tuned prompts degrade, and why the feeling that we are good at prompting may be an illusion created by small test sets
  • The three-stage loop - baseline, evaluate, optimize - and why evaluation has to come before optimization
  • How to define a metric for tasks with a right answer (classification) and for tasks without one (summarization, RAG)
  • How to swap models in one line and choose on measured accuracy, cost, privacy and context window - including when a cheap or local model can do the job you are paying a frontier model for
  • What to do when the model changes under you: rerun the loop instead of reopening a 200-line prompt
  • Where DSPy is worth adopting, and where a plain API call is still the right answer

Why this matters if you build or buy AI: everyone can demo a prompt. Few can show what it scored, on what data, before release - and fewer still can show it is scoring the same way six months later. That evidence trail is the difference between a pilot and a production approval, and it is what your risk and compliance teams will ask for.
Bring your hard questions for the Q&A, especially on the part no book fully solves yet: how to keep monitoring a live LLM system after release.

Friday, September 18 | 12:00-1:30 PM ET | Virtual (Zoom)
Lunch-and-learn format: 10 minutes on why, 40 minutes hands-on, then Q&A held to the end.

About the book: "Building LLM Applications with DSPy: Replacing manual prompts with systematic optimization" is in Manning Early Access, with print publication expected Fall 2026. It builds a classifier, a summarizer, and an agentic RAG chatbot end to end, with evaluation and optimization at every step and no prompt written by hand. https://www.manning.com/books/building-llm-applications-with-dspy
Ebook coupons will be raffled at the event. 45% off any Manning digital product with code TORONTO24.

About the speakers:
Serj Smorodinsky is a DSPy contributor, data scientist, and AI engineer with over ten years of combined experience in software development and data science. His work spans conversational AI for customer service, agentic workflow automation, and LLM evaluation, and he has led teams building chatbots and retrieval-augmented systems for enterprise clients. He teaches agentic systems and data science in production at Nebius Academy.
Brett Kennedy is a data scientist with over thirty years of software development experience and more than ten in data science. He is a regular open source contributor, publishes on Medium and Towards Data Science, and is the author of "Outlier Detection in Python" (Manning).

Serverless Toronto has been bridging the gap between IT and business needs since 2018. 6,000+ members. Past sessions: youtube.serverlesstoronto.org

Related topics

Artificial Intelligence Applications
Amazon Web Services
Cloud Computing
Software Architecture
Enterprise Software

Sponsors

Manning Publications

Manning Publications

Monthly Book / Video / liveProject giveaways.

Magma Inc

Magma Inc

Organizes logistics, speakers, and resources for each event.

You may also like