Skip to content

About us

๐Ÿš€ Weโ€™ve Rebranded: Introducing KSUG.AI! ๐Ÿค–โ˜๏ธ๐Ÿณ

Americas: North and South America

Formerly K8SUG, we are now KSUG.AI โ€“ KubeSmart & AI User Group ๐ŸŽ‰
As the cloud-native landscape continues to evolve, so has our communityโ€™s vision. Our focus has expanded to include not just Kubernetes, but also Multi-Cloud, Cloud Native architectures, and AI/ML workloads.

๐Ÿ“ Whether you're joining us in-person or online from anywhere in the world, KSUG.AI is the place for developers, architects, and platform engineers who are building the future with Kubernetes.

๐Ÿ’ฌ We dive deep into:

  • Multi-cloud and hybrid cloud migration
  • Real-world experiences with OpenShift, vanilla K8s, and managed Kubernetes (on-premises and in the cloud)
  • Building, securing, operating, and managing Kubernetes clusters
  • Backup and disaster recovery strategies for containers
  • And now, the intersection of Kubernetes with AI and ML

๐Ÿ‘ฅ Who should join?
Anyone using or planning to adopt Kubernetesโ€”whether you're in DevOps, Platform Engineering, SRE, or AI/MLโ€”will benefit from the shared knowledge and collaboration in this community.

๐ŸŒ Follow us and get involved!
๐Ÿ”ฅ Explore our community ecosystem

Our reach continues to grow:

Join the movement. Learn, share, and lead with KSUG.AI.

KSUG.AI is an independent community and not affiliated with or endorsed by CNCF. Kubernetes, K8s, Kubestronaut are registered trademarks of The Linux Foundation.

Sponsors

KSUG AI

KSUG AI

KSUG.AI | KubeSmart, Cloud Native & AI Community

Upcoming events

1

See all
  • KCD San Francisco Bay Area 2026

    KCD San Francisco Bay Area 2026

    Computer History Museum, 1401 N Shoreline Blvd, mountain view, CA, US

    Register now to attend! => https://luma.com/0f0h32li <=

    We're excited to invite you to the event organized by KCD SF Bay Area. Make sure you register HERE.

    • Use code: KSUG10 to save 28%
    • Walk-in is NOT allowed

    Latest promotions discovered by our community!

    We're actively seeking awesome K8s and AI topic speakers, amazing volunteers and generous sponsors! Fill out the relevant form ๐Ÿ‘‰ https://ksug.ai

    Join us for the second KCD San Francisco Bay Area, the biggest gathering for the cloud native and Kubernetes community in the Bay Area! KCD San Francisco Bay Area is a full-day event with sessions that explores the latest cloud native technologies and topics like Cloud Native AI, Platform Engineering, Security, Observability, trending cloud native open source projects, and more.

    ๐Ÿ“… Event Details

    Date: Tuesday, September 1, 2026
    Location: Compute History Museum, Mountain View, CA

    Register here to secure your spot! See you all there!

    ๐Ÿ‘‰ Join our Discord and WhatsApp for latest update!

    What you're expecting

    9:00 am Welcome
    9:15 am Keynote - Tim Hockin, Google
    9:30 am The Next Decade of Agents: How Agentic Computing is Reshaping Cloud Native - Ronald Petty, RX-M
    Over the past decade, cloud native transformed how we build, deploy, and operate software. Containers, Kubernetes, and the CNCF ecosystem became the foundation of modern applications. Now another shift is underway. AI agents are evolving from chat interfaces into systems that can reason, plan, use tools, collaborate, and take action. As these capabilities mature, an important question emerges: How will agentic computing reshape cloud native over the next decade? This talk briefly looks back at the evolution of cloud native before exploring how agents may influence Kubernetes and the broader CNCF ecosystem. We'll examine trends in agent communication, identity, memory, observability, governance, and orchestration, and discuss which cloud native concepts may endure, evolve, or give way to new abstractions. Rather than predicting specific products, this session focuses on the long-term patterns that could define the next generation of distributed systems.

    10:00 am Code to Cluster: Abstracting Kubernetes ML Complexity with Michelangelo - Paul Zimmerman & Eric Wang, Uber
    Building machine learning platforms on Kubernetes shouldn't force data scientists to become infrastructure experts. True platform engineering means masking the complexities of cluster manifests, distributed scaling, and fault tolerance so data teams can focus purely on innovation. In this session, Paul Zimmerman and Eric Wang from Uber's Michelangelo team explore how platform engineers can leverage the newly open-sourced Michelangelo framework to build a seamless "Code to Cluster" experience. Weโ€™ll dive into strategies for abstracting infrastructure while retaining native control over Kubernetes primitives, deploying complex stacks via Helm, scaling distributed training with Ray and PyTorch, and using CNCF Cadence Workflow for pipeline resilience. Join us to see how to bridge the gap between abstract code and distributed cloud-native scale.

    10:30 am Platinum and Gold Sponsors
    10:45 am Break
    11:00 am Self-Healing Systems: How LLM Agents Are Reinventing Cloud-Native Disaster Recovery - Akshay Pratinav, Garvit Kataria, & Archana Kataria, Intuit
    Cloud-native systems have outgrown the incident response playbook. Static runbooks, manual operator intervention, and reactive on-call rotations weren't designed for the scale and complexity of modern distributed systems โ€” and the result is longer outages, inconsistent recoveries, and burned-out engineers. This session introduces an agentic AI approach to disaster recovery, where LLM-based agents don't just advise โ€” they act. Multiple collaborating agents work in concert to detect anomalies, reason about root causes, evaluate recovery strategies, and execute remediation directly through your existing Kubernetes and cloud interfaces. Human oversight remains intact through safety controls for high-impact operations, so you keep the guardrails without the bottlenecks. We'll show real-world results: dramatic reductions in MTTR and operational toil, with measurable improvements in recovery consistency and reliability. More importantly, we'll walk through the architecture โ€” how agents are orchestrated, how they reason under uncertainty, and how this system evolves from a passive advisory tool into an autonomous SRE co-pilot. If you're an SRE, platform engineer, or architect wondering where AI fits into your reliability story, this session gives you a concrete, battle-tested answer.

    11:30 am Agentic GitOps: Agent and Sandbox Guardrails for CI/CD - Tamao Nakahara, Guild.ai & Leigh Capili, ControlPlane
    With agentic workflows, kubectl commands can have dangerous consequences. Improper RBAC can allow agents to kubectl delete, and that includes deleting your whole CI with no commits to roll back to! Thatโ€™s why Flux's security-first design is even more relevant for agentic GitOps. We'll cover how to use Flux to confine agents to human-reviewable PRs for all sorts of use cases. Weโ€™ll do this with a kernel-sandboxing tool called `nono`. In addition, for your agent management tool of choice, we'll cover how to manage your agents' sandboxes with Flux so that nefarious (or confused) agents can't destabilize the security policies that you have in place. Weโ€™ll cover safe practices for agentic use cases like: - using Prometheus metrics to trigger resource tuning - troubleshooting and rolling back after HPA crashes - agents requesting additional network access with PRโ€™s for human reviewers Come join in!

    12:00 pm Rocket Your Cloud Native Career: Kubestronauts and the Experts Behind the Exams (Lightning Talk & Birds of a Feather) - AmyJune Hineline, Linux Foundation & Giorgi Keratishvili, EPAM
    Cloud native certifications are shaped by the people who build the exams and the professionals who pursue them. In this lightning talk, AmyJune Hineline will offer a behind-the-scenes look at how subject matter experts help define, develop, and validate certification exams, while Giorgi will share how the Kubestronaut program can support professional growth, continued learning, and community involvement. Together, they will explore two different ways people contribute to and benefit from the cloud native certification landscape and get involved in the CNCF community. Bring your questions, experiences, and certification goals, then continue the conversation with both speakers at the Birds of a Feather table during lunch.

    12:05 pm OpenTelemetry Metrics Just Got 25ร— Faster (Lightning Talk & Birds of a Feather) - Cijo Thomas, Microsoft
    OpenTelemetry's metrics performance has long been a sore point for high-throughput users. Recording a simple Counter with three attributes/labels cost around 50 nanoseconds per call - enough to show up in CPU profiles and rule OpenTelemetry out of the busiest code paths. Pre-resolved metric handles are an old answer. Prometheus client libraries expose labelled-metric handles you can cache, and Windows Performance Counters have used the same pattern for decades. OpenTelemetry now offers it as a first-class API, called bound instruments, and on that same hot path, recording drops to under 2 nanoseconds - roughly 25ร— faster (Measured in OTel Rust Sdk) This lightning talk shows where the speedup comes from, when to reach for it (high-frequency counters with a fixed, known attribute set), and - more importantly - when not to. Used the wrong way, the new fast path can be slower than the original. Flexible by default, fast when you need it.

    12:10 pm OpenChoreo: Developer Platform for both Humans and Agents (Lightning Talk & Birds of a Feather) - Sameera Jayasoma, WSO2
    Kubernetes gives platform teams powerful building blocks. But turning those into a real developer experience takes months of work. OpenChoreo is a complete, open-source developer platform for Kubernetes. It's ready to use from day one, for both humans and agents. In this lightning talk, I'll walk through how OpenChoreo provides development and platform abstractions on top of Kubernetes. It comes with a Backstage-powered developer portal, plus built-in CI/CD, GitOps, and observability. Developers can self-serve deployments without needing to be Kubernetes experts. These same abstractions also work well for AI agents. Agents need to build, deploy, and operate workloads too, often alongside humans. I'll share what it means to design a platform that serves both audiences from the start. If you're building a platform for your team, or thinking about how agents fit into your Kubernetes setup, this talk is for you. You'll leave with a clear picture of what a CNCF Sandbox developer platform looks like when built for both humans and agents.

    12:15 pm Lunch & Birds of a Feather (Grand Hall)
    1:15 pm Panel: At the Crossroads of AI and Cloud Native - Arun Gupta, NVIDIA; Janet Kuo, Google; Joseph Sandoval, Adobe; Rita Zhang, CoreWeave
    1:15 pm Workshop (75 min): Agent workshop with Guild
    Platform teams are in between the company's needs to be competitive through agentic strategies and making sure that the internal developers that they support are supported, productive, and meeting security requirements. This workshop will give participants hands-on experience with spinning up agentic environments quickly and ready for use. We will cover 2 use cases: one for engineering team productivity and one for engineering and platform teams to serve internal business teams. By the end of the workshop, participants will have quickly spun up agents, set up context and skills, and have agentic environments that they can continue to maintain, evolve, and improve as needed.

    2:00 pm Why Is Everyone Still Sending Raw Data? - Julia Furst Morgado, Dash0 & Reese Lee, New Relic
    Most teams running the OTel Collector are using it as a passthrough. Data goes in, data goes out and everything lands in the backend whether it is useful or not. That gets expensive fast and it creates compliance problems when PII ends up in your traces. OTTL, the 12:00emetry Transformation Language, ships with every Collector distribution and lets you filter, redact, and normalize telemetry data before it ever leaves your infrastructure. In this session we'll look at the syntax and architecture, walk through real transformation statements for the transform and filter processors, and cover the cases where OTTL makes more sense than building a custom component.

    2:30 pm Building a Production-Grade LLM Serving Platform on Kubernetes with llm-d, KServe, and Gateway API - Goutham Annem, AWS
    Everyone is racing to deploy LLMs on Kubernetes. Getting a model running is easy, running them efficiently and cost-effective at scale is the real challenge. This talk covers the journey from a vLLM StatefulSet to a production-grade inference platform using llm-d, KServe, and Gateway API Inference Extension. We address three bottlenecks: storage drag from massive model weights, infrastructure lock-in from node affinity, and GPU waste from load balancing that ignores KV-cache locality. See how prefix-cache aware routing delivered 3x throughput and 2x TTFT reduction on Llama 3.1 70B across 4 MI300X GPUs. We also discuss upstream CNCF contributions this work produced and how running at scale hardened these projects for everyone. Walk away with architecture patterns for intelligent LLM routing, real benchmarks, and steps to adopt llm-d + KServe in your clusters.

    3:00 pm Break
    3:15 pm When Packets Disappear in the Cloud: Debugging Kubernetes Inference Workloads - Venkat Gattupalli & Haibing Zhou, OpenAI
    Large scale inference workloads on Kubernetes are extremely sensitive to rare packet loss: one missing packet can cause seconds of tail latency or failed requests. These incidents are hardest when they cross the visibility boundary between platform teams and cloud-provider infrastructure: node kernel, vNIC driver, hypervisor, and network. In this session, we share an incident response story from OpenAI's scale inference. Starting from latency symptoms and TCP retransmissions, we used targeted eBPF tracing to follow packets through the node stack and vNIC driver. Correlating sequence numbers, skb state, DMA mappings, and completions showed affected packets leaving the guest driver path, narrowing the issue toward cloud infrastructure. That evidence enabled a shadow traffic workaround, validated latency improvement, and built rollout confidence while the provider fix was underway.

    3:45 pm Nodes Lie, Caches Lag: Designing a Safe, Multi-Cluster Node Lifecycle Controller - Archana Anand, LinkedIn
    For most teams, draining a node isn't something they built. It's a byproduct of the autoscaler, or a kubectl drain script that works until it doesn't. This session showcases a dedicated, multi-cluster node lifecycle controller built to cordon, drain, retire, and replace nodes across a fleet of hundreds of thousands of nodes, hitting every production edge case along the way. It pairs high-scale architecture with the war stories that shaped it: a two-plane design (workload clusters reconciled against a central management hub) and the multi-handler approval protocol built after racing controllers disrupted live workloads. Attendees will learn how to solve stale cache "lies" (dead nodes reporting Ready), bypass pod eviction loops, and design reversible kill switches so unattended fleet automation is crash-safe, leaving with a production-tested blueprint for resilient node automation.

    4:15 pm Shift-Left for Platform Teams: Kubernetes-Native Infrastructure Testing at EarnIn - Priya Namasivayam, EarnIn & Ole Lensmar, Testkube
    EarnIn operates a high-availability fintech platform on 15+ CNCF projects, including Flux, Karpenter, Kyverno, cert-manager, Linkerd, external-secrets-operator, Argo CD, Velero, and more. When infrastructure components fail silently โ€” cert-manager chains breaking post-upgrade, Velero jobs reporting success without executing โ€” the impact is immediate and customer-facing. To address this, we built a Kubernetes-native testing framework using Testkube that validates our full infrastructure stack on every GitOps-triggered change. Tests are defined as Kubernetes CRDs, integrated into our Flux delivery pipeline, and cover TLS validation, policy enforcement, secret sync, backup integrity, and add-on health across the stack. This session walks through our test architecture, the specific failure classes we catch, and how we operationalized infrastructure testing as a first-class part of our platform engineering workflow โ€” not an afterthought.

    4:45 pm AI SRE: Building Incident-Response Agents That Start the RCA Before You Do - Ishan Shah, PayPal
    PagerDuty goes off. Before a human fully opens the laptop, an AI SRE agent can already be pulling telemetry, checking dashboards, correlating logs, and drafting an incident summary. This talk shows how to design an AI incident-response workflow that integrates with tools like PagerDuty, Datadog, and New Relic to accelerate triage without bypassing safety. What would be covered: โ€ข event trigger from alert to investigation โ€ข gathering evidence from observability systems โ€ข forming a first-pass RCA hypothesis โ€ข drafting timelines and incident summaries โ€ข keeping humans in the approval loop

    5:10 pm Raffle & Closing Comments
    5:20 pm Happy Hour

    ๐Ÿ”– ๐Ž๐ง๐ ๐จ๐ข๐ง๐  ๐ƒ๐ข๐ฌ๐œ๐จ๐ฎ๐ง๐ญ๐ฌ:

    โ˜ธ 30% OFF Kubernetes Certs - Code: 30K8SUG
    โ˜ธ 20% OFF FinOps Certs - Code: KSAI_20

    ๐–๐ก๐ฒ ๐–๐ž ๐ƒ๐จ ๐“๐ก๐ข๐ฌ:

    โœ… Learn ๐…๐€๐’๐“๐„๐‘ โšก
    โœ… Certify ๐’๐Œ๐€๐‘๐“๐„๐‘ ๐Ÿ’ฐ
    โœ… Grow ๐’๐“๐‘๐Ž๐๐†๐„๐‘ ๐Ÿ’ช

    ๐Ÿ๐Ÿ“๐ŸŽ,๐ŸŽ๐ŸŽ๐ŸŽ+ follow KSUG.AI ๐Ÿ”ฅ github.com/ksug-ai

    By registering, you consent to the management of your personal information in accordance with KSUG.AI's Privacy Policy. Additionally, you agree that our sponsors may contact you.

    • Photo of the user
    • Photo of the user
    • Photo of the user
    3 attendees

Group links

Organizers

Yongkang H. is a Super Organizer

Members

4,724
See all

Find us also at