Skip to content

Details

Hybrid event: In-person, Zoom and YouTube
If you want to join remotely, you can submit questions via Zoom Q&A. The zoom link:
Zoom (updated 6:55 pm)
https://acm-org.zoom.us/j/96722167486?pwd=Z8KPqxueHqb9syVRxriwGujYGUXjnm.1
Join via YouTube:

AGENDA
6:30 Door opens, food and networking (we invite honor system contributions)
7:00 SFBayACM upcoming events, introduce the speaker
7:15 Speaker presents.
8:30 - 8:45 finish, depending on Q&A

Join SF Bay ACM Chapter for an insightful discussion on:

### Abstract & Overview

Almost every asynchronous action at Meta passes through a single system most people have never heard of. The Facebook Ordered Queueing Service (FOQS), a fully managed, horizontally scalable priority queue, moves close to a trillion items per day for 300+ engineering teams, and it has become critical infrastructure for AI at Meta's scale. Async LLM inference, Llama serving, GenAI image generation, and AI compute demand control all ride on it.
This talk traces how a queue originally built to absorb massive backlogs and prioritize work across highly heterogeneous producers and consumers grew into the reliability layer beneath Meta's AI stack. It is an impact story rather than a feature tour: how one system came to serve hundreds of teams without them stepping on each other, how it keeps the most important work moving under enormous load, and how it evolved from isolated regional deployments into a globally distributed service that delivers region-level disaster recovery in seconds with zero client-visible downtime. The finale looks at how that same queue now acts as a control plane for shaping AI compute demand.

### Speaker biography:

Jasmit Kaur Saluja is a software engineer at Meta Platforms, where she builds large-scale distributed systems powering products used by billions of people. She is the technical lead of the Facebook Ordered Queueing Service (FOQS), a globally distributed, disaster-ready priority queue processing roughly one trillion items daily, which backs Meta AI's infrastructure at scale. She designed and scaled FOQS, including its cross-region failover architecture, co-presented the work at the Systems @Scale conference, and it is featured on Meta's Engineering Blog. Earlier, I co-authored "vNFS: Maximizing NFS Performance with Compounds and Vectorized I/O," presented at USENIX FAST '17.

Here is her presentatio at Systems @Scale on making a distributed `priority queue` disaster ready:
https://atscaleconference.com/events/systems-scale-fall-2021/?tab=2&item=11

linkedin.com/in/jasmit-kaur-saluja

---
Valley Research Park is a coworking research campus of 104,000 square feet hosting 60+ life science and technology companies. VRP has over 100 dry labs, wet labs, and high power labs sized from 125-15,000 square feet. VRP manages all of the traditional office elements: break rooms, conference rooms, outdoor dining spaces, and recreational spaces.

As a plug-and-play lab space, once companies have secured their next milestone and are ready to expand, VRP has 100+ labs ready to expand into.
https://www.valleyresearchpark.com/

Related topics

Events in Mountain View, CA
Happy Hour
System Administration
IT Infrastructure
Computer Programming
Software Development

You may also like