Skip to content

Details

Note: the yearly conference of Bay Area AI and many other Bay Area developer meetups, Scale By the Bay, is coming back live to Oakland on November 13-15!

Following the PyTorch Conference, we'll have select speakers who are in town at this free meetup. Program to come includes speakers from IBM, Meta and Nvidia.

Registration is required.

IBM is hosting the meetup and you will need an ID matching your full name used for registration at the link above.

Talk 1: Efficiently serving LLMs at scale, Nick Hill, IBM

In this talk I will discuss challenges of serving LLMs efficiently in highly concurrent, multi-user contexts, and some of the optimizations unique to these kinds of models that have emerged over the the last year. These include "continuous batching" of heterogeneous requests. I'll dig into the implementations which involve careful manipulation of tensors with PyTorch.

Nick Hill, IBM. Nick is a Senior Research Engineer focused on scalable serving of large language models. He previously led the architecture and development of distributed machine learning infrastructure supporting key IBM AI cloud products and services including Watson Assistant, Watson Discovery and Watson Natural Language Understanding. He designed and implemented the Model-Mesh serving framework that supports hundreds of thousands of models, now a key component of the KServe open source project. He is also an author of and contributor to other open source projects.

Talk 2: What's New for Python Developer Infrastructure, Omkar Salpekar and Elias Uriegas, Meta

A few updates on the state of the world for PyTorch Developer Infrastructure as well as a lookahead into future projects. Also some info on how we do releases for all of PyTorch at scale.

Omkar Salpekar, Meta AI, Software Engineer
Omkar has worked on PyTorch since 2019, focusing on Distributed Training and Developer Infrastructure. He's worked on PyTorch's support for Data- and Model-Parallel training support for large models, elastic training for large-scale jobs, automating PyTorch's package release infrastructure, and developer tools for PyTorch developers and users.

Elias Uriegas, Meta, Engineering Manager
Eli Uriegas currently supports the PyTorch Developer Infrastructure team. Previously he had worked on Developer Infrastructure at Docker, as well as Rackspace. Eli has interests in enabling developer productivity at scale and without breaking the bank and has been at the forefront of projects that have enabled PyTorch developers to move faster!

Talk 3: Tensor and 2D Parallelism -- Junjie Wang, Xilun Wu,
Iris Zhang, Meta

Junjie Wang, Software Engineer, Meta
Junjie is a software engineer working on PyTorch distributed, especially Tensor Parallel and Sequence Parallel. Working on multiple projects within Meta but eventually found passion in PyTorch.

Xilun Wu, Software Engineer, Meta
Xilun is a software engineer on PyTorch Distributed at Meta Platforms. He has been actively contributing to PyTorch Distributed components including DTensor and TCPStore, as well as third-party dependencies like Gloo. His most recent focus is on scaling TCPStore and ProcessGroup for large training clusters.

Iris Zhang, Software Engineer, Meta
Iris is a software engineer from PyTorch Distributed team at Meta. Iris works on PyTorch DeviceMesh based Distributed APIs as well as Distributed Checkpointing.

Talk 4: AI for Science with Neural Operators, Jean Kossaifi, NVIDIA

Applying AI to science problems is an active research area, aiming to speedup model development and enable better and faster Scientific discovery and engineering design. In many domains, this requires learning mapping between function spaces defined on continuous domains, e.g. to learn spatiotemporal processes or the solution operator to partial differential equations. Neural operators generalize traditional deep learning to do this and can replace simulators while being several orders of magnitude faster. In this talk I will give an introduction these neural operators and how they work and I will give an overview of their application to concrete problems such as weather forecasting.

Jean Kossaifi is a senior research scientist at NVIDIA, working on fundamental ML and algorithms for AI. He has worked extensively on a variety of problems in machine learning and computer vision and is currently focused on applying them to AI for Science. Prior to this he was a research scientist and founding member at the Samsung AI Center in Cambridge, following his PhD in AI at Imperial College. He created and contributes to several Python open source libraries, including TensorLy, neuraloperator, tensorly-quantum and more, all with the goal to make the latest works simple and accessible to all.

Related topics

Events in San Francisco, CA
Artificial Intelligence
Natural Language Processing
Big Data
Semantic Web
Enterprise Search

You may also like