Skip to content

Details

Federated learning development has two major sources of friction: building the workflow correctly and running it efficiently at scale. New users must learn FL-specific concepts, APIs, job layouts, configuration patterns, and validation workflows before they can create a reliable experiment. Experienced users face a different challenge: repeated FL iteration still requires managing candidate jobs, code changes, runs, results, comparisons, and diagnosis.

At the infrastructure layer, scaling FL on shared HPC systems introduces additional complexity around job scheduling, study isolation, containerized execution, GPU allocation, and avoiding designs where orchestration processes unnecessarily occupy accelerator nodes.

This webinar will discuss two areas where NVIDIA FLARE can help reduce these barriers.

First, we will cover Agentic Skills for Federated Workflows. Current FLARE skills package FLARE-specific knowledge into reusable agent workflows, helping users inspect data, convert existing PyTorch code, generate jobs, validate results, and diagnose failures without first mastering the full NVFLARE API surface. We will also discuss the newly added Auto-FL skills, which extend agentic assistance into iterative experiment development by managing candidate jobs, tracking campaign state, comparing outcomes, and guiding follow-up iterations.

Second, we will cover Scaling NVFLARE on HPC Infrastructure. We will discuss how Slurm integration can evolve beyond basic sbatch submission toward a more complete execution model for shared clusters: isolated studies, Apptainer and other container backends, better lifecycle management, and resource patterns that keep GPUs assigned to training work rather than idle orchestration.

Together, these topics show how NVFLARE is evolving across the full path from development to deployment: lowering the learning curve for new users, accelerating iteration for experienced users, and improving large-scale execution on modern HPC systems.

1. Agentic Skills for Federated Workflows

  • Why skills matter: reduce the NVFLARE learning curve by encoding FLARE APIs, job structure, validation steps, and best practices into reusable agent workflows.
  • Current skills: project orientation, PyTorch-to-NVFLARE conversion, data/statistics workflows, job validation, and failure diagnosis.
  • Auto-FL skills: support iterative experiment development through candidate generation, campaign state, result tracking, comparison, and follow-up iteration.
  • Skill development: how skills are authored, tested, and refined using harnesses, evals, deterministic fixtures, and review loops.

Practical outcome: users can start from intent, existing code, or data and reach a validated FL workflow faster.

2. Scaling NVFLARE on HPC Infrastructure

See how an existing PyTorch project can be converted into a validated NVIDIA FLARE job, prepared for execution on a shared Slurm cluster, and optimized through multiple experiment variants with Auto-FL.

  • Why scaling is hard: shared clusters introduce scheduler constraints, study isolation needs, container requirements, GPU allocation issues, and multi-user workload management.
  • Slurm Job Launcher: expanding Slurm integration from job submission to a fuller execution model.
  • Study isolation: separating experiments, artifacts, runtime state, and logs cleanly.
  • Container execution: Apptainer and Pyxis support for reproducible, portable FL workloads.
  • Resource efficiency: avoid holding GPU nodes for orchestration/control processes; keep accelerators focused on training work.

Practical outcome: NVFLARE can better support realistic shared HPC environments and larger federated studies.

Repo: https://github.com/NVIDIA/NVFlare
Speakers: NVFLARE Team
Chester Chen, Peter cnudde, Holger Roth

Join online at
Microsoft Teams meeting
Join: https://teams.microsoft.com/meet/231857209302585?p=saScoNOLEmVRzrg255
Meeting ID: 231 857 209 302 585
Passcode: TZ39Kt9w

***

Dial in by phone
+1 949-570-1120,,615080749# United States, Irvine
Find a local number
Phone conference ID: 615 080 749#

You may also like