Skip to content

Details

Learning with Distributed Data
by Arya Mazumdar

Abstract:
In recent years, large-scale training for machine learning models mandatorily takes place in a massively large distributed system composed of individual computational nodes (e.g., ranges from GPUs to low-end commodity hardware). Such distributed systems are inherently constrained by communication. For example, a computational node 1) may not be able to communicate local information due to limited bandwidth 2) may not want to share information to maintain privacy 2) may intentionally corrupt information as an adversarial attack. What is the compromise in the convergence rate of an optimization algorithm due to these communication constraints? We explore these trade-offs and show that first and second order methods of optimization still work under all such heavy information constraints.

Bio:
Arya Mazumdar is an associate professor in the Halicioglu Data Science Institute of University of California San Diego, with additional affiliation to the Departments of Computer Science and Electrical Engineering. Prior to this, Arya had been a faculty in the College of Information and Computer Sciences in UMass Amherst, and in University of Minnesota, a scientist in Amazon AI/Search, and a postdoctoral scholar at Massachusetts Institute of Technology. Arya obtained his Ph.D. degree from University of Maryland, College Park. Arya is a recipient of multiple awards, including a Distinguished Dissertation Award for his Ph.D. thesis (2011), the NSF CAREER award (2015), an EURASIP JASP Best Paper Award (2020), and the IEEE ISIT Jack K. Wolf Student Paper Award (2010). He currently serves as an Associate Editor for the IEEE Transactions on Information Theory and as an Area editor for Now Publishers Foundation and Trends in Communication and Information Theory series. Arya’s research interests include coding theory, information theory, statistical learning and distributed optimization.

=================
Agenda (Pacific Daylight Time, UTC -07)

  • 5:30 - 5:40 pm -- Gathering and introductions
  • 5:40 - 6:30 pm -- Talk
  • 6:30 - 7:00 pm -- Q & A, discussion

Links to slides and videos of meetup presentations are available on the SDML GitHub repo https://github.com/SanDiegoMachineLearning/talks

=================
Questions?

Join our slack channel or leave a comment below if you have any questions about the group or need clarification on anything.
https://join.slack.com/t/sdmachinelearning/shared_invite/zt-6b0ojqdz-9bG7tyJMddVHZ3Zm9IajJA

You may also like