Skip to content

Details

This is a joint meetup with sfspark.org. Please register there only!

We need a video sponsor for this meetup!

MLeap is an open-source technology that allows Data Scientists and Engineers to deploy ML Pipelines and Models trained in Spark, Scikit-Learn and TensorFlow to a scoring engine instantly.

During our presentation, we will show you how to:

• Train feature transformers (standard scalers, one-hot-encoders, PCA, etc) in Spark

• Train a model in both Spark and TensorFlow

• Deploy both the feature transformation pipeline and the trained models to a cloud-based API server as well as an IoT device

Why MLeap? Data Scientists use a myriad tools to analyze datasets, clean them, build offline models and validate their performance. The resulting scripts are thrown across the wall to Data Engineers and Architects whose job is to bring these pipelines to production. The Engineers are left with the unenviable job of not only reproducing the Data Scientists’ conclusions, but to also scale the resulting pipeline - both of which require a deep understanding of data science itself. As a result, most, if not all, machine learning deployments in the wild end up either too simplistic or take too long to deploy to production.

MLeap solves this problem by providing serialization of ML Pipelines to an MLeap Bundle, which is a graph-based serialization framework built on top of Protobuf 3 and JSON. In addition, MLeap also provides a highly optimized execution engine that doesn’t rely on the Spark-context or scikit-learn, making inference blazing fast and is capable of executing one model or thousands of models in parallel. Today, MLeap is used in production by silicon valley start-ups as well as Fortune 500 companies and we're looking forward to sharing it with the broader community.

Speakers

Hollin Wilkins is a co-founder of Combust, an ML/AI start-up in the Bay Area. He has been working on machine learning infrastructure since 2015, focusing on platforms for data scientists and engineers to rapidly iterate on ML algorithms and pipeline deployments. Previously he worked in the games industry at LindenLab on Blocksworld and Versu, helping to build everything from game UI, to servers, to custom logic languages that drive user experiences. He holds a degree in Biology from Cornell University and spends his time hiking with his dog and snowboarding. Mikhail Semeniuk is a co-founder of Combust, an ML/AI start-up in the Bay Area. Prior to Combust, he held a number of leadership roles in Product Management and Data Science for Shift, TrueCar, and UnitedHealth Group. Mikhail studied Mathematics and Economics at the University of Minnesota. He spent a good part of his career working on both the research and development side of machine learning, which inspired his mission to bridge the gap between data science and engineering. He grew up in Minneapolis and lived in Venice, CA for 6 years where he pursued skydiving and hopes of being a decent surfer, and now resides in the Bay Area.

Related topics

You may also like