combust / combust/mleap

Standardize Serialization format with Spark

Open
#116 0 comments 7 reactions 0 assignees View on GitHub
epic
Dominant language
Scala
Stars
1.5k
Forks
315
PR merge metrics
No merged PRs in 30d

Description

# Standardizing ML Pipeline Serialization

Currently there is a large array of serialization formats for machine learning models:
1. PMML is an XML-based format primarily targeting the JVM for executing ML models
2. Scikit-learn relies on Python pickling to export models
3. Spark has a serialization format based on Parquet and JSON
4. Various other libraries such as Caffe, Torch, MLDB, etc. have their own custom file formats they use to store models with

We propose a serialization format that is highly-extensible, portable across language and platforms, open-source and with a reference implementation in both Scala and Rust. We call this serialization format Bundle.ML.

## Key Features

1. It should be easy for developers to add `custom transformers` in Scala, Java, Python, C, Rust, or any other language
2. The serialization format should be flexible and meet state-of-the-art performance requirements. This means being able to serialize arbitrarily-large random forest, linear, or neural network models.
3. Serialization should be optimized for ML Transformers and Pipelines as seen in Scikit-learn and Spark, but it should also support non-pipeline based frameworks such as H2O
4. Serialization should be accessible for all environments and platforms, including low-level languages like C, C++ and Rust
5. Provide a common, extensible serialization format for any technology to integrate with via custom transformers or core transformers
6. Serialization/Deserialization should be possible with as many technologies as possible to make the models truly portable between different platforms. ie, we should be able to train a pipeline with Scikit-learn then execute it in Spark.

Contributor guide

No contributing guide indexed for this repository

Research direction

No files, tests, or entry points are named. Start by reviewing the existing PMML, Python pickle, and Spark Parquet/JSON formats, then clarify the proposed Bundle.ML scope and reference implementations in Scala and Rust. Done should mean an agreed, extensible, portable serialization format that supports pipelines, custom transformers, and model portability across the named technologies.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, rust, scala, scikit-learn, spark
Domain
data-engineering, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.