MLPipeline does not preserve metadata from JSON pipeline annotation
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 124
- Forks
- 33
- PR merge metrics
- No merged PRs in 30d
Description
- MLBlocks version: 0.4
- Python version: 3.8
Description
I'm trying to get metadata from an MLPipeline object that was present in the JSON pipeline annotation that was loaded. For example, the annotation has a metadata.name field that I'd like to access from the pipeline.
What I Did
In the following example, I would expect that the MLPipeline has a metadata dict which has name key, just like the JSON. But it doesn't.
$ mkdir -p mlprimitives mlpipelines
$ curl -s https://raw.githubusercontent.com/MLBazaar/MLPrimitives/master/mlprimitives/primitives/sklearn.ensemble.RandomForestRegressor.json -o mlprimitives/sklearn.ensemble.RandomForestRegressor.json
$ curl -s https://raw.githubusercontent.com/MLBazaar/MLPrimitives/master/mlprimitives/pipelines/sklearn.ensemble.RandomForestRegressor.json -o mlpipelines/sklearn.ensemble.RandomForestRegressor.json
$ jq .metadata.name mlpipelines/sklearn.ensemble.RandomForestRegressor.json
"sklearn.ensemble.RandomForestRegressor"
$ python
Python 3.8.3 (default, Jul 20 2020, 16:43:14)
[Clang 11.0.3 (clang-1103.0.32.62)] on darwin
Type "help", "copyright", "credits" or "license" for more information.
>>> from mlblocks import load_pipeline, MLPipeline
>>> pipeline = MLPipeline(load_pipeline('sklearn.ensemble.RandomForestRegressor'))
>>> pipeline.metadata
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
AttributeError: 'MLPipeline' object has no attribute 'metadata'
Suggestions
- Explicitly support persisting metadata on the
MLPipelineobject (and presumably on underlyingMLBlockobjects) - Raise an error of the JSON input contains unused (unsupported) keys
- Guarantee that
MLPipeline.loadandMLPipeline.saveare inverse operations, i.e. that no data is lost (currently metadata fields are lost)
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the MLPipeline and load_pipeline entry points, then trace how the JSON pipeline annotation is loaded and how MLPipeline.load and MLPipeline.save handle data. Compare the annotation's metadata.name with the resulting object and saved representation; done requires an agreed behavior for preserving metadata or rejecting unsupported keys.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100