MLMI2-CSSI / MLMI2-CSSI/foundry

Update publishing notebook example dataset

Open
#335 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

documentation and examples good first issue
Dominant language
Python
Stars
88
Forks
17
PR merge metrics
No merged PRs in 30d

Description

As a Publisher, I want to have clear publishing examples that use data that is relevant to me and easy for me to compare my own data to

The current dataset publishing guide notebook uses the iris dataset, which is not only problematic, but also not a very useful guide for people with materials and chemistry datasets.

Assumptions:

  1. The current foundry publishing notebook example can be found here
  2. We have an idea of what dataset we want to replace it with (check in with Ben B)
  3. We anticipate Publishers would want a materials or chemistry related dataset

Acceptance Criteria

  1. Given the publishing notebook, a Publisher can easily understand how the example data relate to the metadata describing that data
  2. When a Publisher runs the notebook, it works without issue
  3. When a Publisher is trying to figure out how to write their own metadata, they can easily reference the metadata in the example notebook for guidance
  4. When a Publisher is trying to run the notebook locally, the downloaded dataset is small and lightweight and does not create bloat

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with examples/publishing-guides/dataset_publishing.ipynb and check in with Ben B about the replacement materials or chemistry dataset. Run the notebook locally and compare the dataset with its metadata. Done means the example is relevant, lightweight, runnable, and clearly demonstrates how publishers should write metadata.

Written by the indexing model from the issue text.

Assessment

Tech stack
jupyter-notebook, python
Domain
data, documentation
Issue type
Documentation
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.