MetOffice / MetOffice/PyPRECIS

Rationalise & organise data for notebooks

Open
#22 2 comments 0 reactions 0 assignees View on GitHub
enhancement paused review needed
Dominant language
Jupyter Notebook
Stars
20
Forks
2
PR merge metrics
No merged PRs in 30d

Description

The data directory associated with the notebooks `/project/precis/worksheets/data` needs rationalising. The aim is to eventually add this to the github repo via git Large File Store (git lfs) so the smaller we can make it the better.

Some things to check:

- [x] Do we need pmsl model data? (Double check it isn't used in any of the exercises)
- [ ] CRU data is the largest single dataset. Can we reduce the size? Do we need to the entire dataset? Can we reduce the area?
- [ ] Would separate directories for variables be better than separate directories for experiments?
- [ ] Improve the data processing workflow so that we don't save rim removed nc versions of each individual pp file. I can't see any good reason for doing this - relates to comments in #23

Contributor guide

Open the contributing guide

Research direction

Audit /project/precis/worksheets/data and the notebook exercises to confirm which datasets are used, starting with the CRU and PMSL data. Review the related workflow discussed in issue #23. Done means the required data is reduced and organised, and the workflow no longer saves unnecessary rim-removed NetCDF files.

Written by the indexing model from the issue text.

Assessment

Tech stack
git, jupyter-notebook, python
Domain
data-engineering
Issue type
Refactor
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.