C-Haines retention - set limits
- Dominant language
- Python
- Stars
- 65
- Forks
- 11
- Avg merge
- 21h 25m
- Merged PRs (30d)
- 70
Description
I've put this story in To be Decided On so that we can get input from P.O. about what this ticket should look like before we prioritize it. Maybe I did wrong?
**Describe the task**
Currently the C-Haines data retention is unlimited. (We're using an object store, so the only limit we have right now is how much $$ we want to pay). We need to set some limit to it.
**Acceptance Criteria**
- [ ] TBD: Something about how many months/years of data to keep.
- [ ] TBD: Is there some kind of gradual reduction in data? E.g. Only keep one HRDPS model run per day? One GDPS model run per day? One RDPS Model run per day?
- [ ] TBD: Do we keep all the historic models? Or do we decide on only keeping HRDPS, because it's the most high res?
**Additional context**
- We can set an absolute time limit to keep all data, or we can have various more fine grained data retention rules. (Like keeping some model runs longer than others, or not keeping some at all?)
- Implication of not keeping some model runs, is that we can't compare how good the GDPS long term prediction was v.s. the short term RDPS/HRDPS.
Contributor guide
Research direction
No files, tests, or entry points are named. First get product-owner decisions on retention duration, model-run reduction, and which historical models to keep; then locate the C-Haines object-store retention configuration. Done means the agreed retention rules are implemented and their effect on model comparison is addressed.
Written by the indexing model from the issue text.
Assessment
- Domain
- data-engineering
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 15/100