JuliaDataCubes / JuliaDataCubes/YAXArrays.jl
YAXArrays seems to download too much data
Nobody has claimed this yet.
- Dominant language
- Julia
- Stars
- 132
- Forks
- 25
- PR merge metrics
- No merged PRs in 30d
Description
I'm trying the example from the docs:
using Zarr, YAXArrays, Dates, DimensionalData
store = "gs://cmip6/CMIP6/ScenarioMIP/DKRZ/MPI-ESM1-2-HR/ssp585/r1i1p1f1/3hr/tas/gn/v20190710/"
g = open_dataset(zopen(store, consolidated=true))
c = g["tas"]
ct = c[Ti=At(Date("2018-08-01"):Day(10):Date("2050-08-01"))]
in_memory = ct.data[:, :, :]
This takes reaally long and fills up all my RAM (32gb).
A few infos:
The selected slice:
Download speed of the julia process
I was expecting it to only download the 328mb, but from the download speed and RAM usage I suspect it's downloading much more data, making it almost impossible to download this part of the dataset...
Am I missing something or is this a bug, or just a limitation of the package?
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reproducing the documented Julia example using open_dataset, zopen, and the ct.data[:, :, :] slice against the linked CMIP6 store. Compare the selected slice size with the data actually downloaded and memory used; done means determining whether this is expected behavior, a package bug, or a documented limitation.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- julia
- Domain
- data, performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100