microsoft / microsoft/PlanetaryComputerExamples
multiple reads of nasa-nex-gddp-cmip6 dataset from MultiZarrToZarr concatenated metadata returns all nans
未关闭
还没有人认领这个 Issue。
- 主要语言
- Jupyter Notebook
- 星标
- 455
- 派生
- 225
- PR 合并指标
- 30 天内没有已合并 PR
描述
Simply running twice cells 13 & 14 of the notebook that read and plot point single variable time-series for a point will reproduce this issue where the first run will have the valid values but second will be all nans. I encountered this when parallelizing reading of the files with dask that results in multiple reads and the unexpected result.
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
调研方向
从 datasets/nasa-nex-gddp-cmip6/nasa-nex-gddp-cmip6-example.ipynb 中的单元格 13 和 14 开始,复现两次点时间序列读取。检查 MultiZarrToZarr 拼接元数据的行为,包括重复读取或并行读取。完成的标准是第二次读取返回与第一次相同的有效值,而不是全部为 NaNs。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- jupyter-notebook
- 领域
- data-engineering
- Issue 类型
- 缺陷
- 难度
- 3/5
- 预计耗时
- 1-2 天
- 活跃度
- 停滞
- 描述清晰度
- 基本清楚
- 新手友好度
- 45/100