multiple reads of nasa-nex-gddp-cmip6 dataset from MultiZarrToZarr concatenated metadata returns all nans

未关闭
#289 2 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
3/5
预计耗时
1-2 天
新手友好度
45/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
停滞
技术栈
jupyter-notebook

调研方向

从 datasets/nasa-nex-gddp-cmip6/nasa-nex-gddp-cmip6-example.ipynb 中的单元格 13 和 14 开始,复现两次点时间序列读取。检查 MultiZarrToZarr 拼接元数据的行为,包括重复读取或并行读取。完成的标准是第二次读取返回与第一次相同的有效值,而不是全部为 NaNs。

由索引模型根据 Issue 内容生成。

描述

Simply running twice cells 13 & 14 of the notebook that read and plot point single variable time-series for a point will reproduce this issue where the first run will have the valid values but second will be all nans. I encountered this when parallelizing reading of the files with dask that results in multiple reads and the unexpected result.

主要语言
Jupyter Notebook
星标
455
派生
225
PR 合并指标
30 天内没有已合并 PR

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

microsoft/PlanetaryComputerExamples 的其他 Issue

查看 microsoft/PlanetaryComputerExamples 的全部 Issue

相似的 Issue

更多 Data Engineering Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。