NASA-IMPACT / NASA-IMPACT/veda-data

Add a OCO-3 SAM Data data layer (high-level steps)

Open
#77 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Jupyter Notebook
Stars
9
Forks
1
Avg merge
1d 9m
Merged PRs (30d)
1

Description

For each dataset, we will follow the following steps:

  1. Identify dataset and where it will be accessed from. Check it's a good source with science team. Ask about specific variables and required spatial and temporal extent. Note most datasets will require back processing (e.g. generating cloud-optimized data for historical data).
  2. If the dataset is ongoing (i.e. new files are continuously added and should be included in the dashboard), design and construct the forward-processing workflow.
    • Each collection will have a workflow which includes discovering data files from the source, generating the cloud-optimized versions of the data and writing STAC metadata.
    • Each collection will have different requirements for both the generation and scheduling of these steps, so a design step much be included for each new collection / data layer.
  3. Verify the COG output with the science team by sharing in a visual interface.
  4. Verify the metadata output with STAC API developers and any systems which may be depending on this STAC metadata (e.g. the front-end dev team).
  5. If the dataset should be backfilled, create and monitor the backward-processing workflow.
  6. Engage the science team to add any required background information on the methodology used to derive the dataset.
  7. Add the dataset to the production dashboard.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No files, tests, or entry points are named. Start by identifying the existing dataset-layer and workflow conventions, then confirm the OCO-3 SAM source, variables, spatial and temporal scope with the science team. Done means the processing, COG and STAC metadata verification, required backfill, background information, and production dashboard integration are complete.

Written by the indexing model from the issue text.

Assessment

Domain
data-engineering
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.