Data Quality Checks/Assessment
Open
affinity integration
enhancement
- Dominant language
- Jupyter Notebook
- Stars
- 2
- Forks
- 1
- PR merge metrics
- No merged PRs in 30d
Description
Example:
- Some datasets have images that are corrupted (for example, can have NaNs)
- Simulated datasets might have boundary "particles" that actually should be ignored
- Files that have artefacts
Contributor guide
Research direction
No files, tests, or entry points are named. Start by locating where datasets and images are loaded or processed, then determine how corrupted images, NaN values, boundary particles, and file artefacts should be identified. Done should mean the agreed data-quality checks are implemented and their results are demonstrated on representative datasets.
Written by the indexing model from the issue text.
Assessment
- Domain
- data, testing-qa
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100