developmentseed / developmentseed/chabud2023

Handle undersampling due to lots of images without burned areas

Open
#12 0 comments 0 reactions 0 assignees View on GitHub
help wanted
Dominant language
Python
Stars
7
Forks
1
PR merge metrics
No merged PRs in 30d

Description

The extra Sentinel-2 imagery dataset provided in https://huggingface.co/datasets/chabud-team/chabud-extra does not contain any burned areas according to https://huggingface.co/datasets/chabud-team/chabud-extra/discussions/1. If we include these datasets in the training, there will be a severe imbalance in the ratio of burned area to unburned area pixels.

Some potential ways to handle the extra data to improve model performance:
- [ ] Loss functions that handle foreground/background classes properly
- [ ] Focal Loss
- [ ] Dice Loss
- [ ] Self-supervised pre-training
- [ ] Develop pretext tasks that make use of the extra data, generate useful embeddings on all the given data, and then fine-tune on images with burned areas only

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.