activeloopai / activeloopai/deeplake
[FEATURE] Detect duplicate samples when adding new data to tensors (images)
Open
enhancement
- Dominant language
- C++
- Stars
- 9.2k
- Forks
- 722
- PR merge metrics
- No merged PRs in 30d
Description
## 🚨🚨 Feature Request
- A new implementation (Improvement, Extension)
Is it possible to discard samples in case they are already present in the dataset? If not, would this be something interesting to implement? I feel like this would make the dataset extension pipeline much easier to use and implement
Contributor guide
Assessment
This issue has not been assessed yet.