Detect duplicate files
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 82
- Forks
- 21
- Avg merge
- 35m
- Merged PRs (30d)
- 1
Description
A dataset submitted to PhysioNet appears to unintentionally include duplicate files:
https://physionet.org/content/cebsdb/1.0.0/
e.g. search for "3e0ccc9e8960427710b47cfbd2531a20fdb73c247bc9f21795e31fecf7069c8d" in https://physionet.org/content/cebsdb/1.0.0/SHA256SUMS.txt
It might be worth adding a duplicate-file detector to the validator.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by locating the validator and reviewing how it processes submitted dataset files and SHA256SUMS.txt. Use the PhysioNet CEBSDB example and the repeated hash 3e0ccc9e8960427710b47cfbd2531a20fdb73c247bc9f21795e31fecf7069c8d as the reference case. Done means the validator detects duplicate files in a dataset submission.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- tooling
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100