tensorflow / tensorflow/datasets
[data request] Ecoset
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 4.6k
- Forks
- 1.6k
- Avg merge
- 3h 54m
- Merged PRs (30d)
- 1
Description
- Name of dataset: Ecoset
- URL of dataset: https://www.kietzmannlab.org/ecoset/
- Dataset repository: https://codeocean.com/capsule/9570390/tree/v1
- License of dataset: Creative Commons Attribution-NonCommercial-ShareAlike 2.0 license (cc-by-nc-sa-2.0)
- Short description of dataset and use case(s):
Ecoset is a large (155GB) image recognition dataset, combining images of objects with appropriate labels (one label per image). Ecoset is intended to provide higher ecological validity than similar image recognition datasets such as ImageNet (ILSVRC). Ecoset contains 1.5 million images from 565 basic level categories, chosen to be both (i) frequent in linguistic usage, and (ii) rated by human observers as concrete (e.g. ‘table’ is concrete, ‘romance’ is not). Additionally, ecoset was tested for label validity with a mislabelling error rate < 5%, as well as filtered to exclude NSFW content.
For more information on the dataset, consider reading the original publication.
⚠️ Folks who would also like to see this dataset in tensorflow/datasets, please thumbs-up so the developers can know which requests to prioritize. ⚠️
And if you'd like to contribute the dataset (thank you!), see our guide to adding a dataset.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with docs/add_dataset.md, then review the Ecoset dataset page, repository, and original publication linked in the issue. Follow the contribution guide to add Ecoset to tensorflow/datasets; done means the dataset contribution meets that guide’s requirements.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, tensorflow
- Domain
- data, machine-learning
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100