tensorflow / tensorflow/datasets
TriviaQA dataset does not contain information if an example comes from wiki or web
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 4.6k
- Forks
- 1.6k
- Avg merge
- 3h 54m
- Merged PRs (30d)
- 1
Description
TriviaQA has two validation/test datasets: Web and Wikipedia.
Even the official leaderboard reports these seperately, so it's important to be able to report separate numbers:
https://competitions.codalab.org/competitions/17208#results
In tfds these two datasets are combined and there is no way of telling which split a validation example originally came from.
Would it be possible to provide a feature, different splits, or different configs to distinguish these datasets?
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by locating the TFDS TriviaQA dataset implementation and its validation split construction. Check how Web and Wikipedia examples are currently combined, then verify that the resulting feature, splits, or configs distinguish the sources so separate evaluation numbers can be reported.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, tensorflow
- Domain
- data, machine-learning
- Issue type
- Feature
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100