tensorflow / tensorflow/datasets

TriviaQA dataset does not contain information if an example comes from wiki or web

Open
#3,506 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

contributions welcome enhancement
Dominant language
Python
Stars
4.6k
Forks
1.6k
Avg merge
3h 54m
Merged PRs (30d)
1

Description

TriviaQA has two validation/test datasets: Web and Wikipedia.

Even the official leaderboard reports these seperately, so it's important to be able to report separate numbers:
https://competitions.codalab.org/competitions/17208#results

In tfds these two datasets are combined and there is no way of telling which split a validation example originally came from.

Would it be possible to provide a feature, different splits, or different configs to distinguish these datasets?

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by locating the TFDS TriviaQA dataset implementation and its validation split construction. Check how Web and Wikipedia examples are currently combined, then verify that the resulting feature, splits, or configs distinguish the sources so separate evaluation numbers can be reported.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, tensorflow
Domain
data, machine-learning
Issue type
Feature
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.