amzn / amzn/explainable-text-vqa

How to map the textVQA-X to textVQA dataset?

Open
#1 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
No language data
Stars
11
Forks
3
PR merge metrics
No merged PRs in 30d

Description

Dear Authors,
Thank you for sharing the dataset, could I ask how is the mapping between the id in textVQA-X about textVQA dataset?
It is a little bit confusing.
In the seg folder, it is indexed from 0 to 19999,what is that index?
And what are train_id and val_id? Are they question ids? Why couldn't I find the question_id in val_id.txt in the TextVQA_0.5.1_val.json?
Or is that you create visual grounding for the 0-19999 images in the textVQA training set and then separate it into train/val set?
Many thanks in advance!

Contributor guide

Open the contributing guide

Research direction

Start by comparing the seg folder with train_id, val_id.txt, and TextVQA_0.5.1_val.json, focusing on how the 0–19999 indexes relate to image or question identifiers. Document the mapping, explain train_id and val_id, and clarify which TextVQA split the visual grounding data uses.

Written by the indexing model from the issue text.

Assessment

Domain
data
Issue type
Documentation
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.