huggingface / huggingface/datasets

CommonSenseQA has missing and inconsistent field names

Open
#4,275 1 comment 0 reactions 1 assignee Claimed by @albertvillanova View on GitHub
dataset bug
Dominant language
Python
Stars
22k
Forks
3.4k
Avg merge
5d 7h
Merged PRs (30d)
17

Description

## Describe the bug
In short, CommonSenseQA implementation is inconsistent with the original dataset.

More precisely, we need to:

1. Add the dataset matching "id" field. The current dataset, instead, regenerates monotonically increasing id.
2. The [“question”][“stem”] field is flattened into "question". We should match the original dataset and unflatten it
3. Add the missing "question_concept" field in the question tree node
4. Anything else? Go over the data structure of the newly repaired CommonSenseQA and make sure it matches the original

## Expected results
Every data item of the CommonSenseQA should structurally and data-wise match the original CommonSenseQA dataset.

## Actual results
TBD

## Environment info
- `datasets` version: 2.1.0
- Platform: macOS-10.15.7-x86_64-i386-64bit
- Python version: 3.8.13
- PyArrow version: 7.0.0
- Pandas version: 1.4.2

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.