facebookresearch / facebookresearch/seamless_interaction

Dataset Missing For Some Nauturalistic Dev/Test

Open
#25 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
413
Forks
31
PR merge metrics
No merged PRs in 30d

Description

### Have you read the Contributing Guidelines?

- [x] I have read the [Contributing Guidelines](https://github.com/facebookresearch/seamless_interaction/blob/main/CONTRIBUTING.md).

### Issue Type

Missing files or data

### Dataset Label

naturalistic

### Dataset Split

dev and test

### Affected Files/Data and Issue Description

After downloading the whole Dataset using offical script, I found that some archieves needed in the assets/filelist.cvs were missing. To be specific, batch 0 in Dev and Test of Naturalistic. The details can be seen in the picture, these archieves are indeed not in the Hugging Face(I've manually checked).

Image

Also, I have some related questions. (1) For many batches among, the last archieve of batch in HuggingFce is typically smaller than others and why they are not been downloaded? (2)And For improvised/dev/0000 batch, it has some extra archieves in the HuggingFace and why are they not beed download and used? (3) Finally, the extra folder in some splits are also not been used, I think a thorough explaination will be a benefit for people to use this dataset. Thanks!!

### Steps to Reproduce

Download the whole dataset using the offical script, download_whole_dataset() in scripts/download_hf.py specifically.

### Additional Context

_No response_

### Self-service

- [ ] I'd be willing to help investigate this data issue. Add comment

Contributor guide

Open the contributing guide

Research direction

Start with scripts/download_hf.py and assets/filelist.cvs, then reproduce the full download for the naturalistic dev and test splits. Check batch 0 against the files listed in the issue and compare the other noted archive and folder differences. Done means determining whether the missing archives are an issue and documenting the observed download behavior and dataset contents.

Written by the indexing model from the issue text.

Assessment

Tech stack
huggingface, python
Domain
data
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.