[Dataset] Add Inner Speech dataset
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 1.1k
- Forks
- 264
- Avg merge
- 1d 13m
- Merged PRs (30d)
- 23
Description
Dataset Information
- Suggested Name for MOABB: Nieto2022
- Full Title: Thinking out loud, an open-access EEG-based BCI dataset for inner speech recognition
- Paradigm: Inner Speech / Imagined Speech
- Category: Could be a new paradigm for MOABB
Link Status
✅ Data Link: https://openneuro.org/datasets/ds003626/versions/2.1.2 (Working)
✅ Paper: https://www.nature.com/articles/s41597-022-01147-2 (Working)
✅ GitHub: https://github.com/N-Nieto/Inner_Speech_Dataset (Working)
Publication Details
| Field | Value |
|---|---|
| Journal | Scientific Data (Nature) |
| Published | February 14, 2022 |
| DOI | 10.1038/s41597-022-01147-2 |
| Authors | Nicolas Nieto, Victoria Peterson, Hugo Leonardo Rufiner, Juan Esteban Kamienkowski, Ruben Spies |
Technical Specifications
| Parameter | Value |
|---|---|
| Number of Subjects | 10 (4 female, 6 male; mean age 34 years) |
| Number of Channels | 128 active EEG + 8 external (EOG/EMG) = 136 total |
| Sampling Frequency | 1024 Hz (original) / 254 Hz (processed) |
| Sessions per Subject | 3 (recorded in one day) |
| Total Trials | 5,640 |
| Trial Duration | 4.5 seconds |
| Recording Duration | >9 hours continuous EEG data |
| Data Format | Raw (BDF files) + Processed (MNE FIF files, epoched) |
| BIDS Format | ✅ Yes |
| Equipment | BioSemi ActiveTwo (24-bit) |
Experimental Design
| Condition | Trials |
|---|---|
| Inner Speech | 2,236 |
| Pronounced Speech | 1,128 |
| Visualized Condition | 2,276 |
Classes: 4 directional Spanish words - "Arriba" (up), "Abajo" (down), "Derecha" (right), "Izquierda" (left)
Description
This dataset was created to advance BCI research targeting the inner speech paradigm. Inner speech refers to the phenomenon of an "inner voice" - the ability to think words without producing audible speech. This paradigm enables the possibility of controlling external devices simply by thinking about commands. The dataset consists of EEG recordings from 10 naive BCI users performing four mental tasks under three different conditions.
Original Suggestion
Created as part of dataset discovery initiative.
Related
Suggested by @vmcru in this comment
This issue is a sub-issue of #1 (Discover new datasets)
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The issue names no repository files, tests, or entry points. Start by reviewing existing dataset integrations and their tests, then use the linked OpenNeuro dataset and paper to define the integration; done means the Nieto2022 dataset is available through MOABB with coverage matching project conventions.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- data, machine-learning
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100