huggingface / huggingface/datasets

Audio dataset is not decoding on 4.1.1

Open
#7,798 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
22k
Forks
3.4k
Avg merge
5d 7h
Merged PRs (30d)
17

Description

### Describe the bug

The audio column remain as non-decoded objects even when accessing them.

```python
dataset = load_dataset("MrDragonFox/Elise", split = "train")
dataset[0] # see that it doesn't show 'array' etc...
```

Works fine with `datasets==3.6.0`

Followed the docs in

- https://huggingface.co/docs/datasets/en/audio_load

### Steps to reproduce the bug

```python
dataset = load_dataset("MrDragonFox/Elise", split = "train")
dataset[0] # see that it doesn't show 'array' etc...
```

### Expected behavior

It should decode when accessing the elemenet

### Environment info

4.1.1
ubuntu 22.04

Related

- https://github.com/huggingface/datasets/issues/7707

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.