facebookresearch / facebookresearch/perception_models

AudioDecoder imported but not available with documented torchcodec==0.1 install

Open
#120 0 comments 1 reaction 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
2.4k
Forks
162
PR merge metrics
No merged PRs in 30d

Description

## Description

In `core/audio_visual_encoder/transforms.py`, the code imports:

```python
from torchcodec.decoders import AudioDecoder
```

However, the documented install instructions specify:

```bash
pip install torchcodec==0.1 --index-url=https://download.pytorch.org/whl/cu124
```

The issue is that **`torchcodec==0.1` does not provide `AudioDecoder`**.
`AudioDecoder` was introduced in later versions of TorchCodec, so following the current install instructions leads to an `ImportError` at runtime when audio processing is enabled.

## Workaround / Fix

I replaced TorchCodec-based audio loading with `torchaudio`, which is already part of the PyTorch stack and avoids the version mismatch.

In `core/audio_visual_encoder/transforms.py`, I replaced `_load_audio` with:

```python
import torchaudio

def _load_audio(self, path: str):
wav, sr = torchaudio.load(path)

# Convert to mono
if wav.shape[0] > 1:
wav = wav.mean(dim=0, keepdim=True)

# Resample if needed
if sr != self.sampling_rate:
wav = torchaudio.functional.resample(wav, sr, self.sampling_rate)

return wav.contiguous()
```

This restores compatibility with `torchcodec==0.1` while keeping the rest of the audio pipeline unchanged.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.