OpenEuroLLM / OpenEuroLLM/Taskboard
Evaluating pipelines for audio extraction
Open
@MariaFjodorowa is already working on this.
Since Jul 17, 2026.
WP3
- Dominant language
- No language data
- Stars
- 3
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
Goal
Choose the best audio extraction pipeline based on a toy audio dataset
Description
Our pipeline candidates:
Datasets can be found on LUMI (/scratch/project_465002891/oellm_pdf_audio_extraction):
- EuroParl
- ParlaSpeech
- probably audio files extracted from CC crawls
Deliverable scope
Benchmarking framework, results on the existing datasets
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.