ml-explore / ml-explore/mlx-examples

[whisper] `mlx_whisper` CLI overwrites output file when multiple audio inputs are provided

Open Beginner friendly
#1,441 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
9k
Forks
1.2k
PR merge metrics
No merged PRs in 30d

Description

Describe the bug

When passing multiple audio files to the mlx_whisper CLI (e.g., mlx_whisper file1.mp3 file2.mp3), all transcriptions are written to the output file of the first input (file1.txt), sequentially overwriting it.

To Reproduce
  1. Run mlx_whisper file1.mp3 file2.mp3
  2. Check the output directory.
  3. Observe that only file1.txt exists and contains the transcription of file2.mp3 (the last processed file). file2.txt is not created.
Expected behavior

Each audio input should produce its own output file (file1.txt, file2.txt, etc.).

Root Cause

In whisper/mlx_whisper/cli.py, output_name is mutated inside the audio loop:

output_name: str = args.pop("output_name")
...
for audio_obj in args.pop("audio"):
    if audio_obj == "-":
        audio_obj = audio.load_audio(from_stdin=True)

        output_name = output_name or "content"
    else:
        output_name = output_name or pathlib.Path(audio_obj).stem
    try:
        result = transcribe(...)
        writer(result, output_name, **writer_args)

On the first iteration, output_name is reassigned to "file1". In subsequent iterations, output_name or pathlib.Path(...) evaluates to "file1", causing all remaining audio files to be written to file1.txt.

Proposed Fix

Use a local variable (e.g., file_output_name) inside the loop rather than reassigning output_name:

for audio_obj in args.pop("audio"):
    if audio_obj == "-":
        audio_obj = audio.load_audio(from_stdin=True)

        file_output_name = output_name or "content"
    else:
        file_output_name = output_name or pathlib.Path(audio_obj).stem
    try:
        result = transcribe(
            audio_obj,
            path_or_hf_repo=path_or_hf_repo,
            **args,
        )
        writer(result, file_output_name, **writer_args)

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in whisper/mlx_whisper/cli.py and reproduce the issue with mlx_whisper file1.mp3 file2.mp3. Inspect how the output name is selected inside the audio loop, then verify that each input produces its own output file, such as file1.txt and file2.txt.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
cli
Issue type
Bug
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Quiet
Clarity
Clearly specified
Newbie friendliness
85/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.