apache / apache/texera

HuggingFace inference operator: uploaded audio is stored as a server temp path that the Python worker refuses to read, so audio upload never works

Open
#6,968 1 comment 0 reactions 1 assignee Claimed by @PG1204 View on GitHub
Dominant language
Scala
Stars
314
Forks
187
Avg merge
1d 21h
Merged PRs (30d)
214

Description

### What happened?

The audio-upload flow is broken end to end:

- The frontend posts the file to `/api/huggingface/upload-audio`, which streams it to `${java.io.tmpdir}/texera-hf-audio` and returns the absolute server path (`HuggingFaceModelResource.scala` ~L171-177). The upload component stores that path in the operator's `audioInput` property (`hugging-face-audio-upload.component.ts` ~L118).
- The generated Python's `_read_audio_input` (`PythonCodegenBase.scala` ~L797-811) deliberately accepts only `data:` and `https` URLs and raises for anything else (filesystem reads were removed as LFI hardening).

Nothing converts the path in between, so every ASR / audio-classification run with an uploaded file fails per-row with "Audio task configuration error: Could not read audio input (ValueError: Unsupported audio input ...)".

The path design is unsound even aside from that: the Python worker need not share a filesystem with the web server in distributed deployments, and the temp file is swept after 1 hour (`AUDIO_TEMP_TTL_MS`) and on JVM exit (`deleteOnExit`), so saved workflows hold dangling references. The leftover file-extension map in `_get_audio_content_type` suggests filesystem support was removed on the Python side without updating the upload flow. Image upload works because that component inlines a `data:image/...` URL instead.

Fix requires a design decision:

- **Option A:** match the image component: read the file in the browser and store a `data:audio/...` URL in `audioInput`. Simple, drops the temp-file endpoints entirely; cost is workflow-JSON bloat (uploads are capped at 25 MiB today).
- **Option B:** store uploads somewhere the worker can reach (e.g. dataset storage) and save a URL that the existing hardened `https` fetch path can consume.

Either way, `/upload-audio`, `/audio-preview`, and the preview component must be updated to match, plus spec coverage for the chosen format.

### How to reproduce?

1. Add the operator with task `automatic-speech-recognition` and upload an audio file in the property panel (the saved value is a server path like `/tmp/texera-hf-audio/hf-audio-XXXX.mp3`).
2. Run the workflow: every row's result column shows "Could not read audio input (ValueError: Unsupported audio input ...)".

### Version/Branch

1.3.0-incubating-SNAPSHOT (main)

### Commit Hash (Optional)

_No response_

### What browsers are you seeing the problem on?

_No response_

### Relevant log output

```shell

```

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.