huggingface / huggingface/optimum-executorch

Integrate input preprocessing into export flow

Open
#152 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
141
Forks
48
Avg merge
26m
Merged PRs (30d)
2

Description

When we export multimodal LLMs we almost always have to call `processor.apply_chat_template` in HF transformers. That call under the hood will preprocess the multimodal inputs.

This is quite annoying during exporting since we can't export `processor.apply_chat_template` directly. One example is this logic: https://github.com/huggingface/transformers/blob/main/src/transformers/models/whisper/processing_whisper.py#L69 it uses a lot of numpy.

Ideally we should ask transformers to write standard processor for all inputs in torch (tracked by https://github.com/huggingface/transformers/issues/40986). In the short term I think optimum-executorch should host some of the common processors like the one in whisper.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.