huggingface / huggingface/optimum-executorch

Enable multimodal vision models

Open
#148 3 comments 0 reactions 2 assignees Claimed by @larryliu0820 View on GitHub
Dominant language
Python
Stars
141
Forks
48
Avg merge
26m
Merged PRs (30d)
2

Description

## Vision model enablement list
- [ ] Aria
- [ ] Aya
- [ ] Blip2
- [ ] Chameleon
- [ ] Cohere2 Vision
- [ ] Deepseek VL (derived from Janus)
- [ ] Deepseek VL Hybrid - same as Deepseek VL
- [ ] Emu3
- [ ] Evolla - N/A, for proteins (?)
- [ ] Florence2 (this does image features -> seq2seq instead of decoder)
- [ ] Fuyu
- [x] Gemma3 - already enabled on Optimum ET
- [ ] Gemma3n
- [ ] Git - not image-text-to-text
- [ ] GLM 4.1V
- [ ] GLM 4.1V MOE - same as GLM 4.1V
- [ ] GOT-OCR2
- [ ] IDEFICS 2
- [ ] IDEFICS 3
- [ ] InstructBLIP
- [ ] InternVL
- [ ] Llama4
- [ ] LLaVA
- [ ] LLaVA-NeXT
- [ ] LLaVA-NeXT-Video
- [ ] LLaVA-OneVision
- [ ] Mistral 3
- [ ] Ovis2
- [ ] PaliGemma
- [ ] PerceptionLM
- [ ] Pix2Struct - not image-text-to-text
- [ ] Pixtral - same as LLaVA
- [ ] Qwen2.5-VL
- [ ] Qwen2-VL
- [ ] Qwen3-VL
- [ ] Qwen3-VL MOE - same as Qwen3-VL
- [ ] #161
- [ ] ViP-LLaVA

## Blocked
- [ ] LFM2-VL - data dependent code in short_conv preventing export
- [ ] apple/FastVLM - model not on Transformers [yet](https://github.com/huggingface/transformers/issues/38765)

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.