Support the model Bagel
- Dominant language
- Python
- Stars
- 133k
- Forks
- 15.7k
- Avg merge
- 1d 10h
- Merged PRs (30d)
- 153
Description
### Feature Idea
https://github.com/ByteDance-Seed/Bagel
An open-source multimodal foundation model with 7B active parameters (14B in total), trained on large-scale interleaved multimodal data. BAGEL outperforms current top open-source VLMs such as Qwen2.5-VL and InternVL-2.5 on the standard multimodal understanding leaderboard, and offers text-to-image quality comparable to powerful professional generators like SD3. Additionally, BAGEL demonstrates qualitative results superior to leading open-source models in classic image editing scenarios. More importantly, it extends to free-form visual composition, multi-view synthesis, and world navigation, features that constitute "world modeling" tasks beyond the scope of previous image editing models.

### Existing Solutions
_No response_
### Other
_No response_
Contributor guide
Assessment
This issue has not been assessed yet.