Comfy-Org / Comfy-Org/ComfyUI

Support the model Bagel

Open
#8,233 1 comment 12 reactions 0 assignees View on GitHub
Feature
Dominant language
Python
Stars
133k
Forks
15.7k
Avg merge
1d 10h
Merged PRs (30d)
153

Description

### Feature Idea

https://github.com/ByteDance-Seed/Bagel

An open-source multimodal foundation model with 7B active parameters (14B in total), trained on large-scale interleaved multimodal data. BAGEL outperforms current top open-source VLMs such as Qwen2.5-VL and InternVL-2.5 on the standard multimodal understanding leaderboard, and offers text-to-image quality comparable to powerful professional generators like SD3. Additionally, BAGEL demonstrates qualitative results superior to leading open-source models in classic image editing scenarios. More importantly, it extends to free-form visual composition, multi-view synthesis, and world navigation, features that constitute "world modeling" tasks beyond the scope of previous image editing models.

![Image](https://github.com/user-attachments/assets/f641768a-e8d8-4260-bc10-0ca061a57cde)

### Existing Solutions

_No response_

### Other

_No response_

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.