ByteDance-Seed / ByteDance-Seed/Bagel

Bagel-based text-image CoT Framework! Thank you for your great work!❤️

Open
#222 1 comment 5 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
6.2k
Forks
545
PR merge metrics
No merged PRs in 30d

Description

Thank for your great work! We have developed a unified image-text CoT framework based on Bagel, named [UniCoT](https://github.com/Fr0zenCrane/UniCoT).

Uni-CoT (Unified Chain-of-Thought) is an innovative vision-text interleaved reasoning framework that extends the classic Chain-of-Thought (CoT) capability from pure text to the multimodal domain combining text and images, effectively empowering large models to achieve interpretable, efficient, and systematic reasoning. Our core insight stems from the observation that humans naturally rely on grasping visual state changes, such as object motion, spatial interactions, and visual causality, during visual understanding. Therefore, Uni-CoT simulates this human reasoning process by systematically completing complex multimodal reasoning tasks through four steps: task planning, subtask execution, self-verification, and result summarization.
On the reasoning-based image generation benchmark [WISE], our proposed Uni-CoT has successfully outperformed mainstream open-source models such as MetaQuery and Bagel, demonstrating exceptional multimodal reasoning performance. Moreover, Uni-CoT has also achieved significant advantages on JourneyDB and complex human written prompt tasks.

Feel free to learn more about Uni-CoT and try it out:
Project Page: https://sais-fuxi.github.io/projects/uni-cot/
Github Repo: [GitHub - Fr0zenCrane/UniCoT: Uni-CoT: Towards Unified Chain-of-Thought Reasoning Across Text and Vis](https://github.com/Fr0zenCrane/UniCoT)
Preliminary Technical Report: [UniCoT/docs/technical_report.md at main · Fr0zenCrane/UniCoT](https://github.com/Fr0zenCrane/UniCoT/blob/main/docs/technical_report.md)

中文版本:

Uni-CoT: 图文交织思维链新范式
Uni-CoT(Unified Chain-of-Thought)是一种创新的图文推理框架,将经典的思维链(Chain-of-Thought,CoT)能力从纯文本扩展到文本与图像相结合的多模态领域,有效帮助大模型实现可解释、高效且系统化的推理能力。我们的核心洞察在于,人类在进行视觉理解时,天然依赖对视觉状态变化的把握,例如物体的运动、空间的互动和视觉因果关系。因此,Uni-CoT 旨在模拟了这一人类推理过程,并通过任务规划、子任务执行、自我校验、结果汇总四个步骤,有条理地完成复杂的图文交织推理。
在基于推理的图片生成基准(benchmark)[WISE]上,我们提出的Uni-CoT已成功超越MetaQuery、Bagel等主流开源模型,展现出卓越的多模态推理性能。此外,在JourneyDB以及由人类撰写的复杂prompt任务中,Uni-CoT同样取得了显著优势。
欢迎大家即刻了解并试用:
项目主页: https://sais-fuxi.github.io/projects/uni-cot/
Github仓库: [GitHub - Fr0zenCrane/UniCoT: Uni-CoT: Towards Unified Chain-of-Thought Reasoning Across Text and Vis](https://github.com/Fr0zenCrane/UniCoT)
前期技术报告: [UniCoT/docs/technical_report.md at main · Fr0zenCrane/UniCoT](https://github.com/Fr0zenCrane/UniCoT/blob/main/docs/technical_report.md)

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue is an announcement for the external UniCoT project and does not identify a Bagel file, test, or entry point to change. Read the linked project page, repository, and technical report for context, but no completion condition or repository task is defined here.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai, machine-learning
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
5/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.