OpenBMB / OpenBMB/MiniCPM-o-Demo

omni full duplex 卡顿 & torch compile

Open
#7 8 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
386
Forks
81
Avg merge
1h 59m
Merged PRs (30d)
3

Description

为什么这个例子说话很流程
https://35.226.63.1:8008/omni

是使用了 flash_attention_2 方式么
目前我是使用sdpa 方式,没有https://35.226.63.1:8008/omni流畅

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the full-duplex behavior at https://35.226.63.1:8008/omni and compare it with the reported SDPA setup. Investigate whether the demo uses flash_attention_2 or torch.compile, then document or isolate the configuration responsible for the smoothness difference; no source files or tests are named in the issue.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
32/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.