OpenBMB / OpenBMB/MiniCPM-V

MiniCPM-o-4_5 transformer推理"video_url": {"url": video_path, "use_audio": True} use_audio默认都为True不可更改么?

Open
#1,072 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
26.4k
Forks
2.1k
Avg merge
14h 39m
Merged PRs (30d)
3

Description

在接入MiniCPM-o-4_5时,视频推理部分,采用OpenAI的结构推理,content.append({
"type": "video_url",
"video_url": {"url": video_path, "use_audio": False}
}) 使用"use_audio": False 发现未生效,我的视频如果没有音频的时候会报错

answer = self.model.chat(msgs=messages, max_new_tokens=32768, omni_mode=True, use_tts_template=False,
File "/usr/local/conda/lib/python3.10/site-packages/torch/utils/_contextlib.py", line 116, in decorate_context
return func(*args, **kwargs)
File "/home/hadoop-aipnlp/.cache/huggingface/modules/transformers_modules/main/modeling_minicpmo.py", line 1146, in chat
content = normalize_content(content)
File "/home/hadoop-aipnlp/.cache/huggingface/modules/transformers_modules/main/utils.py", line 2400, in normalize_content
normalized = normalize_content_item(item)
File "/home/hadoop-aipnlp/.cache/huggingface/modules/transformers_modules/main/utils.py", line 2338, in normalize_content_item
video_frames, audio_segments, stacked_frames = get_video_frame_audio_segments(
File "/home/hadoop-aipnlp/.local/lib/python3.10/site-packages/minicpmo/utils.py", line 414, in get_video_frame_audio_segments
audio_segments = get_audio_segments(
File "/home/hadoop-aipnlp/.local/lib/python3.10/site-packages/minicpmo/utils.py", line 175, in get_audio_segments
video_clip.audio.write_audiofile(temp_audio_file_path, codec="pcm_s16le", fps=sr)
AttributeError: 'NoneType' object has no attribute 'write_audiofile'

看了下https://huggingface.co/openbmb/MiniCPM-o-4_5/blob/main/utils.py#L2338 好像没有支持use_audio 这个参数的传递?这个目前有其他方法解决么?

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the linked Hugging Face utils.py around line 2338 and compare its video normalization path with minicpmo/utils.py, especially get_video_frame_audio_segments and get_audio_segments. Reproduce the failure with a video without audio and verify whether use_audio=False reaches the audio-extraction decision; done means such input no longer calls write_audiofile on a missing audio stream.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.