基于wan2.1,你们有没有可能实现这个 ?
- Dominant language
- Python
- Stars
- 357
- Forks
- 17
- PR merge metrics
- No merged PRs in 30d
Description
wan2.1可以做这么多玩法和控制,那么,你们一定能有技术实现让人物唇形与音频同步了。
以下是腾讯的数字人技术,它可以做多人唱歌,多人相声等丰富的场景和效果。
但缺点是占用98G显存,推理速度极慢。
https://hunyuanvideo-avatar.github.io/
如果字节团队基于 wan2.1 ,来做这个类似的(图生视频)数字人,是否更牛X ?或许更惊艳~
(毕竟 wan2.1 模型推理性能和质量都很好,再配合音频驱动)
有一个类似的,但感觉达不到腾讯那种水准,(供参考)。
https://fantasy-amap.github.io/fantasy-talking/
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reading the ATI implementation and comparing its controllable video-generation scope with the linked HunyuanVideo Avatar and Fantasy Talking references. Define whether the requested result covers audio-driven lip synchronization, image-to-video generation, and multi-person singing or dialogue; the issue names no files, tests, or acceptance criteria.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- machine-learning, python
- Domain
- audio-video-rtc, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 15/100