TEN-framework / TEN-framework/ten-framework
[FEATURE] Can ten-agent do video chat with a normal multi-modal LLM?
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 11.1k
- Forks
- 1.4k
- Avg merge
- 2d 15m
- Merged PRs (30d)
- 22
Description
Description
Thanks for the exciting work for a fancy realtime framework.
But I wonder is there anyway to chat with the bot and make it 'understand' via the camera with a normal VLM, qwen2-vl for example, instead of Gemma-v2v or GPT4 v2v?
Severity
Minor
Additional Information
No response
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The issue names no files, tests, or entry points. Start by locating the TEN-agent camera/video integration and model adapters, then determine whether a normal multimodal LLM such as qwen2-vl can support video chat; the work is complete when that supported path is clearly defined and verified.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- ai, audio-video-rtc
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100