OpenGVLab / OpenGVLab/VideoChat-Flash
[Q] InternVideo 2.5,, TPO
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 530
- Forks
- 20
- PR merge metrics
- No merged PRs in 30d
Description
Hello, thank you for the excellent research and for sharing the code.
I understand that TPO (Task Preference Optimization) has been applied to InternVideo 2.5, and as mentioned in the related paper, it includes three task-specific heads: region, temporal, and mask.
I have two questions regarding this:
- Are these three heads already implemented and integrated into the current InternVideo 2.5 codebase?
- The paper describes a detailed multi-stage training process, but the repository currently provides only inference scripts. Will the training scripts for these heads be released in the future? Alternatively, is there any guidance or reference available to perform supervised fine-tuning (sFT) with these task heads?
Any support or clarification would be greatly appreciated. Thank you again for your valuable contribution!
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
No specific file, test, or entry point is named. Review the repository's existing inference scripts and the paper's TPO description to determine whether the three task heads and training process are present; done would require maintainer clarification or documented guidance.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100