alibaba / alibaba/AliceMind

Fine-tuning video captions

Open
#76 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
2k
Forks
301
PR merge metrics
No merged PRs in 30d

Description

Thanks for your great work!
I', trying to ,modify the image training code to video captioning fine tuning, but there are somethings that doesn't quite clear to me how to modify like using "answer" parameter in MPLUG model.
Could you please release a train framework for this task?

I'm using `vatex_video_caps_dataset` class to load my dataset.

Thanks!

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by locating the image training code, the MPLUG model handling of the "answer" parameter, and the vatex_video_caps_dataset class. Compare the existing dataset and training paths to determine the missing video-caption fine-tuning pieces. Done means a documented, usable training framework for this task, including the required parameter handling.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
15/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.