huggingface / huggingface/transformers
An example for finetuning FLAVA or any VLP multimodel using trainer (for example for classification)
Open
Examples
Feature request
- Dominant language
- Python
- Stars
- 166k
- Forks
- 34.6k
- Avg merge
- 3d 9h
- Merged PRs (30d)
- 281
Description
### Feature request
There is no example of finetuning any VLP model using trainer. I would appreciate an example
### Motivation
The way to use trainers with any Vision and Language pretrained model is not clear.
### Your contribution
None.
Contributor guide
Research direction
Start by reviewing the existing Trainer examples and the FLAVA model entry points to determine how a vision-language classification fine-tuning example should be structured. Done means adding a clear, runnable example that demonstrates fine-tuning a VLP model with Trainer and explains the required data and configuration.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100