huggingface / huggingface/transformers

An example for finetuning FLAVA or any VLP multimodel using trainer (for example for classification)

Open
#18,066 4 comments 0 reactions 0 assignees View on GitHub
Examples Feature request
Dominant language
Python
Stars
166k
Forks
34.6k
Avg merge
3d 9h
Merged PRs (30d)
281

Description

### Feature request

There is no example of finetuning any VLP model using trainer. I would appreciate an example

### Motivation

The way to use trainers with any Vision and Language pretrained model is not clear.

### Your contribution

None.

Contributor guide

Open the contributing guide

Research direction

Start by reviewing the existing Trainer examples and the FLAVA model entry points to determine how a vision-language classification fine-tuning example should be structured. Done means adding a clear, runnable example that demonstrates fine-tuning a VLP model with Trainer and explains the required data and configuration.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.