huggingface / huggingface/smollm

Question Regarding SmolVLM2-2.2B-Instruct Fine-tuning

Open
#83 0 comments 1 reaction 0 assignees View on GitHub
Image Video
Dominant language
Python
Stars
3.9k
Forks
315
Avg merge
1m
Merged PRs (30d)
1

Description

Thank you for your excellent work. I’m currently attempting to reproduce the fine-tuning phase of SmolVLM2-2.2B-Instruct, as reported in Table 1 of the SmolVLM paper. However, I’m unsure about the dataset details. Could you clarify whether the fine-tuning used the entire content of the following datasets:

LLaVA-OneVision, M4-Instruct, Mammoth, LLaVA-Video-178K,FineVideo,Video-STAR,VRipt,Vista-400K,MovieChat,ShareGPT4Video

Or were only subsets of these datasets sampled? If it’s the latter, I would greatly appreciate any information on how the sampling was done (like what is the llava-onevision/other).
Additionally, if possible, could you share more details about the fine-tuning procedure (e.g., hyperparameters, epoch, video sampling strategies)?

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the SmolVLM paper's Table 1 and the SmolVLM2-2.2B-Instruct fine-tuning materials referenced by the project. Document whether the listed datasets were used in full or sampled, the sampling details, and the training procedure, including hyperparameters, epochs, and video sampling strategies.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.