huggingface / huggingface/smollm
Question Regarding SmolVLM2-2.2B-Instruct Fine-tuning
- Dominant language
- Python
- Stars
- 3.9k
- Forks
- 315
- Avg merge
- 1m
- Merged PRs (30d)
- 1
Description
Thank you for your excellent work. I’m currently attempting to reproduce the fine-tuning phase of SmolVLM2-2.2B-Instruct, as reported in Table 1 of the SmolVLM paper. However, I’m unsure about the dataset details. Could you clarify whether the fine-tuning used the entire content of the following datasets:
LLaVA-OneVision, M4-Instruct, Mammoth, LLaVA-Video-178K,FineVideo,Video-STAR,VRipt,Vista-400K,MovieChat,ShareGPT4Video
Or were only subsets of these datasets sampled? If it’s the latter, I would greatly appreciate any information on how the sampling was done (like what is the llava-onevision/other).
Additionally, if possible, could you share more details about the fine-tuning procedure (e.g., hyperparameters, epoch, video sampling strategies)?
Contributor guide
No contributing guide indexed for this repository
Research direction
Start with the SmolVLM paper's Table 1 and the SmolVLM2-2.2B-Instruct fine-tuning materials referenced by the project. Document whether the listed datasets were used in full or sampled, the sampling details, and the training procedure, including hyperparameters, epochs, and video sampling strategies.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100