AnswerDotAI / AnswerDotAI/fsdp_qlora
Is it possible to fine-tune a Vision Language Model (VLM)?
Open
- Dominant language
- Jupyter Notebook
- Stars
- 1.6k
- Forks
- 201
- PR merge metrics
- No merged PRs in 30d
Description
Hi there,
Just wondering, does this repo support fine-tuning a Vision Language Model (VLM), e.g https://huggingface.co/microsoft/Phi-3.5-vision-instruct?
Many thanks for any help, and for this amazing lib!
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue names no repository files, tests, or training entry points; start by inspecting the existing training workflow and the referenced Hugging Face model. Done is a documented, reproducible answer establishing whether this repository supports fine-tuning that VLM.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- huggingface, jupyter-notebook
- Domain
- ai, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100