facebookresearch / facebookresearch/VLM3
Question about joint multitask training
- Dominant language
- Jupyter Notebook
- Stars
- 426
- Forks
- 14
- PR merge metrics
- No merged PRs in 30d
Description
Great work!
I noticed that 4 separate checkpoints are released, one for each task: depth estimation, object-level 3D understanding, pixel correspondence, and camera pose estimation.
Since a central theme of the paper is that VLM3 enables a generalist VLM to learn diverse 3D tasks, I was wondering why the team chose to train and release separate task-specific models rather than SFT a single model jointly on all four tasks.
Was this simply an experimental choice?
Also, do you expect a single jointly trained VLM³ model to perform as well across the four tasks?
Thanks!
Contributor guide
Research direction
Read the paper and the descriptions of the four released task-specific checkpoints first. Document why separate models were chosen and whether joint SFT is expected to match their performance across all four tasks.
Written by the indexing model from the issue text.
Assessment
- Domain
- computer-vision, machine-learning
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100