facebookresearch / facebookresearch/VLM3

Question about joint multitask training

Open
#5 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
426
Forks
14
PR merge metrics
No merged PRs in 30d

Description

Great work!

I noticed that 4 separate checkpoints are released, one for each task: depth estimation, object-level 3D understanding, pixel correspondence, and camera pose estimation.

Since a central theme of the paper is that VLM3 enables a generalist VLM to learn diverse 3D tasks, I was wondering why the team chose to train and release separate task-specific models rather than SFT a single model jointly on all four tasks.

Was this simply an experimental choice?
Also, do you expect a single jointly trained VLM³ model to perform as well across the four tasks?

Thanks!

Contributor guide

Open the contributing guide

Research direction

Read the paper and the descriptions of the four released task-specific checkpoints first. Document why separate models were chosen and whether joint SFT is expected to match their performance across all four tasks.

Written by the indexing model from the issue text.

Assessment

Domain
computer-vision, machine-learning
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.