facebookresearch / facebookresearch/sam2
Using multiple points for video prediction
- Dominant language
- Jupyter Notebook
- Stars
- 19.9k
- Forks
- 2.5k
- PR merge metrics
- No merged PRs in 30d
Description
Hello, when I use multiple points for video prediction, I find that there are always a lot of noisy masks in the first frame, but not in the following frames. Is there any way to improve this situation?

At the same time, the video does not seem to support the application of returning a mask (mulit mask output) for each point?
Contributor guide
Research direction
The issue names no file, test, or entry point; start with the repository's example notebooks and reproduce video prediction using multiple points. Compare the first-frame masks with later frames and check how per-point mask output is handled; done means the requested behavior is reproduced, explained, and its supported outcome is established.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- jupyter-notebook
- Domain
- computer-vision, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100