facebookresearch / facebookresearch/sam2

Using multiple points for video prediction

Open
#371 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
19.9k
Forks
2.5k
PR merge metrics
No merged PRs in 30d

Description

Hello, when I use multiple points for video prediction, I find that there are always a lot of noisy masks in the first frame, but not in the following frames. Is there any way to improve this situation?
![image](https://github.com/user-attachments/assets/4feb0b09-1e3e-4516-a4b8-85bda86cf15a)
At the same time, the video does not seem to support the application of returning a mask (mulit mask output) for each point?

Contributor guide

Open the contributing guide

Research direction

The issue names no file, test, or entry point; start with the repository's example notebooks and reproduce video prediction using multiple points. Compare the first-frame masks with later frames and check how per-point mask output is handled; done means the requested behavior is reproduced, explained, and its supported outcome is established.

Written by the indexing model from the issue text.

Assessment

Tech stack
jupyter-notebook
Domain
computer-vision, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.