facebookresearch / facebookresearch/segment-anything

Seemingly missing reshaping operation in point prompt encoding in PromptEncoder

Open
#786 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
54.9k
Forks
6.4k
PR merge metrics
No merged PRs in 30d

Description

https://github.com/facebookresearch/segment-anything/blob/main/segment_anything/modeling/prompt_encoder.py#L81-L85

If I read it correctly, the shape of points input is `[B, N, 2]`, where B is the batch size and N is the number of points per image. The padding ensures that the point prompt also contains the 2d coordinates of two points to make it compatible with the box prompt. Without reshaping operation before the `torch.cat` operation, wouldn't the shape become `[B, N + 1, 2]` after the padding. This doesn't feel right. Since this PromptEncoder is used in the SAM2 as well, it seems to impact both models.

Please correct me if I misunderstand any part of this.

Thank you!

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.