facebookresearch / facebookresearch/segment-anything

Decomposed relative position embedding in attention layer of ImageEncoder

Open
#438 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
54.9k
Forks
6.4k
PR merge metrics
No merged PRs in 30d

Description

I am looking for more information on why this is added in the attention layer of the image encoder, I understand it acts as a residual connection to the original positional embeddings but I can't find mentions of it in the paper, has any ablation been conducted of its effect on VITs? Has in been studied in another publication?

Contributor guide

Open the contributing guide

Research direction

Start by examining the image encoder's attention layer and its decomposed relative position embedding, then compare that implementation with the paper mentioned in the issue. Search the cited literature for related ViT studies or ablations; done means documenting the rationale, relevant publications, and any evidence about its effect.

Written by the indexing model from the issue text.

Assessment

Domain
computer-vision, machine-learning
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.