facebookresearch / facebookresearch/segment-anything
Decomposed relative position embedding in attention layer of ImageEncoder
- Dominant language
- Jupyter Notebook
- Stars
- 54.9k
- Forks
- 6.4k
- PR merge metrics
- No merged PRs in 30d
Description
I am looking for more information on why this is added in the attention layer of the image encoder, I understand it acts as a residual connection to the original positional embeddings but I can't find mentions of it in the paper, has any ablation been conducted of its effect on VITs? Has in been studied in another publication?
Contributor guide
Research direction
Start by examining the image encoder's attention layer and its decomposed relative position embedding, then compare that implementation with the paper mentioned in the issue. Search the cited literature for related ViT studies or ablations; done means documenting the rationale, relevant publications, and any evidence about its effect.
Written by the indexing model from the issue text.
Assessment
- Domain
- computer-vision, machine-learning
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100