facebookresearch / facebookresearch/sam3
Question about the custom PE-L+ backbone in SAM 3
- Dominant language
- Python
- Stars
- 11.7k
- Forks
- 1.8k
- PR merge metrics
- No merged PRs in 30d
Description
I’m interested in the backbone architecture of SAM 3. In the paper, you mentioned using a PE-L+ encoder (450M vision parameters), but I couldn't find this specific version in the original Perception Encoder paper, which only describes a 320M "Large" version.
Could you clarify if this is a custom version optimized for SAM 3? Also, why weren't the larger models (like PE coreG or spatialG) used? Did the custom L+ version perform better than them on the SA-Co benchmark?
Thanks!
Contributor guide
Research direction
Read the SAM 3 paper and the original Perception Encoder paper, then compare their described model sizes and variants. Use the SA-Co benchmark context mentioned in the issue to determine whether PE-L+ is custom and how it compares with PE coreG or spatialG; done means documenting clear answers to these architecture and performance questions.
Written by the indexing model from the issue text.
Assessment
- Domain
- ai, machine-learning
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 15/100