facebookresearch / facebookresearch/sam3

Question about the custom PE-L+ backbone in SAM 3

Open
#429 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
11.7k
Forks
1.8k
PR merge metrics
No merged PRs in 30d

Description

I’m interested in the backbone architecture of SAM 3. In the paper, you mentioned using a PE-L+ encoder (450M vision parameters), but I couldn't find this specific version in the original Perception Encoder paper, which only describes a 320M "Large" version.
Could you clarify if this is a custom version optimized for SAM 3? Also, why weren't the larger models (like PE coreG or spatialG) used? Did the custom L+ version perform better than them on the SA-Co benchmark?
Thanks!

Contributor guide

Open the contributing guide

Research direction

Read the SAM 3 paper and the original Perception Encoder paper, then compare their described model sizes and variants. Use the SA-Co benchmark context mentioned in the issue to determine whether PE-L+ is custom and how it compares with PE coreG or spatialG; done means documenting clear answers to these architecture and performance questions.

Written by the indexing model from the issue text.

Assessment

Domain
ai, machine-learning
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
15/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.