bethgelab / bethgelab/supersanity
Gemini-2.5-Flash on VSC-Repeat
- Dominant language
- Python
- Stars
- 16
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
The paper states: “Our findings do not imply that there has been no progress in improving spatial supersensing. Cambrian-S reports that Gemini-2.5-Flash, used as a generic long-context video model, achieves around 41.5% accuracy on 60-minute VSR videos—similar in magnitude to the Cambrian-S model with predictive memory on the same split—despite having no benchmark-specific inference design exposed to users.”
I think an interesting follow-up experiment would be to evaluate Gemini-2.5-Flash on VSC-Repeat to see whether it already shows a preliminary ability to build a stable internal map that survives revisits to the same environment.
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue names no files, tests, or entry points. Start by locating the existing VSC-Repeat evaluation path and determine how Gemini-2.5-Flash could be evaluated; done means reporting its results on VSC-Repeat and comparing them with the cited benchmark context.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- computer-vision, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100