bethgelab / bethgelab/supersanity

Gemini-2.5-Flash on VSC-Repeat

Open
#2 4 comments 1 reaction 0 assignees View on GitHub
Dominant language
Python
Stars
16
Forks
0
PR merge metrics
No merged PRs in 30d

Description

The paper states: “Our findings do not imply that there has been no progress in improving spatial supersensing. Cambrian-S reports that Gemini-2.5-Flash, used as a generic long-context video model, achieves around 41.5% accuracy on 60-minute VSR videos—similar in magnitude to the Cambrian-S model with predictive memory on the same split—despite having no benchmark-specific inference design exposed to users.”

I think an interesting follow-up experiment would be to evaluate Gemini-2.5-Flash on VSC-Repeat to see whether it already shows a preliminary ability to build a stable internal map that survives revisits to the same environment.

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue names no files, tests, or entry points. Start by locating the existing VSC-Repeat evaluation path and determine how Gemini-2.5-Flash could be evaluated; done means reporting its results on VSC-Repeat and comparing them with the cited benchmark context.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
computer-vision, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.