facebookresearch / facebookresearch/jepa
is it possible to learn an action model (or the action's effects) with v-jepa ?
- Dominant language
- Python
- Stars
- 4.1k
- Forks
- 422
- PR merge metrics
- No merged PRs in 30d
Description
Hello,
I would like to know if it is possible to add the knowledge of the actions performed by an agent into the architecture.
From my understanding the unmasked part of the image and the coordinates of the masked parts are given as input to the predictor (which predicts the masked parts). So, as I understand, the prediction predicts static elements (parts of the same image) and not next states.
Would it be possible, instead, to make Jepa to predict next images, given a present image and an action ? Or, can the actual implementation be used to produce representations that would fit in this downstream task (i.e. obtaining the "effects" of an action onto an image) ?
Thanks a lot
Contributor guide
Research direction
The issue names no files, tests, or entry points. Start by reviewing the V-JEPA predictor inputs and architecture to assess whether action-conditioned next-state prediction is supported; the requested outcome is a feasibility decision or a defined design for learning action effects.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100