facebookresearch / facebookresearch/jepa

Question about encoder choice for downstream tasks

Open
#86 2 comments 3 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
4.1k
Forks
424
PR merge metrics
No merged PRs in 30d

Description

Question about encoder choice for downstream tasks

In the video classification evaluation code, I noticed that the target encoder (y-encoder) is being used for downstream tasks instead of the context encoder (x-encoder). This seems different from other self-supervised learning approaches:

  1. Most SSL methods like MOCO, SimCLR, and BYOL use their main/query encoder for downstream tasks rather than the momentum/target encoder.

  2. In V-JEPA, the y-encoder has stop_gradient applied during training, which intuitively suggests the x-encoder might be more suitable for downstream tasks since it learns to predict comprehensive features from partial information.

Looking at the implementation, I noticed the following code for loading checkpoints:

checkpoint = torch.load(pretrained, map_location='cpu')
try:
    pretrained_dict = checkpoint[checkpoint_key]
except Exception:
    pretrained_dict = checkpoint['encoder']

While the code primarily uses the target_encoder (through checkpoint_key), it seems there's a fallback option to use 'encoder'. This suggests that using the context encoder might still be possible, though not prioritized.

I'm curious about:

  1. Was there experimental evidence showing that the target encoder consistently performs better than the context encoder for downstream tasks?
  2. If so, was this the reason for prioritizing the target encoder in the implementation?
  3. Are there specific characteristics of V-JEPA that make the target encoder more suitable for downstream tasks, unlike other SSL approaches?

Would appreciate any insights into this design choice. Thanks!

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the video classification evaluation code around the shown torch.load checkpoint logic, including checkpoint_key, target_encoder, and the encoder fallback. Investigate the rationale and any experimental evidence for the selected encoder, then document how V-JEPA's encoders relate to downstream-task performance and what evidence supports the choice.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.