google-research / google-research/vision_transformer

Correct imagenet top-1 accuracy uploaded pretrained weights of hybrid ViT?

Open
#39 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
12.7k
Forks
1.5k
PR merge metrics
No merged PRs in 30d

Description

Hi, Thanks so much for the great work!
I tried to restore 'imagenet21k+imagenet2012_R50+ViT-B_16.npz' and got 83.41% imagenet Top-1 accuracy, is that a correct accuracy? I directly resize the test images into 384x384 without a crop, I'm not quite sure whether it is a correct operation.

Contributor guide

Open the contributing guide

Research direction

No file, test, or entry point is named. Start by checking the evaluation and preprocessing used for imagenet21k+imagenet2012_R50+ViT-B_16.npz, especially direct 384x384 resizing versus cropping. Done means confirming whether 83.41% is expected and documenting the correct evaluation procedure.

Written by the indexing model from the issue text.

Assessment

Domain
machine-learning
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.