google-research / google-research/vision_transformer
Correct imagenet top-1 accuracy uploaded pretrained weights of hybrid ViT?
Open
- Dominant language
- Jupyter Notebook
- Stars
- 12.7k
- Forks
- 1.5k
- PR merge metrics
- No merged PRs in 30d
Description
Hi, Thanks so much for the great work!
I tried to restore 'imagenet21k+imagenet2012_R50+ViT-B_16.npz' and got 83.41% imagenet Top-1 accuracy, is that a correct accuracy? I directly resize the test images into 384x384 without a crop, I'm not quite sure whether it is a correct operation.
Contributor guide
Research direction
No file, test, or entry point is named. Start by checking the evaluation and preprocessing used for imagenet21k+imagenet2012_R50+ViT-B_16.npz, especially direct 384x384 resizing versus cropping. Done means confirming whether 83.41% is expected and documenting the correct evaluation procedure.
Written by the indexing model from the issue text.
Assessment
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100