facebookresearch / facebookresearch/ImageBind
Confuse about ImageNet1k results
Open
- Dominant language
- Python
- Stars
- 9.1k
- Forks
- 842
- PR merge metrics
- No merged PRs in 30d
Description
Wonderful work!
In Table 2, the top-1 accuray of ImageNet1k is 77.7%, which is higher than CLIP(OpenCLIP) by 2.2%(2.0%). But ImageBind did not train the vision encoder and text encoder, so what make results different or anything I miss?
Contributor guide
Assessment
This issue has not been assessed yet.