huggingface / huggingface/pytorch-image-models
[FEATURE] Add ViT weights: RADIO
- Dominant language
- Python
- Stars
- 37.1k
- Forks
- 5.2k
- Avg merge
- 1d 13h
- Merged PRs (30d)
- 34
Description
https://github.com/NVlabs/RADIO
The code and model weights of paper *[CVPR 2024] AM-RADIO: Agglomerative Vision Foundation Model - Reduce All Domains Into One* has been released by Nvidia
> RADIO , a new vision foundation model (actually a new vit pretrained weight), excels across visual domains, serving as a superior replacement for vision backbones. Integrating CLIP variants, DINOv2, and SAM through distillation, it preserves unique features like text grounding and segmentation correspondence.

Contributor guide
Assessment
This issue has not been assessed yet.