huggingface / huggingface/pytorch-image-models

[FEATURE] Add ViT weights: RADIO

Open
#2,177 4 comments 7 reactions 0 assignees View on GitHub
enhancement
Dominant language
Python
Stars
37.1k
Forks
5.2k
Avg merge
1d 13h
Merged PRs (30d)
34

Description

https://github.com/NVlabs/RADIO

The code and model weights of paper *[CVPR 2024] AM-RADIO: Agglomerative Vision Foundation Model - Reduce All Domains Into One* has been released by Nvidia

> RADIO , a new vision foundation model (actually a new vit pretrained weight), excels across visual domains, serving as a superior replacement for vision backbones. Integrating CLIP variants, DINOv2, and SAM through distillation, it preserves unique features like text grounding and segmentation correspondence.

![image](https://github.com/huggingface/pytorch-image-models/assets/19152032/b084f0c3-0930-4163-9022-20d1f4ab82b9)

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.