facebookresearch / facebookresearch/ImageBind
Get the multimodal embeddings
Open
- Dominant language
- Python
- Stars
- 9.1k
- Forks
- 842
- PR merge metrics
- No merged PRs in 30d
Description
Thank you for the great model.
I wonder how can I get the multimodat embedding of different inputs like image and its caption usign Imagebind?
if I can get that then how can it be compared to CLIP?
Contributor guide
Assessment
This issue has not been assessed yet.