addyosmani / addyosmani/chatty

[feature] WebAI - Add support for SmolVLM 256M (& 500M)

Ouverte
#82 1 commentaire 0 réactions 0 personnes assignées Voir sur GitHub
enhancement good first issue help wanted
Langage dominant
TypeScript
Étoiles
837
Forks
103
Métriques de merge des PR
Aucune PR mergée en 30 j

Description

> SmolVLM 256M (& 500M): The world’s smallest multimodal model. Designed for efficiency and perfect for on-device applications

Demo: https://huggingface.co/spaces/HuggingFaceTB/SmolVLM-256M-Instruct-WebGPU
Twitter: https://x.com/xenovacom/status/1882435994160447587
Examples: https://github.com/huggingface/transformers.js-examples/tree/main/smolvlm-webgpu

It would be great to extend our models with SmolVLM.

Image

There are some UX improvements we could also make at the same time such as, when you use a prompt with an image:

Fixing the [Object] text following a selected image.

Image

Double-checking there are no specific issues with multi-modal support. I was running into issues with our other multimodal model (couldn't get it working), unsure if this was model specific.

Image

Image

Guide de contribution

Ouvrir le guide de contribution

Évaluation

Cette issue n'a pas encore été évaluée.

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.