google / google/gemma.cpp

Add SmolLM2 support — reuse the Qwen3 converter path?

Offen
#982 5 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Vorherrschende Sprache
C++
Sterne
7k
Forks
660
Ø Merge
20 Std. 43 Min.
Gemergte PRs (30 T.)
33

Beschreibung

Hi @jan-wassenberg — I'd like to contribute SmolLM2 (135M / 360M / 1.7B) support, and wanted to check the approach with you before writing any code.

SmolLM2 is plain Llama-style: RMSNorm, SwiGLU, RoPE theta 10000, no biases, no QK-norm, tied embeddings, byte-level BPE, ChatML turn tokens. As far as I can tell that needs **no new kernels** — everything is already on the Qwen3 path.

Sketch:

1. `python/convert_from_safetensors.py` — the HF tensor names are identical to Qwen3's, so `export_qwen3_lm_sbs` almost works as-is. The only blockers are the `has_qk_norm` assert and deriving `head_dim` from `q_norm.weight`. Plus a `smollm2-*` dispatch prefix.
2. `gemma/configs.{h,cc}` — new `Model` enum values + config functions.
3. `gemma/tokenizer.cc` — reuse the Qwen3 branch (same `<|im_start|>` / `<|im_end|>`).
4. `gemma/gemma.cc` — `HasEmbeddingScaling()` has to return false for it.

Questions:

- Is a third family outside Gemma/Qwen welcome here, or would you rather keep the model list narrow?
- Prefer a family-neutral `export_llama_style_lm_sbs` that both Qwen3 and SmolLM2 route through, or a separate function?
- For `HasEmbeddingScaling`, would you rather grow the per-family check, or add a `ModelConfig` field?

Happy to send a PR if the direction sounds right.

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

Read python/convert_from_safetensors.py and the existing Qwen3 path first, then inspect gemma/configs.{h,cc}, gemma/tokenizer.cc, and gemma/gemma.cc for model registration, token handling, and embedding scaling. Done means the maintainers have chosen the family and converter design, and SmolLM2 135M, 360M, and 1.7B support is implemented without new kernels.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
cpp, python
Bereich
machine-learning
Issue-Typ
Feature
Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Aktivitätsstatus
Aktiv
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
38/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.