pytorch / pytorch/executorch

no_rope_layer_interval is declared but never applied, and SmolLM3's own config sets it

Open
#22,043 2 comments 0 reactions 2 assignees View on GitHub

@MukeshK17 is already working on this.

Since Aug 24, 2026.

module: llm triaged
Dominant language
Python
Stars
5k
Forks
1.2k
Avg merge
2d 10h
Merged PRs (30d)
581

Description

ModelArgs declares no_rope_layer_interval, and examples/models/smollm3/3b_config.json sets it to 4, but nothing on the examples/models/llama export path reads it. On main (e4576d0) the name appears exactly once across the 32 Python files of that path:

examples/models/llama/model_args.py:119:    no_rope_layer_interval: Optional[int] = (

__post_init__ does not touch it, and Attention.forward applies rope unconditionally (attention.py:529, attention.py:563) with no layer-index test. The two places that do honour it are backends/mlx/llm/et_attention.py:111 and examples/qualcomm/oss_scripts/llama/model/static_llama.py:288.

So a model that alternates RoPE and NoPE layers exports without complaint and comes out with rope on every layer.

Reproducing

examples/models/smollm3/3b_config.json and examples/models/smollm3/convert_weights.py are both in the tree, so SmolLM3-3B looks ready to export. Converting its weights and exporting through export_llm with that params json (XNNPACK, 8da4w + 8-bit embedding, 1852.2 MB) gives a .pte that loads and runs, then does this:

prompt: <|im_start|>user\nWhat is the capital of France?<|im_end|>\n<|im_start|>assistant\n
output: Okay, so you want to know what what what what what what what what what what what
        what what what what what what what what what what what what what what what what

prompt: <|im_start|>user\nWhat is 17 times 4?<|im_end|>\n<|im_start|>assistant\n
output: I think you are looking for an answer that can be given that can be be be be be
        determ determ determ determ determ determ determ determ determ determ#ae determ#

It starts as English and collapses, which is what a positional-encoding mismatch looks like: the first few tokens are fine and the drift grows with position.

What would help

Either implement the interval on this path, or reject a params json that sets a field the chosen backend does not implement. The silent version is the expensive one — the export succeeds, the file is the right size, and the model is simply a different model.

Measured with the executorch 1.4.0 wheel; the code paths quoted above are byte-identical on main at e4576d0.

cc @larryliu0820 @mergennachin @cccclai @helunwencser @jackzhxng @digantdesai

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.