huggingface / huggingface/optimum-executorch

Enable deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B

Open
#26 0 comments 0 reactions 0 assignees View on GitHub
bug
Dominant language
Python
Stars
141
Forks
48
Avg merge
26m
Merged PRs (30d)
2

Description

The bos_token_id doesn't match between the model config and its tokenizer. It happens on those using Qwen as the base model. Opened an discussion here: https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B/discussions/25

It may not fit on device w/o quantization, but exporting the llama based deepseek-R1 to ExecuTorch works just fine, e.g. setting model_id to `deepseek-ai/DeepSeek-R1-Distill-Llama-8B`

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.