[Feature]: Qwen3.8-Flash-Next
Nobody has claimed this yet.
- Dominant language
- C++
- Stars
- 1.9k
- Forks
- 152
- Avg merge
- 4h 14m
- Merged PRs (30d)
- 11
Description
Suggestion Description
Devs are starting to run this model in CPU @ 6tps in Llama CPP and 12tps in CUDA in a RTX 3060 12Gb. Please, could you bring this model to XDNA2?
More open models are coming and they can even run at decent tps in CPU. By the paper XDNA2 would be able to handle these new models at a good pace.
Operating System
Ubuntu 26.04
GPU
Ryzen AI 9 HX 370
ROCm Component
No response
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The issue names no files, tests, or entry points. Start by reviewing how FastFlowLM currently supports models on XDNA2 and compare that with the Qwen3.8-Flash-Next model and its Llama CPP or CUDA usage; done means the model runs on the reported Ryzen AI hardware through XDNA2.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- ai
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 30/100