[Bug]: Qwen3MoE + Engines - repeating text issue
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 14.7k
- Forks
- 2.8k
- Avg merge
- 2d 23h
- Merged PRs (30d)
- 489
Description
System Info
H100, TRT-LLM 1.2.0rc8.
Who can help?
Qwen3MoE is supported of engines style. A migration to torchflow is not supported from our end as we are working on a
While the first ~10 tokens are okay, the quality seems to suffer over time.
{"stream": true, "messages": [{"role": "user", "content": "Tell me everything you know about optimized inference."}], "max_tokens": 512, "temperature": 0.5}
yields
"<think>\nOkay, the user is asking for everything I know about optimized inference. I need to make sure I cover all aspects. Let me start by recalling what optimized inference means. It's about improving the efficiency of models during inference, right? So, I should explain the concept, maybe the different techniques, like model compression, quantization, pruning, etc. Also, maybe the hardware aspects, like using specialized hardware for inference. Then, the applications, like in real-time systems, or in mobile devices. Also, the benefits, like reduced latency, lower power consumption, etc. Maybe the challenges, like maintaining accuracy while optimizing. Also, the tools, like frameworks such as TensorFlow Lite, PyTorch, etc. Maybe the research, like the papers or studies. Also, the examples, like specific models or cases. I need to structure this in a coherent way. Let me check if I have all the points. Maybe I should start with the definition, then the techniques, then the hardware, then applications, then benefits, then challenges, then tools, then research, then examples. Also, maybe the different types of optimization, like model optimization, data optimization, etc. Also, the methods, like the different approaches. Maybe the the process, like the steps involved. Also, the the importance, like why it's important. I need to make sure I don't miss any key points. Also, maybe the the current trends, like the latest developments. Also, the the the future, like the potential. I need to make sure I have a comprehensive answer. Let me think about the possible topics. Maybe the the the different types of optimization, like the model optimization, the data optimization, the algorithm optimization. Also, the the the different methods, like the model pruning, the quantization, the model compression. Also, the the the different hardware, like the specialized hardware, like the GPU, the TPU, etc. Also, the the the different tools, like the frameworks, like the TensorFlow Lite, the PyTorch, etc. Also, the the the different applications, like the real-time, like the mobile, etc. Also, the the the different benefits, like the reduced latency, the lower power, etc. Also, the the the different challenges, like the maintaining accuracy, the the trade-off, etc. Also, the the the different research, like the papers, like the studies, etc. Also, the the the different examples, like the specific models, like the example, etc. Also, the the the different types of optimization, like the model optimization, the data optimization, the algorithm optimization. Also, the the the different methods, like the model pruning, the quantization, the model compression. Also, the the the different hardware, like the specialized hardware, like the GPU, the TPU, etc. Also, the the the different tools, like the frameworks, like the TensorFlow Lite, the PyTorch, etc. Also, the the different applications, like the real-time, like the mobile, etc. Also, the the different benefits, like the reduced latency, the lower power, etc. Also, the the different challenges, like the maintaining accuracy, the the trade-off, etc. Also, the the the different research, like the papers, like the studies, etc. Also, the the the different examples, like the specific models, like the example, etc. I need to make sure I cover all these. Maybe I should start with the definition, then the techniques, then the hardware, then applications, then benefits, then challenges, then tools, then research, then examples. Also, the the different types of optimization, like the model optimization, the data optimization, the algorithm optimization. Also, the the different methods, like the model pruning, the quantization, the model compression. Also, the the different hardware, like the specialized hardware, like the GPU, the TPU, etc. Also, the the different tools, like the frameworks, like the TensorFlow Lite, the PyTorch, etc. Also, the different applications, like the real-time, like the mobile, etc. Also, the different benefits, like the reduced latency, the lower power, etc. Also, the different challenges, like the maintaining accuracy, the the trade-off, etc. Also, the the different research, like the papers, like the studies, etc. Also, the the different examples, like the specific models, like the example, etc. I need to make sure I have a comprehensive answer. Let me check if I have all the points. Maybe I should start with the definition, then the techniques, then the hardware, then applications, then benefits, then challenges, then tools, then research, then examples. Also, the different types of optimization, like the model optimization, the data optimization, the algorithm optimization. Also, the different methods, like the model pruning, the quantization, the model compression. Also, the different hardware, like the specialized hardware, like the GPU, the TPU, etc. Also, the different tools, like the frameworks, like the TensorFlow Lite, the PyTorch, etc. Also, the different applications, like the real-time, like the mobile, etc. Also, the different benefits, like the reduced latency, the lower power, etc. Also, the different challenges, like the maintaining accuracy, the the trade-off, etc. Also, the the different research, like the papers, like the studies, etc. Also, the the different examples, like the specific models, like the example, etc. I need to make sure I cover all these. Maybe I should start with the definition, then the techniques, then the hardware, then applications, then benefits, then challenges, then tools, then research, then examples. Also, the different types of optimization, like the model optimization, the data optimization, the algorithm optimization. Also, the different methods, like the model pruning, the quantization, the model compression. Also, the different hardware, like the specialized hardware, like the GPU, the TPU, etc. Also, the different tools, like the frameworks, like the TensorFlow Lite, the PyTorch, etc. Also, the different applications, like the real-time, like the mobile, etc. Also, the different benefits, like the reduced latency, the lower power, etc. Also, the different challenges, like the maintaining accuracy, the the trade-off, etc. Also, the the different research, like the papers, like the studies, etc. Also, the the different examples, like the specific models, like the example, etc. I need to make sure I have a comprehensive answer. Let me check if I have all the points. Maybe I should start with the definition, then the techniques, then the hardware, then applications, then benefits, then challenges, then tools, then research, then examples. Also, the different types of optimization, like the model optimization, the data optimization, the algorithm optimization. Also, the different methods, like the model pruning, the quantization, the model compression. Also, the different hardware, like the specialized hardware, like the GPU, the TPU, etc. Also, the different tools, like the frameworks, like the TensorFlow Lite, the PyTorch, etc. Also, the different applications, like the real-time, like the mobile, etc. Also, the different benefits, like the reduced latency, the lower power, etc. Also, the different challenges, like the maintaining accuracy, the the trade-off, etc. Also, the the different research, like the papers, like the studies, etc. Also, the the different examples, like the specific models, like the example, etc. I need to make sure I cover all these. Maybe I should start with the definition, then the techniques, then the hardware, then applications, then benefits, then challenges, then tools, then research, then examples. Also, the different types of optimization, like the model optimization, the data optimization, the algorithm optimization. Also, the different methods, like the model pruning, the quantization, the model compression. Also, the different hardware, like the specialized hardware, like the GPU, the TPU, etc. Also, the different tools, like the frameworks, like the TensorFlow Lite, the PyTorch, etc. Also, the different applications, like the real-time, like the mobile, etc. Also, the different benefits, like the reduced latency, the lower power, etc. Also, the different challenges, like the maintaining accuracy, the the trade-off, etc. Also, the the different research, like the papers, like the studies, etc. Also, the the different examples, like the specific models, like the example, etc. I need to make sure I have a comprehensive answer. Let me check if I have all the points. Maybe I should start with the definition, then the techniques, then the hardware, then applications, then benefits, then challenges, then tools, then research, then examples. Also, the different types of optimization, like the model optimization, the data optimization, the algorithm optimization. Also, the different methods, like the model pruning, the quantization, the model compression. Also, the different hardware, like the specialized hardware, like the GPU, the TPU, etc. (And so on, never stops)
Repo:
- what leads to the problem here? It seems like there is a precision / modeling issue.
- This happens for TP2 bf16 engines as well as TP2 fp16 engines.
- It also happens for modelopt fp8 engines
- i also tried with MoEConfig.RENORMILIZATION
- I also tried strongly_typed bf16
It feels like the MoE Router or some layer suffers from a precision issue.
Any help for supporting Qwen3MoE would be apprechiated.
Information
- The official example scripts
- My own modified scripts
Tasks
- An officially supported task in the
examplesfolder (such as GLUE/SQuAD, ...) - My own task or dataset (give details below)
Reproduction
Please ping us in Nvidia Baseten Slack if you have any ideas.
Expected behavior
actual behavior
additional notes
Before submitting a new issue...
- Make sure you already searched for relevant issues, and checked the documentation and examples for answers to frequently asked questions.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reproducing the repeating-text behavior with the uploaded Hugging Face engine commit and the TP2 bf16, fp16, modelopt fp8, renormalization, and strongly typed bf16 configurations described in the report. Compare the outputs and narrow the failure to the MoE router or another layer; done means identifying the precision or modeling cause and documenting a verified correction.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- machine-learning, performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100