NVIDIA / NVIDIA/TensorRT-LLM

[Usage]: When using TensorRT-LLM with gpt-oss, our system prompt is getting modified!!!!!!

Open
#9,188 4 comments 0 reactions 1 assignee View on GitHub

@LinPoly is already working on this.

Since Nov 17, 2025.

question
Dominant language
Python
Stars
14.7k
Forks
2.8k
Avg merge
2d 23h
Merged PRs (30d)
489

Description

System Info

When using gpt-oss-120b, our system prompt is getting modified.

This is what is getting appended:

Knowledge cutoff:** 2024-06-30 \nCurrent date: 2025-11-15\n\nReasoning: high\n\nChannels: analysis, commentary, final. The channel must be included for every message.

this is not the correct implementation of the harmony format, and it is causing a lot of errors in the output.

in particular, we keep seeing this:

[TRT-LLM] [W] ⚠️ Harmony parsing fell back to raw text decoding

along with a lot of garbled text, i.e.

 ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ..... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ..... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... . ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... . ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ..... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ..... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... . ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... . ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ..... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ..... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... . ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... . ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ...

and this:

[11/15/2025-17:11:47] [TRT-LLM] [W] Failed to parse harmony messages from tokens: %s Unknown role: assistantfinal
[11/15/2025-17:11:47] [TRT-LLM] [W] Failed to parse harmony output: %s. Raw output: %s Harmony parsing failed: Unknown role: assistantfinal analysis<|message|>We need to interpret the user message:

we are just passing in a messages array: [{"role": "system", "content": ...,}, {"role": "user", "content": ...}]

These server endpoints should be stateless, you should not be modifying (or even touching) our prompts.

How do we fix this efficiently?

This is our launch script:

trtllm-serve /huggingface_hub/models/openai/gpt_oss_120b/ \
  --host 0.0.0.0 \
  --port 8355 \
  --max_batch_size 4 \
  --extra_llm_api_options /cache/extra_llm_api_config.yml \
  --max_num_tokens 131072 \
  --max_seq_len 131072

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.