NVIDIA / NVIDIA/TensorRT-Edge-LLM

enable_thinking=True produces reasoning text without literal <think>...</think> tags — _ThinkingStateMachine cannot separate it from the final answer

Open
#113 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
563
Forks
135
Avg merge
14h 13m
Merged PRs (30d)
1

Description

Environment

  • Hardware: Jetson AGX Thor, 128GB unified memory, JetPack 7.2, CUDA 13.2
  • TensorRT-Edge-LLM version: (check with cat ~/TensorRT-Edge-LLM/VERSION or git log)
  • Model: Qwen3-14B, quantized to NVFP4 via tensorrt_edgellm.scripts.quantize
  • Server: experimental OpenAI-compatible server (experimental/server),
    custom persistent wrapper around LLM.generate_stream()
  • Engine built with: llm_build --maxInputLen 40960 --maxKVCacheCapacity 40960

Description

When generating with SamplingParams(enable_thinking=True, ...), the model
produces reasoning content as plain text but does not emit literal
<think> / </think> tags
in the streamed or final output.

api_server.py defines _ThinkingStateMachine (line ~219) specifically to
parse THINK_OPEN_TAG = "<think>" / THINK_CLOSE_TAG = "</think>" boundaries
in streamed text — but since the engine never emits these tags, the state
machine has nothing to key on and reasoning text is indistinguishable from
the final answer in raw output.

Steps to reproduce

  1. Quantize Qwen3-14B to NVFP4, export to ONNX, build TRT engine
    (standard pipeline per Quick Start Guide)
  2. Serve with a persistent Python wrapper calling
    llm.generate_stream(messages, SamplingParams(enable_thinking=True, ...))
  3. Send a simple prompt, e.g. "what is 2+2"
  4. Inspect raw streamed/concatenated token text

Expected

Raw output should contain <think>\n...reasoning...\n</think>\n\n2 + 2 = 4
so _ThinkingStateMachine (or any consumer) can reliably split reasoning
from the final answer.

Actual

Raw output is plain text with no <think> tags at all, e.g.:

Because reasoning itself frequently contains multiple \n\n-separated
paragraphs, there is no reliable text-based heuristic to find the boundary
between reasoning and the final answer — _ThinkingStateMachine is
unusable in this configuration.

Things I checked

  • Swapped generation_prompt and generation_prompt_thinking in
    processed_chat_template.json — output behavior with respect to tag
    presence did not change (only whether reasoning happens at all changes
    with enable_thinking=True/False, not whether tags are emitted).
  • Confirmed enable_thinking=False reliably produces clean, tag-free,
    reasoning-free output — this works as expected, just disables the
    feature entirely rather than exposing it cleanly.

Question

Is <think> tag emission expected to work with the persistent
LLM.generate_stream() Python API + a manually-quantized NVFP4 engine,
or is tag emission only supported through a specific server entrypoint /
chat template configuration we may be missing? If there's a known-correct
recipe for enable_thinking=True + reliable <think> tag output on a
custom-quantized Qwen3 engine, a pointer would be very helpful.

Happy to provide the full processed_chat_template.json, exact
quantize/export/build commands, or raw debug logs on request.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with experimental/server/api_server.py, especially _ThinkingStateMachine, and trace how LLM.generate_stream() handles enable_thinking=True. Compare that behavior with the generation_prompt and generation_prompt_thinking entries in processed_chat_template.json. Done means determining whether tag emission is supported for this API and custom engine configuration, and documenting the correct entrypoint or configuration if one exists.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai, backend-api-design
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.