NVIDIA / NVIDIA/TensorRT-Edge-LLM
Qwen3-VL-4B-Instruct TensorRT vs PyTorch layer-wise hidden states misalignment: layers 0–15 normal, layers 16+ cosine similarity drops to ~0.9 with inconsistent generation results
@fans-nv is already working on this.
Since Mar 3, 2026.
- Dominant language
- Python
- Stars
- 563
- Forks
- 135
- Avg merge
- 14h 13m
- Merged PRs (30d)
- 1
Description
Describe the bug
When comparing the layer-wise hidden states of the Qwen3-VL-4B-Instruct (VLM) model between the TensorRT (TRT) implementation and the official PyTorch implementation, a critical misalignment issue was identified:
- Layers 0–15 are well aligned, with cosine similarity close to 1, indicating consistent feature extraction.
- Starting from layer 16, the cosine similarity drops sharply to around 0.9, and the maximum absolute difference reaches over 6500, leading to severe misalignment of feature vectors.
In contrast, under the same testing setup (same input, environment, and evaluation method), the Qwen3-0.6B (pure LLM) model shows perfect alignment across all 26 layers between TensorRT and PyTorch, with cosine similarity close to 1 and negligible differences.
Impact
Blocker: The TensorRT implementation of the Qwen3-VL-4B-Instruct model produces outputs that are semantically similar but not identical to the official PyTorch version. This misalignment makes the TRT model unreliable for production deployment, as feature consistency is critical for VLM task accuracy.
Expected behavior
The TensorRT implementation of Qwen3-VL-4B-Instruct should fully align with the official PyTorch version in terms of hidden states and generated outputs:
- Cosine similarity of hidden states ≥ 0.99 across all 35 decoder layers.
- Max and mean differences of hidden states remain at a negligible level, consistent with the pure LLM (Qwen3-0.6B) performance.
- All layers pass the alignment check (marked as ✅ in the comparison table).
- Generated text from TensorRT should be identical to that from PyTorch.
Alignment Results - Qwen3-0.6B (Pure LLM) - Perfect Alignment
Layer-wise hidden states comparison between PyTorch and TensorRT shows full alignment across all layers:
================================================================================
Layer | Description | Cosine Similarity | Max Difference | Mean Difference | Status
-----+---------------------------+------------------+----------------+----------------+-------
0 | Embedding Output | 1.000005 | 0.000000 | 0.000000 | ✅
1 | Decoder Layer 0 Output | 0.999989 | 0.019531 | 0.000987 | ✅
2 | Decoder Layer 1 Output | 0.999988 | 0.046875 | 0.001506 | ✅
3 | Decoder Layer 2 Output | 1.000008 | 16.000000 | 0.003611 | ✅
4 | Decoder Layer 3 Output | 1.000009 | 16.000000 | 0.004102 | ✅
5 | Decoder Layer 4 Output | 1.000010 | 16.000000 | 0.004324 | ✅
6 | Decoder Layer 5 Output | 1.000011 | 16.000000 | 0.004904 | ✅
7 | Decoder Layer 6 Output | 1.000013 | 16.000000 | 0.005195 | ✅
8 | Decoder Layer 7 Output | 1.000012 | 16.000000 | 0.005722 | ✅
9 | Decoder Layer 8 Output | 1.000014 | 16.000000 | 0.006104 | ✅
10 | Decoder Layer 9 Output | 1.000016 | 16.000000 | 0.006457 | ✅
11 | Decoder Layer 10 Output | 1.000014 | 16.000000 | 0.007333 | ✅
12 | Decoder Layer 11 Output | 1.000014 | 16.000000 | 0.007791 | ✅
13 | Decoder Layer 12 Output | 1.000015 | 16.000000 | 0.008002 | ✅
14 | Decoder Layer 13 Output | 1.000014 | 16.000000 | 0.008343 | ✅
15 | Decoder Layer 14 Output | 1.000015 | 16.000000 | 0.008700 | ✅
16 | Decoder Layer 15 Output | 1.000015 | 16.000000 | 0.009891 | ✅
17 | Decoder Layer 16 Output | 1.000015 | 16.000000 | 0.011886 | ✅
18 | Decoder Layer 17 Output | 1.000013 | 16.000000 | 0.015036 | ✅
19 | Decoder Layer 18 Output | 1.000011 | 16.000000 | 0.018898 | ✅
20 | Decoder Layer 19 Output | 1.000008 | 16.000000 | 0.023356 | ✅
21 | Decoder Layer 20 Output | 1.000005 | 16.000000 | 0.029656 | ✅
22 | Decoder Layer 21 Output | 1.000004 | 16.000000 | 0.036687 | ✅
23 | Decoder Layer 22 Output | 1.000004 | 16.000000 | 0.043885 | ✅
24 | Decoder Layer 23 Output | 1.000004 | 16.000000 | 0.053605 | ✅
25 | Decoder Layer 24 Output | 1.000006 | 16.000000 | 0.065863 | ✅
26 | Decoder Layer 25 Output | 1.000003 | 16.000000 | 0.078584 | ✅
27 | Decoder Layer 26 Output | 0.999999 | 16.000000 | 0.088963 | ✅
28 | Layer 27 + Final Norm | 0.999933 | 1.171875 | 0.028502 | ✅
Generated Text Comparison (Qwen3-0.6B)
TensorRT: The capital of the United States is Washington, D.C.
PyTorch: The capital of the United States is Washington, D.C.
✅ Generated text is identical.
- Qwen3-VL-4B-Instruct (VLM) - Severe Misalignment
Layer-wise hidden states comparison between PyTorch and TensorRT (first 520 input tokens):
================================================================================
Layer | Description | Cosine Similarity | Max Difference | Mean Difference | Status
-----+---------------------------+------------------+----------------+----------------+-------
0 | Decoder Layer 0 | 0.995587 | 3.082031 | 0.044694 | ✅
1 | Decoder Layer 1 | 0.996698 | 7.390625 | 0.050343 | ✅
2 | Decoder Layer 2 | 0.997200 | 10.750000 | 0.056863 | ✅
3 | Decoder Layer 3 | 0.997302 | 12.742188 | 0.056604 | ✅
4 | Decoder Layer 4 | 0.997765 | 14.984375 | 0.055809 | ✅
5 | Decoder Layer 5 | 0.998015 | 14.406250 | 0.053952 | ✅
6 | Decoder Layer 6 | 1.001310 | 24.000000 | 0.052265 | ✅
7 | Decoder Layer 7 | 1.001245 | 24.000000 | 0.051395 | ✅
8 | Decoder Layer 8 | 1.001192 | 24.000000 | 0.051749 | ✅
9 | Decoder Layer 9 | 1.001075 | 24.000000 | 0.050262 | ✅
10 | Decoder Layer 10 | 1.000968 | 24.000000 | 0.051166 | ✅
11 | Decoder Layer 11 | 1.000891 | 24.000000 | 0.052188 | ✅
12 | Decoder Layer 12 | 1.000843 | 24.000000 | 0.053888 | ✅
13 | Decoder Layer 13 | 1.000813 | 24.000000 | 0.055465 | ✅
14 | Decoder Layer 14 | 1.000752 | 24.000000 | 0.056307 | ✅
15 | Decoder Layer 15 | 1.000753 | 24.000000 | 0.061572 | ✅
16 | Decoder Layer 16 | 0.903065 | 6570.000000 | 0.088749 | ❌
17 | Decoder Layer 17 | 0.903608 | 6561.000000 | 0.096210 | ❌
18 | Decoder Layer 18 | 0.903693 | 6562.000000 | 0.104558 | ❌
19 | Decoder Layer 19 | 0.903894 | 6563.000000 | 0.113052 | ❌
20 | Decoder Layer 20 | 0.903870 | 6565.000000 | 0.121492 | ❌
21 | Decoder Layer 21 | 0.904010 | 6566.000000 | 0.131104 | ❌
22 | Decoder Layer 22 | 0.904153 | 6568.000000 | 0.141548 | ❌
23 | Decoder Layer 23 | 0.904475 | 6569.000000 | 0.166450 | ❌
24 | Decoder Layer 24 | 0.904809 | 6571.000000 | 0.206467 | ❌
25 | Decoder Layer 25 | 0.905046 | 6571.000000 | 0.231294 | ❌
26 | Decoder Layer 26 | 0.905575 | 6571.000000 | 0.265599 | ❌
27 | Decoder Layer 27 | 0.905945 | 6571.000000 | 0.300187 | ❌
28 | Decoder Layer 28 | 0.906246 | 6570.000000 | 0.349266 | ❌
29 | Decoder Layer 29 | 0.906787 | 6570.000000 | 0.413752 | ❌
30 | Decoder Layer 30 | 0.907706 | 6568.000000 | 0.500248 | ❌
31 | Decoder Layer 31 | 0.908276 | 6565.000000 | 0.614202 | ❌
32 | Decoder Layer 32 | 0.908615 | 6565.000000 | 0.728069 | ❌
33 | Decoder Layer 33 | 0.909470 | 6568.000000 | 0.849629 | ❌
34 | Decoder Layer 34 | 0.978219 | 1632.000000 | 0.968672 | ⚠️
35 | Decoder Layer 35 | 0.935589 | 71.046875 | 0.533357 | ❌
Generated Text Comparison (Qwen3-VL-4B-Instruct)
The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set TRANSFORMERS_VERBOSITY=info for more details.
Input tokens: 520
TRT (610 chars):
This is a close-up, charming photograph of a red panda resting its head on a wooden structure.
Key details:
- Subject: A red panda, characterized by its reddish-brown fur, white facial markings around the eyes and muzzle, and fluffy white-tipped ears.
- Pose: The panda is looking directly a
PT (604 chars):
This is a close-up, charming photograph of a red panda resting its head on a wooden structure.
Key details:
- Subject: A red panda with a distinctive reddish-brown coat, white markings around its eyes, on its muzzle, and on the tips of its ears. Its dark, expressive eyes and small black nose ar
torch.Size([1976, 1536])
Torch pixel_values shape: torch.Size([1976, 1536])
❌ Mismatch at character 135:
TRT: 'bject**: A red panda, characterized by i'
PT: 'bject**: A red panda with a distinctive '
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.