NVIDIA / NVIDIA/TensorRT-LLM

Problem serving nvidia/DeepSeek-V3-0324-FP4 on 8xH200

Open
#6,038 4 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
14.7k
Forks
2.8k
Avg merge
2d 23h
Merged PRs (30d)
489

Description

Hi,

I'm trying to deploy huggingface nvidia/DeepSeek-V3-0324-FP4 on 8xH200 server. Both trtllm-serve, vLLM, and sglang are not able to serve this model because of the GPU arch not support this model.

I noticed that the supported device in README of huggingface is:

Inference:
Engine: TensorRT-LLM
Test Hardware: B200

So, my questions is:

  1. how could I run this model on NVIDIA H200?
  2. If not, is there any document that I can convert model of deepseek v3 0324 myself to run FP4 on H200?

Thank you!

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the Hugging Face model README and the trtllm-serve, vLLM, and sglang entry points mentioned in the report, checking their GPU architecture and FP4 support for H200. Done would require a confirmed way to serve the model on 8xH200 or documentation of a working conversion path, including any limitations.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai, backend
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.