b200 fp4 minimax vllm single node / B200 FP4 MiniMax vLLM 单节点

Open
#951 7 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
35/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Quiet
Tech stack
huggingface, python

Research direction

Start by locating the existing vLLM benchmark entry point and configurations for B200 or MI355 single-node runs. Add coverage for nvidia/MiniMax-M2.5-NVFP4 with TP=2 and TP=4, covering concurrency 4 through 64, then verify the benchmark completes on the intended hardware; resolve whether B200 or MI355 is the target.

Written by the indexing model from the issue text.

Description

@claude vllm add minimax fp4 for mi355 tp4 & tp2 from conc 4 to 64

huggingface model nvidia/MiniMax-M2.5-NVFP4

中文说明

请求在 B200 上使用 vLLM 添加 MiniMax FP4 单节点基准测试,支持 TP=4 和 TP=2,并发范围从 4 到 64。HuggingFace 模型为 nvidia/MiniMax-M2.5-NVFP4

Dominant language
Python
Stars
1.7k
Forks
303
Avg merge
1d 13h
Merged PRs (30d)
284

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from SemiAnalysisAI/InferenceX

All issues in SemiAnalysisAI/InferenceX

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.