PaddlePaddle / PaddlePaddle/FastDeploy

如何提高推理响应时间?

Open
#3,269 0 comments 0 reactions 1 assignee View on GitHub

@gzy19990617 is already working on this.

Since Aug 8, 2025.

Dominant language
Python
Stars
3.7k
Forks
756
Avg merge
19h 28m
Merged PRs (30d)
4

Description

环境:
8张昆仑芯P800(96GB显存)
执行参数如下:
python -m fastdeploy.entrypoints.openai.api_server
--model /Work/deepseek32b
--port 8188
--metrics-port 8181
--engine-worker-queue-port 8182
--tensor-parallel-size 8
--max-model-len 16384
--max-num-seqs 64
--max-num-batched-tokens 16384
--kv-cache-ratio 0.8
--enable-chunked-prefill
--gpu-memory-utilization 0.85
--graph-optimization-config '{"use_cudagraph":true,"graph_opt_level":1}'
--reasoning-parser qwen3

性能详见附件,平均端到端响应速度需要143秒,这速度太慢了。
调整什么参数能提高响应速度。

性能分析报告.xlsx

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.