bytedance / bytedance/Dolphin

如何提高推理服务冷启动的速度

Open
#97 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
9k
Forks
775
PR merge metrics
No merged PRs in 30d

Description

按理参数才700M的大小,相比olmocr这个7B的参数量小很多,实际在使用的过程中,模型加载的速度并没有明显的优势,都需要30s左右的时间

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue names no files, tests, or entry points. Begin by locating the inference-service startup and model-loading path, then measure its cold-start stages against the reported roughly 30 seconds; the issue does not define a target completion time.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.