alibaba / alibaba/clusterdata

Why does evaluator for an inference job consume so much time in the cluster-trace-gpu-v2020?

Open
#197 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
2.2k
Forks
482
PR merge metrics
No merged PRs in 30d

Description

1.as shown in the picture"evaluator" is for inference job ,and the "runtime" is giant:
![1695265375378](https://github.com/alibaba/clusterdata/assets/34826891/9f678eb7-cf60-4706-8fb4-d3fb9cb00b5b)
2.in the paper(MLaaS in the Wild: Workload Analysis and Scheduling
in Large-Scale Heterogeneous GPU Clusters),Figure 4a,the taskrun time is also begin 10s
image
inference job such as Image classification do not need 10s, so, there is no any such job in the cluster? and what is the job consume so much time ?

thank you very much!

Contributor guide

No contributing guide indexed for this repository

Research direction

Review the evaluator and runtime fields in cluster-trace-gpu-v2020, then compare their meanings with Figure 4a of the cited paper. Determine what the evaluator represents, why inference runtimes can be long, and whether the dataset contains the questioned jobs; document the explanation.

Written by the indexing model from the issue text.

Assessment

Domain
data
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.