alibaba / alibaba/clusterdata

GPU Memory 'max_gpu_wrk_mem' seems to be more than the actual GPU type in GPU'20 trace ?

Open
#208 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
2.2k
Forks
482
PR merge metrics
No merged PRs in 30d

Description

For example in GPU 2020 trace

the job 'e5d6d5b546bff61f93b47ebf' has **max_gpu_wrk_mem** '44.289062' but the gpu type is V100 where the memory capacity should be 16GB or at max 32GB !?

Should I assume that once the **max_gpu_wrk_mem** > **GPU_type_capcity**, the worker encounters OOM ?

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the GPU 2020 trace entry for job e5d6d5b546bff61f93b47ebf and compare max_gpu_wrk_mem with the stated V100 capacity. Determine whether the value is a measurement, an aggregate, or an indication of OOM, and document the interpretation of values above the GPU capacity.

Written by the indexing model from the issue text.

Assessment

Domain
data
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.