4paradigm / 4paradigm/k8s-vgpu-scheduler

Is there a way to monitor vGPU with DCGM?

Open
#18 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Go
Stars
595
Forks
100
PR merge metrics
No merged PRs in 30d

Description

DCGM exporter is not picking the pods that are using vGPU, making it hard to to track utilization of the pods.
is there any workaround to monitor GPU utilization with vGPU?
is there a way to get the mapping between the vGPU and the actual GPU IDs?

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue asks about monitoring vGPU with DCGM and mapping vGPU to physical GPU IDs. Start by examining the DCGM exporter code and the vGPU device plugin's internal mapping logic. Look for existing metrics exposure or pod labeling. 'Done' would be a documented method or code change enabling DCGM to track vGPU-utilizing pods.

Written by the indexing model from the issue text.

Assessment

Tech stack
docker, kubernetes
Domain
cloud, devops, observability
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.