4paradigm / 4paradigm/k8s-vgpu-scheduler

如何在Prometheus里监控gpu的使用情况

Open
#40 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Go
Stars
595
Forks
100
PR merge metrics
No merged PRs in 30d

Description

包括每个pod的gpu使用率,显存。有相关指标的介绍吗?

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue asks about exposing GPU metrics (usage, memory) per pod in Prometheus for the k8s-vgpu-scheduler. First, examine the existing device plugin code to see what metrics are already collected. Look for Prometheus client integration or metrics endpoints. Research how other GPU device plugins (like NVIDIA's) expose metrics. Determine what new metrics need to be defined and how to label them by pod. Check if there are existing tests for metrics.

Written by the indexing model from the issue text.

Assessment

Tech stack
kubernetes, prometheus
Domain
cloud, devops, observability
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.