elastic / elastic/integrations

[nvidia-gpu]: Not possible to get metrics for all nodes if running in k8s cluster

Open
#16,334 0 comments 0 reactions 0 assignees View on GitHub
Integration:nvidia_gpu needs:triage Team:Obs-InfraObs
Dominant language
Handlebars
Stars
333
Forks
647
Avg merge
2d 17h
Merged PRs (30d)
225

Description

### Integration Name

NVIDIA GPU Monitoring [nvidia_gpu]

### Dataset Name

_No response_

### Integration Version

0.4.1

### Agent Version

8.17.10

### Agent Output Type

elasticsearch

### Elasticsearch Version

8.17.1

### OS Version and Architecture

Ubuntu 22.04 LTS

### Software/API Version

_No response_

### Error Message

_No response_

### Event Original

_No response_

### What did you do?

Configure the nvidia-gpu integration to scrape metrics from our nvidia-gpu-operator service.
This results, because its a K8s service that acts as a loadbalancer of course, we only get metrics from the randomized GPU worker nodes.

### What did you see?

metrics from randomized GPUs (5 instead of 8 worker nodes currently)

### What did you expect to see?

Metrics from all GPU servers and not only from a randomized pattern

### Anything else?

_No response_

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.