4paradigm / 4paradigm/k8s-vgpu-scheduler

Vgpu的限制问题

Aperta
#28 1 commento 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
Go
Stelle
595
Fork
100
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

6月前更新的libvgpu.so。可以工作,在pytorch上工作正常,超出显存大小会正常报错。但是在tensorflow上不正常,显存限制不正常,可以超出切分的大小而不报错。

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Direzione di ricerca

The issue involves the libvgpu.so library and its interaction with TensorFlow. Start by examining the vGPU device plugin code, particularly the memory management and error handling for TensorFlow. Look for tests or examples of memory limit enforcement. The goal is to understand why TensorFlow does not respect the memory limit and to make it fail appropriately when exceeding the allocated vGPU memory.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
go, kubernetes, pytorch, tensorflow
Ambito
ai-infra-agents, cloud, devops, infrastructure
Tipo di issue
Bug
Difficoltà
4/5
Tempo stimato
3-5 giorni
Stato di attività
Ferma
Chiarezza
Abbastanza chiara
Idoneità per principianti
35/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.