NVIDIA / NVIDIA/Personal-AI-Router
[Feature]: Multiple ollama instances on single host support
Dieses Issue hat noch niemand übernommen.
- Vorherrschende Sprache
- Go
- Sterne
- 1.4k
- Forks
- 250
- Ø Merge
- 23 Std. 27 Min.
- Gemergte PRs (30 T.)
- 1
Beschreibung
Area
Routing and scheduling
User problem
I have a system with two GPUs from different vendors, one being a nvidia GPU. I've had to setup two ollama instances locally on different ports to get this to work with different settings. I'd like to be able to expose both ollama instances to the cluster
Desired outcome
Both ollama instances show up for a node (or it appears as two nodes)
Alternatives considered
No response
Compatibility and security implications
No response
Validation approach
Launch two ollama instances on different ports, ideally with specific GPUs enabled for each
Confirmations
- I searched existing issues for duplicates.
- I agree to follow the Code of Conduct.
Beitragsleitfaden
Erste Schritte
- Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
- Forke das Repository und arbeite in einem Branch.
- Öffne einen Pull Request, der die Issue-Nummer nennt.
Rechercherichtung
Beginne mit den Einstiegsstellen für Routing und Scheduling, die Ollama-Instanzen erkennen und verfügbar machen. Verfolge, wie der Ollama-Endpunkt und die Konfiguration eines Knotens dargestellt werden, und ermittle anschließend, ob zwei Instanzen als separate Knoten oder als Endpunkte auf einem Knoten verfügbar gemacht werden sollten. Überprüfe dies, indem du zwei Ollama-Instanzen auf unterschiedlichen Ports mit separaten GPU-Einstellungen startest und bestätigst, dass beide im Cluster erscheinen.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- go, ollama
- Bereich
- ai, backend
- Issue-Typ
- Feature
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Aktivitätsstatus
- Aktiv
- Klarheit
- Größtenteils klar
- Anfängerfreundlichkeit
- 52/100