mudler / mudler/LocalAI

Expose per-model memory usage (the per-process VRAM walk already computes it, then discards the split)

Open
#11,970 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement
Dominant language
Go
Stars
49.2k
Forks
4.5k
Avg merge
1d 3m
Merged PRs (30d)
239

Description

Is your feature request related to a problem? Please describe.
There is no way to ask LocalAI how much memory each loaded model holds, so a
downstream scheduler cannot split a shared GPU into "what the engine holds" and
"what everything else holds".

I run chat, embeddings, transcription and image generation through one LocalAI
container on one consumer card, and admit background work only when there is room
left for a user's request. Without a per-model figure that subtraction is
unavailable and background work has to be admitted unweighed.

Checked on v4.9.0 (f7ad3f70eb5d8a0ddf80e08557f0d7df28cf032e), image
localai/localai:v4.9.0-gpu-nvidia-cuda-12, RTX 3060, two models resident:

  • GET /system — ids and backend names only, no figures.
  • GET /api/pssize and size_vram are literal 0. Filed separately as a bug.
  • GET /backend/monitor?model=granite-4.1-8b — 500 here, and gopsutil RSS/VMS
    is host process memory rather than device memory even when it works:
    rpc error: code = Unimplemented desc = , then failed local node process sample: no grpc backend found for llama-cpp/models/granite-4.1-8b/granite-4.1-8b-Q4_K_M.gguf.bin
  • GET /api/resources — device aggregate and per-GPU totals, which is genuinely
    useful and I'll switch to it, but it cannot attribute usage to a model.
  • POST /api/models/vram-estimate — a prediction from weight files, not a reading.
    It can't see the actual offload split, and being wrong in the optimistic
    direction is what crashes the box.
  • GET /metrics — Go runtime metrics only.

Describe the solution you'd like
Per-model resident memory on /system, as optional per-entry fields:

{"loaded_models":[{"id":"granite-4.1-8b","backend":"llama-cpp","size_vram":5100000000}]}

Omitting the field (or null) where a backend can't report it is much better than 0,
for the reason in the linked bug. Weights alone would already be a large improvement
over nothing; a KV-cache split would be ideal but isn't required. /system vs. a new
endpoint vs. populating /api/ps — no preference, whichever fits the API surface.

Most of the machinery looks like it's already there on Linux. drmFdInfoUsageByRenderNode
walks /proc/[0-9]*/fd for DRM render fds and reads drm-total-<region> per PID:

https://github.com/mudler/LocalAI/blob/v4.9.0/pkg/xsysinfo/drmfdinfo.go#L48-L78

but then does out[renderName] += ..., aggregating per GPU and dropping the per-PID
breakdown. Since each backend is its own gRPC process, keeping the PID key and joining
it against the loaded-model list would give per-model VRAM with no new probing. That
path covers amdgpu/i915/xe on kernel >= 5.19 and not NVIDIA's proprietary driver, which
doesn't emit those keys — which is exactly why I'm asking for an optional per-entry
field rather than a guaranteed one. Please don't let uneven backend support block it:
partial coverage that says so is far more useful than none.

Describe alternatives you've considered

  • nvidia-smi --query-compute-apps — can't work across vendors, which is a large part
    of why LocalAI is attractive here. It also returns no rows at all in-container on a
    WSL2 host, while the card correctly reports 10849/12288 MiB used.
  • Summing /backend/monitor — rejected even with the 500 fixed: host RSS/VMS says
    nothing about VRAM and double-counts mapped weights.
  • Trusting /api/ps — would over-commit the card, per the linked bug.
  • Aggregate device reading plus a measured "everything else" peak — what I do today.
    It works but is coarse: the engine's own share is never known.
  • Estimating from the GGUF plus a modelled KV cache, i.e. what vram-estimate does —
    a guess wearing a measurement's clothes, and wrong in the direction that crashes.

Additional context
Searched existing issues for size_vram, api/ps, GPUMemoryInfo and vram-in-title and found nothing covering this. I'm not able to take the PR, but happy to test a change against this setup and to report what each backend does or doesn't report.

Written by my beloved Claude Code :)

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with pkg/xsysinfo/drmfdinfo.go, especially drmFdInfoUsageByRenderNode and its /proc DRM fd walk, then trace how /system assembles loaded model entries. Determine how the per-PID information can be associated with loaded backend processes without changing unsupported-driver behavior. Done means /system can expose an optional per-model memory field, omitting it when unavailable, with existing aggregate reporting preserved.

Written by the indexing model from the issue text.

Assessment

Tech stack
go, grpc, linux
Domain
api, backend, observability
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
68/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.