Compile fails for all novel specs: health reports gpu_services:{}, finetune worker gets 500 from compiler on loopback
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 45/100
Direzione di ricerca
Start by reproducing the failure with the /api/v1/health, /api/v1/compile, and /api/v1/compile/async endpoints, then trace the SDK's paw.compile() timeout and CompileResponse handling. Check the REST API documentation and the /api/openapi.json entry point. Done should distinguish unavailable compiler workers from slow compiles and clarify the cached result in SDK or troubleshooting output.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Summary
All compiles of novel specs have been failing for ~62 hours. /api/v1/health reports "status": "healthy" but with "gpu_services": {} — no GPU workers registered. The finetune worker's error surfaces the underlying cause: a 500 from the standard compiler on loopback.
Metadata endpoints (/models/compilers, /compile/precheck) work normally, so this looks like the vLLM compiler services rather than the API host.
Reproduction
import httpx
# Any spec that is not already cached:
httpx.post("https://programasweights.com/api/v1/compile",
json={"spec": "Return ONLY the word RED or BLUE. Say RED if the input mentions fire."},
timeout=110)
# -> httpx.ReadTimeout
68 characters, no unusual content. The SDK's paw.compile() hits the same thing and fails at its hardcoded timeout=120.0.
The useful error
/api/v1/compile only ever hangs, but /compile/async (finetune) returns a job whose error field is readable:
POST /api/v1/compile/async {"spec": "...", "compiler": "paw-ft-bs48"}
-> 202 {"job_id": "...", "status": "queued", "compiler_snapshot": "paw-ft-bs48-20260530"}
GET /api/v1/compile/<job_id>
-> [ 0.1s] compiling pct=0.0
-> [ 111.4s] failed
error: Server error '500 Internal Server Error' for url
'http://127.0.0.1:18000/api/v1/compile'
Reproduced twice, ~24h apart (first run failed at 322s, second at 111s).
What is and isn't working
| Endpoint | Result |
|---|---|
GET /api/v1/health |
200 healthy, but gpu_services: {}, queue_depth: 0 |
GET /api/v1/models/compilers |
200 |
POST /api/v1/compile/precheck |
200, cached: false, correct compiler_snapshot |
POST /api/v1/compile |
ReadTimeout |
POST /api/v1/compile/async + poll |
failed, loopback 500 above |
uptime_s was 154214 → 223474 across my checks (~62h), so the API host has not restarted in that window.
Ruled out
- Not the spec.
precheckreturns 200 withcached: falseon it. A 68-char spec on an unrelated topic fails identically. - Not spec size. Mine is 1408 chars / ~352 tokens against
maxLength: 16000and the documented 5120-token limit. - Not rate limiting. That returns 429; this is a timeout or a 500.
- Not the client. Plain
httpxbypassing the SDK behaves the same.
Two docs issues found along the way
/api/v1/openapi.json404s; the schema is at/api/openapi.json. Worth a mention in the REST API docs.- Cache hits look like success.
"Classify sentiment as positive or negative"returns 202 instantly because it is cached from the docs' own example. I initially took that as evidence my spec was at fault. Acached: trueflag is already inCompileResponse— surfacing it in the SDK output, or noting it in the troubleshooting table, would save others the same wrong turn.
Suggestions
- Have
/api/v1/healthreport unhealthy (or awarningsentry) whengpu_servicesis empty, since compile cannot succeed in that state. - Return a 503 from
/api/v1/compileinstead of hanging when no compiler worker is available — the current behaviour is indistinguishable from a slow compile. - Consider making the SDK's 120s compile timeout configurable; there is currently no env override.
Environment
programasweights==0.4.4, Python 3.9, macOS. Anonymous (no PAW_API_KEY).
- Lingua principale
- Python
- Stelle
- 332
- Fork
- 31
- Merge medio
- 2g 3h
- PR unite (30g)
- 1
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di programasweights/programasweights-python
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
Tutte le issue di programasweights/programasweights-python
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
-
🐛 Bug 🔔 Pending processing
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
jumpserver/jumpserver#17584 ·