Compile fails for all novel specs: health reports gpu_services:{}, finetune worker gets 500 from compiler on loopback

Offen
#5 1 Kommentar 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

Bewertung

Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Anfängerfreundlichkeit
45/100
Issue-Typ
Bug
Klarheit
Größtenteils klar
Aktivitätsstatus
Aktiv
Tech-Stack
python
Bereich
api, backend

Rechercherichtung

Start by reproducing the failure with the /api/v1/health, /api/v1/compile, and /api/v1/compile/async endpoints, then trace the SDK's paw.compile() timeout and CompileResponse handling. Check the REST API documentation and the /api/openapi.json entry point. Done should distinguish unavailable compiler workers from slow compiles and clarify the cached result in SDK or troubleshooting output.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Beschreibung

Summary

All compiles of novel specs have been failing for ~62 hours. /api/v1/health reports "status": "healthy" but with "gpu_services": {} — no GPU workers registered. The finetune worker's error surfaces the underlying cause: a 500 from the standard compiler on loopback.

Metadata endpoints (/models/compilers, /compile/precheck) work normally, so this looks like the vLLM compiler services rather than the API host.

Reproduction

import httpx
# Any spec that is not already cached:
httpx.post("https://programasweights.com/api/v1/compile",
           json={"spec": "Return ONLY the word RED or BLUE. Say RED if the input mentions fire."},
           timeout=110)
# -> httpx.ReadTimeout

68 characters, no unusual content. The SDK's paw.compile() hits the same thing and fails at its hardcoded timeout=120.0.

The useful error

/api/v1/compile only ever hangs, but /compile/async (finetune) returns a job whose error field is readable:

POST /api/v1/compile/async  {"spec": "...", "compiler": "paw-ft-bs48"}
-> 202 {"job_id": "...", "status": "queued", "compiler_snapshot": "paw-ft-bs48-20260530"}

GET /api/v1/compile/<job_id>
-> [   0.1s] compiling pct=0.0
-> [ 111.4s] failed
   error: Server error '500 Internal Server Error' for url
          'http://127.0.0.1:18000/api/v1/compile'

Reproduced twice, ~24h apart (first run failed at 322s, second at 111s).

What is and isn't working

Endpoint Result
GET /api/v1/health 200 healthy, but gpu_services: {}, queue_depth: 0
GET /api/v1/models/compilers 200
POST /api/v1/compile/precheck 200, cached: false, correct compiler_snapshot
POST /api/v1/compile ReadTimeout
POST /api/v1/compile/async + poll failed, loopback 500 above

uptime_s was 154214 → 223474 across my checks (~62h), so the API host has not restarted in that window.

Ruled out

  • Not the spec. precheck returns 200 with cached: false on it. A 68-char spec on an unrelated topic fails identically.
  • Not spec size. Mine is 1408 chars / ~352 tokens against maxLength: 16000 and the documented 5120-token limit.
  • Not rate limiting. That returns 429; this is a timeout or a 500.
  • Not the client. Plain httpx bypassing the SDK behaves the same.

Two docs issues found along the way

  1. /api/v1/openapi.json 404s; the schema is at /api/openapi.json. Worth a mention in the REST API docs.
  2. Cache hits look like success. "Classify sentiment as positive or negative" returns 202 instantly because it is cached from the docs' own example. I initially took that as evidence my spec was at fault. A cached: true flag is already in CompileResponse — surfacing it in the SDK output, or noting it in the troubleshooting table, would save others the same wrong turn.

Suggestions

  • Have /api/v1/health report unhealthy (or a warnings entry) when gpu_services is empty, since compile cannot succeed in that state.
  • Return a 503 from /api/v1/compile instead of hanging when no compiler worker is available — the current behaviour is indistinguishable from a slow compile.
  • Consider making the SDK's 120s compile timeout configurable; there is currently no env override.

Environment

programasweights==0.4.4, Python 3.9, macOS. Anonymous (no PAW_API_KEY).

Vorherrschende Sprache
Python
Sterne
332
Forks
31
Ø Merge
2 T. 3 Std.
Gemergte PRs (30 T.)
1

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Erste Schritte

  1. Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
  3. Forken Sie das Repository und arbeiten Sie in einem Branch.
  4. Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.

Mehr aus programasweights/programasweights-python

Alle Issues in programasweights/programasweights-python

Ähnliche Issues

Weitere Issues zu Python

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.