Compile fails for all novel specs: health reports gpu_services:{}, finetune worker gets 500 from compiler on loopback

Ouverte
#5 1 commentaire 0 réactions 0 personnes assignées Voir sur GitHub

Personne n'a encore pris cette issue.

Évaluation

Difficulté
4/5
Temps estimé
3-5 jours
Accessibilité débutants
45/100
Type d'issue
Bug
Clarté
Plutôt claire
Activité
Active
Stack technique
python
Domaine
api, backend

Piste de recherche

Start by reproducing the failure with the /api/v1/health, /api/v1/compile, and /api/v1/compile/async endpoints, then trace the SDK's paw.compile() timeout and CompileResponse handling. Check the REST API documentation and the /api/openapi.json entry point. Done should distinguish unavailable compiler workers from slow compiles and clarify the cached result in SDK or troubleshooting output.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Description

Summary

All compiles of novel specs have been failing for ~62 hours. /api/v1/health reports "status": "healthy" but with "gpu_services": {} — no GPU workers registered. The finetune worker's error surfaces the underlying cause: a 500 from the standard compiler on loopback.

Metadata endpoints (/models/compilers, /compile/precheck) work normally, so this looks like the vLLM compiler services rather than the API host.

Reproduction

import httpx
# Any spec that is not already cached:
httpx.post("https://programasweights.com/api/v1/compile",
           json={"spec": "Return ONLY the word RED or BLUE. Say RED if the input mentions fire."},
           timeout=110)
# -> httpx.ReadTimeout

68 characters, no unusual content. The SDK's paw.compile() hits the same thing and fails at its hardcoded timeout=120.0.

The useful error

/api/v1/compile only ever hangs, but /compile/async (finetune) returns a job whose error field is readable:

POST /api/v1/compile/async  {"spec": "...", "compiler": "paw-ft-bs48"}
-> 202 {"job_id": "...", "status": "queued", "compiler_snapshot": "paw-ft-bs48-20260530"}

GET /api/v1/compile/<job_id>
-> [   0.1s] compiling pct=0.0
-> [ 111.4s] failed
   error: Server error '500 Internal Server Error' for url
          'http://127.0.0.1:18000/api/v1/compile'

Reproduced twice, ~24h apart (first run failed at 322s, second at 111s).

What is and isn't working

Endpoint Result
GET /api/v1/health 200 healthy, but gpu_services: {}, queue_depth: 0
GET /api/v1/models/compilers 200
POST /api/v1/compile/precheck 200, cached: false, correct compiler_snapshot
POST /api/v1/compile ReadTimeout
POST /api/v1/compile/async + poll failed, loopback 500 above

uptime_s was 154214 → 223474 across my checks (~62h), so the API host has not restarted in that window.

Ruled out

  • Not the spec. precheck returns 200 with cached: false on it. A 68-char spec on an unrelated topic fails identically.
  • Not spec size. Mine is 1408 chars / ~352 tokens against maxLength: 16000 and the documented 5120-token limit.
  • Not rate limiting. That returns 429; this is a timeout or a 500.
  • Not the client. Plain httpx bypassing the SDK behaves the same.

Two docs issues found along the way

  1. /api/v1/openapi.json 404s; the schema is at /api/openapi.json. Worth a mention in the REST API docs.
  2. Cache hits look like success. "Classify sentiment as positive or negative" returns 202 instantly because it is cached from the docs' own example. I initially took that as evidence my spec was at fault. A cached: true flag is already in CompileResponse — surfacing it in the SDK output, or noting it in the troubleshooting table, would save others the same wrong turn.

Suggestions

  • Have /api/v1/health report unhealthy (or a warnings entry) when gpu_services is empty, since compile cannot succeed in that state.
  • Return a 503 from /api/v1/compile instead of hanging when no compiler worker is available — the current behaviour is indistinguishable from a slow compile.
  • Consider making the SDK's 120s compile timeout configurable; there is currently no env override.

Environment

programasweights==0.4.4, Python 3.9, macOS. Anonymous (no PAW_API_KEY).

Langage dominant
Python
Étoiles
332
Forks
31
Merge moyen
2 j 3 h
PR mergées (30 j)
1

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Autres issues de programasweights/programasweights-python

Toutes les issues de programasweights/programasweights-python

Issues similaires

Plus d'issues Python

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.