programasweights / programasweights/programasweights-python

Compile fails for all novel specs: health reports gpu_services:{}, finetune worker gets 500 from compiler on loopback

Open
#5 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
332
Forks
31
Avg merge
2d 3h
Merged PRs (30d)
1

Description

Summary

All compiles of novel specs have been failing for ~62 hours. /api/v1/health reports "status": "healthy" but with "gpu_services": {} — no GPU workers registered. The finetune worker's error surfaces the underlying cause: a 500 from the standard compiler on loopback.

Metadata endpoints (/models/compilers, /compile/precheck) work normally, so this looks like the vLLM compiler services rather than the API host.

Reproduction

import httpx
# Any spec that is not already cached:
httpx.post("https://programasweights.com/api/v1/compile",
           json={"spec": "Return ONLY the word RED or BLUE. Say RED if the input mentions fire."},
           timeout=110)
# -> httpx.ReadTimeout

68 characters, no unusual content. The SDK's paw.compile() hits the same thing and fails at its hardcoded timeout=120.0.

The useful error

/api/v1/compile only ever hangs, but /compile/async (finetune) returns a job whose error field is readable:

POST /api/v1/compile/async  {"spec": "...", "compiler": "paw-ft-bs48"}
-> 202 {"job_id": "...", "status": "queued", "compiler_snapshot": "paw-ft-bs48-20260530"}

GET /api/v1/compile/<job_id>
-> [   0.1s] compiling pct=0.0
-> [ 111.4s] failed
   error: Server error '500 Internal Server Error' for url
          'http://127.0.0.1:18000/api/v1/compile'

Reproduced twice, ~24h apart (first run failed at 322s, second at 111s).

What is and isn't working

Endpoint Result
GET /api/v1/health 200 healthy, but gpu_services: {}, queue_depth: 0
GET /api/v1/models/compilers 200
POST /api/v1/compile/precheck 200, cached: false, correct compiler_snapshot
POST /api/v1/compile ReadTimeout
POST /api/v1/compile/async + poll failed, loopback 500 above

uptime_s was 154214 → 223474 across my checks (~62h), so the API host has not restarted in that window.

Ruled out

  • Not the spec. precheck returns 200 with cached: false on it. A 68-char spec on an unrelated topic fails identically.
  • Not spec size. Mine is 1408 chars / ~352 tokens against maxLength: 16000 and the documented 5120-token limit.
  • Not rate limiting. That returns 429; this is a timeout or a 500.
  • Not the client. Plain httpx bypassing the SDK behaves the same.

Two docs issues found along the way

  1. /api/v1/openapi.json 404s; the schema is at /api/openapi.json. Worth a mention in the REST API docs.
  2. Cache hits look like success. "Classify sentiment as positive or negative" returns 202 instantly because it is cached from the docs' own example. I initially took that as evidence my spec was at fault. A cached: true flag is already in CompileResponse — surfacing it in the SDK output, or noting it in the troubleshooting table, would save others the same wrong turn.

Suggestions

  • Have /api/v1/health report unhealthy (or a warnings entry) when gpu_services is empty, since compile cannot succeed in that state.
  • Return a 503 from /api/v1/compile instead of hanging when no compiler worker is available — the current behaviour is indistinguishable from a slow compile.
  • Consider making the SDK's 120s compile timeout configurable; there is currently no env override.

Environment

programasweights==0.4.4, Python 3.9, macOS. Anonymous (no PAW_API_KEY).

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the failure with the /api/v1/health, /api/v1/compile, and /api/v1/compile/async endpoints, then trace the SDK's paw.compile() timeout and CompileResponse handling. Check the REST API documentation and the /api/openapi.json entry point. Done should distinguish unavailable compiler workers from slow compiles and clarify the cached result in SDK or troubleshooting output.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
api, backend
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.