programasweights / programasweights/programasweights-python
Compile fails for all novel specs: health reports gpu_services:{}, finetune worker gets 500 from compiler on loopback
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 332
- Forks
- 31
- Avg merge
- 2d 3h
- Merged PRs (30d)
- 1
Description
Summary
All compiles of novel specs have been failing for ~62 hours. /api/v1/health reports "status": "healthy" but with "gpu_services": {} — no GPU workers registered. The finetune worker's error surfaces the underlying cause: a 500 from the standard compiler on loopback.
Metadata endpoints (/models/compilers, /compile/precheck) work normally, so this looks like the vLLM compiler services rather than the API host.
Reproduction
import httpx
# Any spec that is not already cached:
httpx.post("https://programasweights.com/api/v1/compile",
json={"spec": "Return ONLY the word RED or BLUE. Say RED if the input mentions fire."},
timeout=110)
# -> httpx.ReadTimeout
68 characters, no unusual content. The SDK's paw.compile() hits the same thing and fails at its hardcoded timeout=120.0.
The useful error
/api/v1/compile only ever hangs, but /compile/async (finetune) returns a job whose error field is readable:
POST /api/v1/compile/async {"spec": "...", "compiler": "paw-ft-bs48"}
-> 202 {"job_id": "...", "status": "queued", "compiler_snapshot": "paw-ft-bs48-20260530"}
GET /api/v1/compile/<job_id>
-> [ 0.1s] compiling pct=0.0
-> [ 111.4s] failed
error: Server error '500 Internal Server Error' for url
'http://127.0.0.1:18000/api/v1/compile'
Reproduced twice, ~24h apart (first run failed at 322s, second at 111s).
What is and isn't working
| Endpoint | Result |
|---|---|
GET /api/v1/health |
200 healthy, but gpu_services: {}, queue_depth: 0 |
GET /api/v1/models/compilers |
200 |
POST /api/v1/compile/precheck |
200, cached: false, correct compiler_snapshot |
POST /api/v1/compile |
ReadTimeout |
POST /api/v1/compile/async + poll |
failed, loopback 500 above |
uptime_s was 154214 → 223474 across my checks (~62h), so the API host has not restarted in that window.
Ruled out
- Not the spec.
precheckreturns 200 withcached: falseon it. A 68-char spec on an unrelated topic fails identically. - Not spec size. Mine is 1408 chars / ~352 tokens against
maxLength: 16000and the documented 5120-token limit. - Not rate limiting. That returns 429; this is a timeout or a 500.
- Not the client. Plain
httpxbypassing the SDK behaves the same.
Two docs issues found along the way
/api/v1/openapi.json404s; the schema is at/api/openapi.json. Worth a mention in the REST API docs.- Cache hits look like success.
"Classify sentiment as positive or negative"returns 202 instantly because it is cached from the docs' own example. I initially took that as evidence my spec was at fault. Acached: trueflag is already inCompileResponse— surfacing it in the SDK output, or noting it in the troubleshooting table, would save others the same wrong turn.
Suggestions
- Have
/api/v1/healthreport unhealthy (or awarningsentry) whengpu_servicesis empty, since compile cannot succeed in that state. - Return a 503 from
/api/v1/compileinstead of hanging when no compiler worker is available — the current behaviour is indistinguishable from a slow compile. - Consider making the SDK's 120s compile timeout configurable; there is currently no env override.
Environment
programasweights==0.4.4, Python 3.9, macOS. Anonymous (no PAW_API_KEY).
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reproducing the failure with the /api/v1/health, /api/v1/compile, and /api/v1/compile/async endpoints, then trace the SDK's paw.compile() timeout and CompileResponse handling. Check the REST API documentation and the /api/openapi.json entry point. Done should distinguish unavailable compiler workers from slow compiles and clarify the cached result in SDK or troubleshooting output.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- api, backend
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 45/100