Compile fails for all novel specs: health reports gpu_services:{}, finetune worker gets 500 from compiler on loopback
Ninguém assumiu esta issue ainda.
Avaliação
- Dificuldade
- 4/5
- Tempo estimado
- 3-5 dias
- Facilidade para iniciantes
- 45/100
Direção de pesquisa
Start by reproducing the failure with the /api/v1/health, /api/v1/compile, and /api/v1/compile/async endpoints, then trace the SDK's paw.compile() timeout and CompileResponse handling. Check the REST API documentation and the /api/openapi.json entry point. Done should distinguish unavailable compiler workers from slow compiles and clarify the cached result in SDK or troubleshooting output.
Escrita pelo modelo de indexação a partir do texto da issue.
Descrição
Summary
All compiles of novel specs have been failing for ~62 hours. /api/v1/health reports "status": "healthy" but with "gpu_services": {} — no GPU workers registered. The finetune worker's error surfaces the underlying cause: a 500 from the standard compiler on loopback.
Metadata endpoints (/models/compilers, /compile/precheck) work normally, so this looks like the vLLM compiler services rather than the API host.
Reproduction
import httpx
# Any spec that is not already cached:
httpx.post("https://programasweights.com/api/v1/compile",
json={"spec": "Return ONLY the word RED or BLUE. Say RED if the input mentions fire."},
timeout=110)
# -> httpx.ReadTimeout
68 characters, no unusual content. The SDK's paw.compile() hits the same thing and fails at its hardcoded timeout=120.0.
The useful error
/api/v1/compile only ever hangs, but /compile/async (finetune) returns a job whose error field is readable:
POST /api/v1/compile/async {"spec": "...", "compiler": "paw-ft-bs48"}
-> 202 {"job_id": "...", "status": "queued", "compiler_snapshot": "paw-ft-bs48-20260530"}
GET /api/v1/compile/<job_id>
-> [ 0.1s] compiling pct=0.0
-> [ 111.4s] failed
error: Server error '500 Internal Server Error' for url
'http://127.0.0.1:18000/api/v1/compile'
Reproduced twice, ~24h apart (first run failed at 322s, second at 111s).
What is and isn't working
| Endpoint | Result |
|---|---|
GET /api/v1/health |
200 healthy, but gpu_services: {}, queue_depth: 0 |
GET /api/v1/models/compilers |
200 |
POST /api/v1/compile/precheck |
200, cached: false, correct compiler_snapshot |
POST /api/v1/compile |
ReadTimeout |
POST /api/v1/compile/async + poll |
failed, loopback 500 above |
uptime_s was 154214 → 223474 across my checks (~62h), so the API host has not restarted in that window.
Ruled out
- Not the spec.
precheckreturns 200 withcached: falseon it. A 68-char spec on an unrelated topic fails identically. - Not spec size. Mine is 1408 chars / ~352 tokens against
maxLength: 16000and the documented 5120-token limit. - Not rate limiting. That returns 429; this is a timeout or a 500.
- Not the client. Plain
httpxbypassing the SDK behaves the same.
Two docs issues found along the way
/api/v1/openapi.json404s; the schema is at/api/openapi.json. Worth a mention in the REST API docs.- Cache hits look like success.
"Classify sentiment as positive or negative"returns 202 instantly because it is cached from the docs' own example. I initially took that as evidence my spec was at fault. Acached: trueflag is already inCompileResponse— surfacing it in the SDK output, or noting it in the troubleshooting table, would save others the same wrong turn.
Suggestions
- Have
/api/v1/healthreport unhealthy (or awarningsentry) whengpu_servicesis empty, since compile cannot succeed in that state. - Return a 503 from
/api/v1/compileinstead of hanging when no compiler worker is available — the current behaviour is indistinguishable from a slow compile. - Consider making the SDK's 120s compile timeout configurable; there is currently no env override.
Environment
programasweights==0.4.4, Python 3.9, macOS. Anonymous (no PAW_API_KEY).
- Linguagem predominante
- Python
- Estrelas
- 332
- Forks
- 31
- Merge médio
- 2d 3h
- PRs com merge (30d)
- 1
Guia de contribuição
Nenhum guia de contribuição indexado para este repositório
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de programasweights/programasweights-python
-
Dificuldade 5/5 Mais de uma semana Facilidade para iniciantes 25/100
Todas as issues de programasweights/programasweights-python
Issues semelhantes
-
bug
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 86/100
zostera/django-bootstrap4#894 ·
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 78/100
use-agent-os/agent-os#3276 ·
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 88/100
zephyrproject-rtos/zephyr#119726 ·
-
area/auth bug comp/agent P3 platform/discord type/security
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 88/100
NousResearch/hermes-agent#117848 ·
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 82/100
zilliztech/memsearch#759 ·