Compile fails for all novel specs: health reports gpu_services:{}, finetune worker gets 500 from compiler on loopback

Aberta
#5 1 comentário 0 reações 0 responsáveis Ver no GitHub

Ninguém assumiu esta issue ainda.

Avaliação

Dificuldade
4/5
Tempo estimado
3-5 dias
Facilidade para iniciantes
45/100
Tipo de issue
Bug
Clareza
Razoavelmente clara
Status de atividade
Ativa
Stack de tecnologia
python
Domínio
api, backend

Direção de pesquisa

Start by reproducing the failure with the /api/v1/health, /api/v1/compile, and /api/v1/compile/async endpoints, then trace the SDK's paw.compile() timeout and CompileResponse handling. Check the REST API documentation and the /api/openapi.json entry point. Done should distinguish unavailable compiler workers from slow compiles and clarify the cached result in SDK or troubleshooting output.

Escrita pelo modelo de indexação a partir do texto da issue.

Descrição

Summary

All compiles of novel specs have been failing for ~62 hours. /api/v1/health reports "status": "healthy" but with "gpu_services": {} — no GPU workers registered. The finetune worker's error surfaces the underlying cause: a 500 from the standard compiler on loopback.

Metadata endpoints (/models/compilers, /compile/precheck) work normally, so this looks like the vLLM compiler services rather than the API host.

Reproduction

import httpx
# Any spec that is not already cached:
httpx.post("https://programasweights.com/api/v1/compile",
           json={"spec": "Return ONLY the word RED or BLUE. Say RED if the input mentions fire."},
           timeout=110)
# -> httpx.ReadTimeout

68 characters, no unusual content. The SDK's paw.compile() hits the same thing and fails at its hardcoded timeout=120.0.

The useful error

/api/v1/compile only ever hangs, but /compile/async (finetune) returns a job whose error field is readable:

POST /api/v1/compile/async  {"spec": "...", "compiler": "paw-ft-bs48"}
-> 202 {"job_id": "...", "status": "queued", "compiler_snapshot": "paw-ft-bs48-20260530"}

GET /api/v1/compile/<job_id>
-> [   0.1s] compiling pct=0.0
-> [ 111.4s] failed
   error: Server error '500 Internal Server Error' for url
          'http://127.0.0.1:18000/api/v1/compile'

Reproduced twice, ~24h apart (first run failed at 322s, second at 111s).

What is and isn't working

Endpoint Result
GET /api/v1/health 200 healthy, but gpu_services: {}, queue_depth: 0
GET /api/v1/models/compilers 200
POST /api/v1/compile/precheck 200, cached: false, correct compiler_snapshot
POST /api/v1/compile ReadTimeout
POST /api/v1/compile/async + poll failed, loopback 500 above

uptime_s was 154214 → 223474 across my checks (~62h), so the API host has not restarted in that window.

Ruled out

  • Not the spec. precheck returns 200 with cached: false on it. A 68-char spec on an unrelated topic fails identically.
  • Not spec size. Mine is 1408 chars / ~352 tokens against maxLength: 16000 and the documented 5120-token limit.
  • Not rate limiting. That returns 429; this is a timeout or a 500.
  • Not the client. Plain httpx bypassing the SDK behaves the same.

Two docs issues found along the way

  1. /api/v1/openapi.json 404s; the schema is at /api/openapi.json. Worth a mention in the REST API docs.
  2. Cache hits look like success. "Classify sentiment as positive or negative" returns 202 instantly because it is cached from the docs' own example. I initially took that as evidence my spec was at fault. A cached: true flag is already in CompileResponse — surfacing it in the SDK output, or noting it in the troubleshooting table, would save others the same wrong turn.

Suggestions

  • Have /api/v1/health report unhealthy (or a warnings entry) when gpu_services is empty, since compile cannot succeed in that state.
  • Return a 503 from /api/v1/compile instead of hanging when no compiler worker is available — the current behaviour is indistinguishable from a slow compile.
  • Consider making the SDK's 120s compile timeout configurable; there is currently no env override.

Environment

programasweights==0.4.4, Python 3.9, macOS. Anonymous (no PAW_API_KEY).

Linguagem predominante
Python
Estrelas
332
Forks
31
Merge médio
2d 3h
PRs com merge (30d)
1

Guia de contribuição

Nenhum guia de contribuição indexado para este repositório

Primeiros passos

  1. Leia a issue inteira e depois o guia de contribuição do projeto.
  2. Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
  3. Faça um fork do repositório e trabalhe em uma branch.
  4. Abra um pull request que referencie o número da issue.

Mais de programasweights/programasweights-python

Todas as issues de programasweights/programasweights-python

Issues semelhantes

Mais issues de Python

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.