Compile fails for all novel specs: health reports gpu_services:{}, finetune worker gets 500 from compiler on loopback

オープン
#5 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
45/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
活発
技術スタック
python
領域
api, backend

調査の方向性

Start by reproducing the failure with the /api/v1/health, /api/v1/compile, and /api/v1/compile/async endpoints, then trace the SDK's paw.compile() timeout and CompileResponse handling. Check the REST API documentation and the /api/openapi.json entry point. Done should distinguish unavailable compiler workers from slow compiles and clarify the cached result in SDK or troubleshooting output.

索引モデルが issue の本文から書いたものです。

説明

Summary

All compiles of novel specs have been failing for ~62 hours. /api/v1/health reports "status": "healthy" but with "gpu_services": {} — no GPU workers registered. The finetune worker's error surfaces the underlying cause: a 500 from the standard compiler on loopback.

Metadata endpoints (/models/compilers, /compile/precheck) work normally, so this looks like the vLLM compiler services rather than the API host.

Reproduction

import httpx
# Any spec that is not already cached:
httpx.post("https://programasweights.com/api/v1/compile",
           json={"spec": "Return ONLY the word RED or BLUE. Say RED if the input mentions fire."},
           timeout=110)
# -> httpx.ReadTimeout

68 characters, no unusual content. The SDK's paw.compile() hits the same thing and fails at its hardcoded timeout=120.0.

The useful error

/api/v1/compile only ever hangs, but /compile/async (finetune) returns a job whose error field is readable:

POST /api/v1/compile/async  {"spec": "...", "compiler": "paw-ft-bs48"}
-> 202 {"job_id": "...", "status": "queued", "compiler_snapshot": "paw-ft-bs48-20260530"}

GET /api/v1/compile/<job_id>
-> [   0.1s] compiling pct=0.0
-> [ 111.4s] failed
   error: Server error '500 Internal Server Error' for url
          'http://127.0.0.1:18000/api/v1/compile'

Reproduced twice, ~24h apart (first run failed at 322s, second at 111s).

What is and isn't working

Endpoint Result
GET /api/v1/health 200 healthy, but gpu_services: {}, queue_depth: 0
GET /api/v1/models/compilers 200
POST /api/v1/compile/precheck 200, cached: false, correct compiler_snapshot
POST /api/v1/compile ReadTimeout
POST /api/v1/compile/async + poll failed, loopback 500 above

uptime_s was 154214 → 223474 across my checks (~62h), so the API host has not restarted in that window.

Ruled out

  • Not the spec. precheck returns 200 with cached: false on it. A 68-char spec on an unrelated topic fails identically.
  • Not spec size. Mine is 1408 chars / ~352 tokens against maxLength: 16000 and the documented 5120-token limit.
  • Not rate limiting. That returns 429; this is a timeout or a 500.
  • Not the client. Plain httpx bypassing the SDK behaves the same.

Two docs issues found along the way

  1. /api/v1/openapi.json 404s; the schema is at /api/openapi.json. Worth a mention in the REST API docs.
  2. Cache hits look like success. "Classify sentiment as positive or negative" returns 202 instantly because it is cached from the docs' own example. I initially took that as evidence my spec was at fault. A cached: true flag is already in CompileResponse — surfacing it in the SDK output, or noting it in the troubleshooting table, would save others the same wrong turn.

Suggestions

  • Have /api/v1/health report unhealthy (or a warnings entry) when gpu_services is empty, since compile cannot succeed in that state.
  • Return a 503 from /api/v1/compile instead of hanging when no compiler worker is available — the current behaviour is indistinguishable from a slow compile.
  • Consider making the SDK's 120s compile timeout configurable; there is currently no env override.

Environment

programasweights==0.4.4, Python 3.9, macOS. Anonymous (no PAW_API_KEY).

主要言語
Python
スター
332
フォーク
31
平均マージ
2日 3時間
マージ済み PR(30日)
1

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

programasweights/programasweights-python のほかの issue

programasweights/programasweights-python の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。