antirez / antirez/ds4

GLM-5.3 branch: compaction fails when think-mode generation overshoots the soft limit — analysis + 4 local patches

Đang mở
#906 1 bình luận 0 reaction 0 người được giao Xem trên GitHub
Ngôn ngữ chính
C
Star
22.3k
Fork
2.1k
Merge trung bình
1 ngày 3 giờ
Pull request đã merge (30 ngày)
4

Mô tả

```markdown
GLM-5.3 branch: compaction fails when think-mode generation overshoots the soft limit — analysis + 4 local patches

### GLM-5.3 branch: compaction fails ("not enough context left") when a long
think-mode generation overshoots the soft limit — analysis + 4 local patches

**Environment**
- Branch: `glm-5.3-flash` (a60a2a0), built from source, macOS, M3 Ultra 256 GB
- Model: GLM-5.3-Flash Q4_K GGUF, agent mode, `-c 100000`, think mode = HIGH
- Workload: long multi-step engineering analysis (heavy per-turn thinking)

**Observed failure**

After a long session the agent hit:

COMPACTING soft limit before user turn: summarizing durable task state
ds4-agent: not enough context left to request compaction summary

and the live session was invalidated. The auto-save line confirmed the session
died at **99999 / 100000 tokens**.

**Root cause (from reading ds4_agent.c on this branch)**

1. `agent_worker_should_compact()` runs *before* generation and measures
`w->transcript.len`.
2. In think mode the generated thinking tokens are appended to the transcript
*after* that check. GLM-5.3 at HIGH effort regularly emits 10k+ thinking
tokens per round on hard analysis tasks, and there is no thinking budget.
3. So one generation can push occupancy from ~84% to ~100% between checks.
When compaction finally fires, `summary_room = ctx - prompt.len - 1 < 256`
and `agent_worker_compact()` invalidates the session.

DeepSeek V4 Flash hides this because its thinking is much shorter; GLM-5.3's
longer, more thorough reasoning exposes the gap. The first few compactions
worked fine — it only breaks once thinking volume per round grows.

**Local patches (all in ds4_agent.c, ~60 lines total, running in production here)**

- **P1 — trigger on the real ledger.** `should_compact` uses
`max(transcript.len, ds4_session_pos(...))`, and the soft percentage is
env-tunable (`DS4_AGENT_COMPACT_SOFT_PCT`, default still 85).
- **P2 — reserve compaction fuel per generation.** Clamp each round's
`max_tokens` to `room - 1 - max(AGENT_COMPACT_MIN_FREE_TOKENS, room/6)` so a
single generation can never consume the fuel the next compaction needs.
- **P3 — degrade instead of dying.** If `summary_room < 256`, summarize only
the recent tail (last half of the transcript) instead of invalidating the
session.
- **P4 — mechanical thinking-strip before LLM compaction.** At the soft limit,
remove completed think-block spans from the transcript (keep the most recent
block verbatim for continuity), rebuild the session via
`agent_worker_sync_tokens` — no model call at all — and only fall through to
LLM summarization if pressure remains.

**Production result**

THINK-STRIP dropping 60316 old thinking tokens: 98561 -> 38245
(re-prefill of 38,245 tokens at ~405 tok/s, ~94 s, session continued)

After the patches, repeated hard thinking rounds no longer kill the session;
"not enough context left to request compaction summary" has not recurred.

Finally — a sincere thank you. This branch turned a single 256 GB Mac into a
machine that runs serious, hours-long local reasoning for real engineering
work every day — no cloud bill, no quota, full privacy. Making
state-of-the-art models run well on ordinary people's hardware, and keeping
pushing the compute/quality balance forward, is genuinely democratizing work.
Happy to give something back with these patches.

Happy to clean this up and open a PR if the approach looks acceptable.
Everything above was observed on the `glm-5.3-flash` branch; I have not
reproduced it on `main` with DeepSeek models (their thinking is shorter).
Related discussion: #868.

```

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.