antirez / antirez/ds4

--layers maps the shard as N disjoint Metal buffers even when the range spans the whole file (~77x decode slowdown)

Đang mở
#845 2 bình luận 0 reaction 0 người được giao Xem trên GitHub
Ngôn ngữ chính
C
Star
22.4k
Fork
2.1k
Merge trung bình
2 ngày 13 giờ
Pull request đã merge (30 ngày)
5

Mô tả

---

## Summary

When `ds4-server` is started with `--layers`, the Metal model map is split into one shared buffer **per tensor span** instead of a single overlapping mapping. On the same machine, with the same binary and the same model family, this costs roughly **77x in decode throughput**.

The restriction appears to be a no-op in terms of *what* is mapped — the reported span is essentially the entire shard file — so the cost looks like pure mapping fragmentation rather than a smaller working set.

## Environment

- Apple Silicon Mac Studio, 512 GB unified memory, Metal backend
- Two hosts, pipeline parallelism (`--role coordinator` / `--role worker`)
- `--ctx 393216`, `--threads 8`, `--prefill-chunk 4096`
- Model split into two shards by layer range

## What the log shows

With `--layers`:

```
restricting metal model map to layers 0:30 (156 spans, 426.09 GiB tensor span)
...
mapped 436323.72 MiB from offset 5.06 MiB
...
156 disjoint shared buffers across 156 tensor spans
```

The second half of the split reports the same shape:

```
153 disjoint shared buffers across 153 tensor spans
```

**The restriction restricts nothing.** The shard file is 436,323.72 MiB ≈ **426.10 GiB**, and the mapped tensor span is **426.09 GiB** — the ~10 MiB difference matches the 5.06 MiB header offset. The engine is laying down the same bytes in 156 pieces instead of one.

Without `--layers`, on **the same machine and the same binary**, a single-process run of a full model maps as:

```
1 overlapping shared buffers
```

## Measured impact

| configuration | mapping | decode |
|---|---|---|
| single process, no `--layers` | `1 overlapping shared buffers` | **10.01 t/s** |
| coordinator, `--layers 0:30` | `156 disjoint shared buffers` | **0.13 t/s** |

Same box. Same binary. ~77x.

## The network is not involved

Per-token telemetry from the coordinator, repeated identically for every token:

```
distributed telemetry: request=2 hop=0 layers=31:60 pos=11 tokens=1
eval=2744.908ms downstream_wait=0.000ms forward_send=0.000ms
```

`downstream_wait` and `forward_send` are **exactly zero**, not merely small. Wire traffic is ~0.11 MiB in / ~0.49 MiB out per decode step. The time is entirely compute:

- worker, 30 layers: ~2,742 ms/token (σ < 7 ms over 8 tokens)
- coordinator, 31 layers: ~4,965 ms/token (by difference from 7,707 ms/token total)

That is roughly **91–160 ms per layer per token**, on hardware where the same engine runs a whole model at 33 t/s.

Both halves are equally affected, which rules out one bad host.

## Not memory, not cold cache

Residency time fell from 58,183 ms to 10,000 ms across a restart (warm file cache) and the decode rate did not change at all. Paging and load time are not the cause.

## Reproducibility

Reproduced on **two separate dates across five runs**, including a `--ssd-streaming` variant, which did not help:

```
day 1, 07:20 3 tokens / 1667.469 s = 0.002 t/s
day 1, 08:45 8 tokens / 266.808 s = 0.03 t/s
day 1, 08:54 7 tokens / 574.693 s = 0.01 t/s
day 2, 03:xx 8 tokens / 61.65 s = 0.13 t/s
```

The same `156 disjoint … 156 tensor spans` line is present in the day-1 logs, so this is not specific to a single launch or to startup ordering (restarting with corrected startup order changed nothing).

## Why `--tensor-parallel` is not a workaround here

`--role coordinator` rejects a start without `--layers`:

```
--role coordinator requires --layers
```

and `--tensor-parallel` requires the full model on both hosts rather than per-host shards, so on a sharded deployment there is no way to reach a non-fragmented mapping through the CLI.

## Suggested fix

Either coalesce adjacent tensor spans when building the restricted map, or skip the restriction entirely when the requested layer range already covers the whole file — which, per the numbers above, is the common case for a per-shard file.

## Minor, separate

The binary exits with **rc=0 on argument-validation errors**. `./ds4-server --role coordinator -m --ctx 4096` prints `--role coordinator requires --layers` and still returns 0, so scripts cannot detect a failed launch by exit status and must grep the log. Two probes cost nothing because validation runs before the single-instance guard and before model load — that part is good.

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Hướng nghiên cứu

Start at the --layers restricted-map construction and trace how layer ranges become Metal shared buffers. Reproduce the reported 156-disjoint-buffer case, compare it with the one-buffer full-model path, and verify mapping counts and decode throughput after coalescing spans or skipping a full-file restriction. Separately, inspect the argument-validation path for its exit status.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
c
Lĩnh vực
ai-infra-agents, performance
Loại issue
Lỗi
Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức độ hoạt động
Sôi nổi
Độ rõ ràng
Khá rõ ràng
Mức phù hợp với người mới
52/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.