贡献 / #1600

#1600 Layer split: with `STRATA_SPLIT_RING` set, every borrowed slot keeps its RAM copy (a 104k prompt read no longer pulls 10 GB from the model file)

open · @noon-at-cgn · 0 评论 · 去 GitHub 看

BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

说明

## Title
Issue: Related: #1190 (this PR depends on it and is branched from its head; please merge it first). Near it: #1417 changes `part_slots` in the same area of `generate.cpp`.

## Summary
**Why it matters:** with `STRATA_SPLIT_RING=384` the prompt path borrowed 5,396 cache slots but the RAM copy kept only 4,662. Each long prompt then re-read the other 734 slots' experts from the model file in every chunk. The cause is an ordering bug: the lend regions are sized with `Prefill::bytes_needed` before the ring override is applied, although `Prefill::set_ring_override` says it must come first. Now the ring is decided first, and a start-up line shows borrowed against kept so a mismatch is visible.

| 2x RTX 3080, layer split 23, 104k-token prompt, resident RAM mode | before | after |
|---|---|---|
| slots borrowed / kept in RAM | 5,396 / 4,662 | 5,396 / 5,396 |
| blob reads from the model file per request | 413 | 0 |
| expert data read from the GGUF per request | 10,345 MB | 0 MB |
| pinned RAM for the copies | 13.61 GiB | 15.76 GiB (+2.15) |

Nothing changes without a split or without `--resident-experts`. With the override unset the lend lines are identical before and after.

## What changed
- `src/program/generate.cpp`: the ring choice moves into `decide_split_ring()`, called before the stage lend regions are sized (the old call site stays and does nothing a second time). The tail-slot count for a region and for the borrow is one function. New line: `strata serve: lend sizing: the prompt path borrows N slots (ring R slots at chunk C), K of them keep their experts in RAM too`, with a `WARNING` when K is below N.
- `include/strata/program/split_ring.hpp` (new): `split_ring_slots`, `split_lend_slots`. `src/program/split_ring_test.cpp` (new) + `CMakeLists.txt`: CPU test (23 checks): the copy keeps every borrowed slot with the ring decided first, not with the old order.
- `docs/MULTI_GPU.md`: one paragraph.

## Extra Notes
Machine: 2x RTX 3080 20 GB (220 W), Xeon E5-2696 v4, 90 GiB container, UD-Q4_K_XL, `--prefill auto:16384`, one restart per arm; the patched binary has none of our other patches. MemAvailable minimum during the reads: 11.98 GiB with the fix, 13.34 GiB without, both above the 2 GiB floor.
**Not claimed:** a speed gain. Medians of three 104k reads were 2,803 (before) and 2,907 (after) tok/s, but the before arm's spread is 2,540-2,897 (its first read faulted 4.2 GB in from NVMe) and it has no A/A restart. 25k and 51k reads are equal within 1-2% (they touch none of the unkept slots). On a box with less spare page cache the 10 GB per request would be disk traffic; here the page cache absorbed it.
The setting our own deployment uses, `STRATA_PREFILL_RING=384`, was never affected. 413 = 7 x 59 (7 equal chunks) is our inference, not logged. Not tested: both ring variables unset.
Checked: `split_ring_test`, #1190's `file_expert_source_test` and `split_lend_complement_test`, CUDA build on #1190's base. Not built: HIP, SYCL. This is a bug fix, not an opt-in: with a split and the resident RAM mode the sizing order changes for every ring choice.

本站相关内容

相关页面的快捷入口。