Pull requests / #639

#639 layer split: with explicit split points each GPU loads only its own layers' dense weights

closed · @JeanP00l · 0 comentários · No GitHub

BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants

Descrição

## What

With a layer split every GPU keeps a full copy of the dense weights (~3.4 GB for the Coder), although each stage only reads its own layers. This PR lets each GPU load only its layers' dense weights when the split points are explicit (`--layer-split 27`), and gives the freed VRAM to the expert cache.

- The canonical arena of each GPU skips the `blk.N.*` tensors outside its layer range (read from the pack's `index.txt`). PLE tensors stay on every GPU.
- `NativeDense::set_layer_range` limits the native projections the next `load` uploads in the same way.
- `--layer-split auto` and single-GPU runs are unchanged. `STRATA_STAGE_TRIM=0` keeps the full copies.
- One log line per GPU says which layers it loads, and a later stage logs its free VRAM before its cache is sized.

## Measured

2x AMD Instinct MI50 16 GB (gfx906, see #638), Coder IQ1_M, 128K context, `--layer-split 27`:

| | full copies | own layers only |
|---|---|---|
| experts held in VRAM | 8,819 of 12,288 | 10,626 of 12,288 |
| VRAM hit rate in decode | 95% | 98% |
| decode, 17K-token agent prompt | 39.2 tok/s | 41.7 tok/s |

The trimmed tensors are never read by the GPU that skips them.

## Status

Re-run on 2x MI50 with this branch on current `main`, results in the comment below.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

No site

Links install, modelos, releases.