Pull requests / #639
#639 layer split: with explicit split points each GPU loads only its own layers' dense weights
closed · @JeanP00l · 0 comments · View on GitHub
BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants
Description
## What With a layer split every GPU keeps a full copy of the dense weights (~3.4 GB for the Coder), although each stage only reads its own layers. This PR lets each GPU load only its layers' dense weights when the split points are explicit (`--layer-split 27`), and gives the freed VRAM to the expert cache. - The canonical arena of each GPU skips the `blk.N.*` tensors outside its layer range (read from the pack's `index.txt`). PLE tensors stay on every GPU. - `NativeDense::set_layer_range` limits the native projections the next `load` uploads in the same way. - `--layer-split auto` and single-GPU runs are unchanged. `STRATA_STAGE_TRIM=0` keeps the full copies. - One log line per GPU says which layers it loads, and a later stage logs its free VRAM before its cache is sized. ## Measured 2x AMD Instinct MI50 16 GB (gfx906, see #638), Coder IQ1_M, 128K context, `--layer-split 27`: | | full copies | own layers only | |---|---|---| | experts held in VRAM | 8,819 of 12,288 | 10,626 of 12,288 | | VRAM hit rate in decode | 95% | 98% | | decode, 17K-token agent prompt | 39.2 tok/s | 41.7 tok/s | The trimmed tensors are never read by the GPU that skips them. ## Status Re-run on 2x MI50 with this branch on current `main`, results in the comment below. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.