Pull requests / #1273
#1273 serve: session files work with --peer-device
open · @pspranger-throw · 0 comments · View on GitHub
Server & APINVIDIA / CUDAModels & quants
Description
The blanket refusal in the SAVE/RESTORE path is gone. The peer tier is a pure in-memory expert cache rebuilt at load from the expert source files (the model fingerprint's inputs); it holds no session state, so nothing is missing from a saved file and no per-run peer artifact can drift. Verified live: save and restore a conversation with --peer-device 1 active. Tested live on an RTX 3090 + RTX 2060S box, IQ3_S 3.5 bpw, `--peer-device 1` active (the peer filled 3,571 experts / 6.78 GiB and rebalanced during the runs): - Shallow cycle: a 10,462-token session saved through the slot API (396 MB, 314 ms); the engine stopped and started; restored in 423 ms; the next turn reused 10,495 of 10,515 tokens (99.8%) in 0.8 s, where the cold prefill was ~13 s. - Deep cycle, same peer shape: a 198,830-token session (3.05 GiB) saved in 1.4 s, restored at start in 3.5 s; the next turn reused 198,859 of 198,879 tokens (99.99%) in 1.8 s, where the cold read was ~177 s. - Through a gateway switch (39,060 tokens, peer active): 833 MB saved in 959 ms, restored in 1.0 s, 99.95% reused in 0.7 s. - End-to-end workflow: while the main IQ3_S session was live, sub-agents ran on other engines (different local models) that took over the GPUs and RAM; switching back to the main session was always warm.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.