Pull requests / #1523
#1523 LoRA adapter support with per-request on/off switch
open · @agrogov · 0 comments · View on GitHub
BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentation
Description
## Summary
Optional LoRA adapter support in the CUDA engine (llama.cpp's adapter GGUF format), with a per-request on/off switch.
- `--lora FILE` / `--lora-scaled FILE:SCALE[,...]` load a llama.cpp LoRA GGUF (`general.type = adapter`, `adapter.type = lora`; f32/f16/bf16). It is applied to the mixer output projections: `ssm_out` on GDN layers and `attn_output` on QSA layers. For each adapted layer, `y += s·B(A·x)` with `s = scale·alpha/rank`; several files stack their ranks.
- **Per request:** the engine key is `lora=0|1`, and the server takes `"lora": true|false` in the request body (OpenAI and Anthropic APIs) or `"sampling": {"lora": ...}` in the config. The switch is a device flag the kernels read, so the captured CUDA graphs serve both states with no re-capture. Turned off, the output is bit-identical to the engine without an adapter.
- **State:** the conversation cache, batch slots and session files are keyed by the switch, next to `cvec`. Session files stay backward compatible: a new bit in the existing flag word, and the fingerprint changes only when an adapter is loaded. While batch slots are decoding, a request for the other state is refused.
- **Not adapted:** the MTP draft layer (it only drafts; the verified tokens are the adapted model's).
## Commits
1. `kernels`: the LoRA kernels plus the `lora_parity` test.
- Decode/verify: two kernels per 8 tokens with per-layer scratch; capturable.
- Prompt chunks: tiled, split-K.
2. `core`: hooks in decode (`layer.cpp`), verify (`verify.cpp`) and the prompt path (`prefill.cpp`).
3. `conversation`: cache/session keying.
4. `generate`: loader, CLI flags, request key, layer-split replication.
5. `serve`: the `"lora"` request field, plus a unit test.
6. `perf(lora)`: the decode/verify up kernel is a programmatic dependent launch. It starts while the down kernel runs and loads its column of B, then waits for h. Rank > 64 and HIP use the plain launch. `lora_parity` also times a captured 48-layer graph.
## Test results
Tested on current `main` (v0.1.40.4, `6674a00`) with this branch applied, and before that on v0.1.40.2 with the same results. One NVIDIA B200, CUDA 12.8, sm_100. Model: Qwen4Exp IQ2_XS. Adapter: rank 50 on all 48 layers, 80 MiB VRAM.
- **`lora_parity`:** all checks pass.
- Error against double precision ≤ 8e-7 for 1–300 tokens, fp32 and fp16 activations.
- Off is bitwise unchanged, eager and from a graph captured while on.
- Switching back on re-applies through the same graph.
- **End to end:** one engine answers `lora=1`, `lora=0`, `lora=1`; a second engine without an adapter answers the same prompts. Greedy, with and without MTP:
- `lora=0` gives the same tokens as no adapter.
- Both `lora=1` runs give the same tokens.
- `lora=1` differs from `lora=0` where the adapter changes behaviour.
- **Speed** (256 generated tokens, prompts of about 2,700 tokens, 3 runs each; no MTP is the mean of 3 such benchmarks):
| | decode, MTP | decode, no MTP | prompt read / 1K tokens |
|---|---|---|---|
| no adapter | 213.2 tok/s | 118.2 tok/s | 216 ms |
| loaded, `lora=0` | 210.4 (−1.3%) | 114.9 (−2.8%) | 209 ms |
| loaded, `lora=1` | 206.1 (−3.3%) | 111.7 (−5.5%) | 227 ms (+5%) |
Before commit 6, `lora=1` ran at 202.2 tok/s with MTP and 110.1 without. Each setting reads different prompt texts, so prompt read moves by a few percent between settings even where the prompt path is unchanged.
- **Latency** (derived from the same runs: decode as ms per generated token, prompt read as the mean wall time for prompts of about 2,700 tokens, normalized to 1K tokens):
| | per token, MTP | per token, no MTP | prompt read / 1K tokens, MTP | prompt read / 1K tokens, no MTP |
|---|---|---|---|---|
| no adapter | 4.69 ms | 8.46 ms | 216 ms | 216 ms |
| loaded, `lora=0` | 4.75 ms (+0.06) | 8.70 ms (+0.24) | 209 ms | 209 ms |
| loaded, `lora=1` | 4.85 ms (+0.16) | 8.95 ms (+0.49) | 229 ms | 227 ms |
- **Kernel latency per adapted layer** (`lora_parity`, rank 50, 6144 → 2560; there are 48 adapted layers per forward pass). Decode replays its layers from CUDA graphs, so the in-graph figure is the one decode pays:
| tokens per call | path | one call | in a 48-layer graph |
|---|---|---|---|
| 1 | decode | 6.2 µs (was 10.2) | 4.74 µs (was 6.06) |
| 4 | verify window | 10.2 µs (was 12.3) | |
| 8 | verify window | 13.3 µs (was 16.4) | 11.48 µs (was 13.05) |
| 300 | prompt chunk | 47.1 µs | |
- **`serve/test_server.py`:** the new `test_lora_key` passes. The 3 errors that remain also occur on the untouched base: a socket timeout and Python 3.9 `staticmethod`.
## Not included
- No SYCL/HIP port, no web UI toggle, no VRAM accounting in the layer-split auto planner (80 MiB at rank 50).
- Docs (`docs/DETAILS.md`) and the benchmark scripts/results are kept out of this PR; they live on the fork's `feat/lora-adapter` branch.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.