Pull requests / #1581

#1581 CUDA prefill VRAM loans and SM75 Q4/native-IQ opt-ins

open · @c3c · 0 评论 · 在 GitHub 查看

BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindows

描述

# CUDA prefill VRAM loans and SM75 Q4/native-IQ opt-ins

tldr; massive PP speed win on my 8GB VRAM laptop GPU

## Summary

This branch adds two manual Turing prompt-compute paths and one manual,
request-scoped CUDA memory feature:

1. An SM75 Q4 prompt-attention tensor-core path, selected with
   `STRATA_PROMPT_ATTN_Q4_TC=1` and compiled with the qualified 64-cell tile.
2. A native IQ expert path for supported IQ3/IQ4 expert formats, selected with
   `STRATA_PF_FUSED_NATIVE_SM75=1`. The separately manual
   `STRATA_PF_FUSED_NATIVE_SM75_ADAPTIVE=1` enables the per-format SM75 shape
   probe; unsupported layers remain on the existing path.
3. Request-scoped VRAM loans for the serialized, single-NVIDIA-GPU server path.
   They temporarily use decode-only VRAM for the expert cache during prefill,
   then restore the verifier, native head and MTP resources before generation.
   A repeat-prefill option applies the same transaction to later long messages.

The Q4-attention and native-IQ paths are SM75-specific. The VRAM-loan feature
does not require SM75: it can apply to other NVIDIA CUDA GPUs with virtual
memory management and the supported single-GPU server configuration. The
compute-path measurements below use an RTX 2080 Super Max-Q; the loan is also
qualified separately on an RTX 4070 Ti.

All switches are off when absent. Existing behavior is therefore unchanged for
users who do not opt in.

## Benchmark platform

- GPU: NVIDIA GeForce RTX 2080 Super Max-Q, 8 GiB, 105 W limit, SM75
- CPU/RAM: Intel i7-10875H, 64 GiB RAM
- PCIe: PCIe 3.0 x16; measured host-to-device bandwidth about 12.2 GB/s
- OS/build: Windows, CUDA 13.4, MSVC, Release build with SM75 SASS
- Model: Qwen3.8-Flash-Next IQ3_XXS, native GGUF plus IQ3_XXS expert pack
- Context: 65,536 tokens
- MTP: enabled, `--spec 4`, same draft model and profile in each matched arm
- Speed projection: enabled and held constant
- Sampling: temperature 0 for benchmark requests
- PP/TG are the engine-reported rates from the timing line.

The long-prompt qualification used approximately 31,258 fresh prompt tokens
and a 512-token completion cap. Each run was serialized, started with a fresh
engine, and checked for normal completion. No parallel Strata servers were
used.

## Results

### End-to-end results

Each row is the median of three fresh-engine runs over 31,258 prompt tokens and
512 generated tokens. Baseline and opt-in runs use the same v0.1.41 source
revision, model, prompt and runtime settings.

| KV | Build | Prompt plan | PP (tok/s) | TG (tok/s) | PP range |
| --- | --- | ---: | ---: | ---: | ---: |
| Q4_0 | Feature flags off | 2,304 / 227 | 422.5 | 23.4 | 422.4--424.4 |
| Q4_0 | All applicable opt-ins | 4,864 / 512 | **865.2** | **23.6** | 863.5--881.0 |
| INT8 | Feature flags off | 1,792 / 227 | 435.8 | 22.9 | 435.6--436.1 |
| INT8 | Native IQ and loans | 4,352 / 512 | **854.7** | **21.2** | 849.3--865.5 |

The Q4_0 result is **+104.8% PP** with effectively unchanged TG. The INT8
result is **+96.1% PP** and **-7.4% TG**; its TG range was 21.1--23.1 and stayed
within the 10% guardrail. These are complete-stack results, not percentages for
individual features: the opt-ins also change prompt chunk, staging-ring and
expert-cache geometry.

The Q4_0 opt-in runs used:

```json
{
  "STRATA_BF16_TC": "1",
  "STRATA_PROMPT_ATTN_Q4_TC": "1",
  "STRATA_PF_FUSED": "1",
  "STRATA_PF_FUSED_NATIVE": "1",
  "STRATA_PF_FUSED_NATIVE_SM75": "1",
  "STRATA_PF_FUSED_NATIVE_SM75_ADAPTIVE": "1",
  "STRATA_PREFILL_ELASTIC_LOAN": "1",
  "STRATA_PREFILL_MTP_LOAN": "1",
  "STRATA_PREFILL_HEAD_LOAN": "1",
  "STRATA_PREFILL_RETAIN_STARTUP_CHUNK": "1",
  "STRATA_SELECT_SIMT": "1"
}
```

`STRATA_PREFILL_REPEAT_LOAN=1` is used only for the subsequent-message
qualification below; it does not participate in a fresh engine's first prompt.
For subsequent messages, `STRATA_PREFILL_REPEAT_MIN_TOKENS` sets the minimum
number of uncached prompt tokens that triggers releasing and reloading the
decode-only resources. It defaults to `1024` and is clamped to a minimum of
`256`. The threshold does not govern the first prompt: the startup MTP/head
loan is an engine-start on/off choice because those resources are deferred
before the first request length is known.

### Individual feature evidence

- **SM75 Q4 prompt attention:** with native IQ, loans and the 4,864/512 plan
  held constant, enabling only this path changed the three-run median from
  613.6 to 849.8 PP (**+38.5%**). QSA attention fell from about 19.2 to 5.1
  seconds.
- **SM75 native IQ:** the matched native-off median was 749.1 PP; the fixed
  native path measured 829.4 PP (**+10.7%**). This comparison also changed the
  plan from 5,376/227 to 4,864/512, so it is not a pure kernel-only percentage.
- **Prompt VRAM loans:** there is no clean three-run, same-planner loan-only
  comparison, so this PR does not assign the loan a percentage. Its observable
  effect is enabling the larger cache/chunk geometry and restoring all deferred
  decode resources before generation.

The existing `STRATA_SELECT_SIMT=1` option can compound with these changes: an
exact full-stack A/B changed 849.8 to 879.9 PP (**+3.54%**). SIMT is not part of
this PR.

### Subsequent-message loan

One serialized-session qualification used INT8 KV, a 131,072-token context and
128-MiB VRAM segments. After a 60-token warm-up, a 32,458-token
subsequent prompt selected 3,328/512 and completed at **787.0 PP / 23.1 TG**
with 512 generated tokens. The request released and restored the verifier,
native output head, MTP experts and draft head without a CUDA/runtime error.

For example, a configuration that only takes repeat loans when at least 50,000
new prompt tokens remain after prompt-cache reuse can use:

```json
{
  "STRATA_PREFILL_ELASTIC_LOAN": "1",
  "STRATA_PREFILL_MTP_LOAN": "1",
  "STRATA_PREFILL_HEAD_LOAN": "1",
  "STRATA_PREFILL_REPEAT_LOAN": "1",
  "STRATA_PREFILL_REPEAT_MIN_TOKENS": "50000"
}
```

### Cross-GPU VRAM-loan qualification

The loan is not SM75-specific. A separate test used an RTX 4070 Ti
(12 GiB, SM89), INT8 KV, a 196,608-token context, a 512-slot staging ring and
the same v0.1.41 source revision. Loan-off and loan-on runs were interleaved,
serialized, and started from a fresh engine. The 36k result is a three-run
median and the 72k result is a five-run median; both used a 512-token
completion cap.

| Fresh prompt | Loan | Prompt chunk | PP (tok/s) | TG (tok/s) | GPU prefill |
| ---: | --- | ---: | ---: | ---: | ---: |
| 36,061 tokens | Off | 8,448--8,960 | **2,562.9** | **54.8** | 13.811 s |
| 36,061 tokens | On | 15,104 | 2,539.4 | 52.9 | **12.987 s** |
| 72,061 tokens | Off | 8,960 | 2,617.0 | 54.1 | 27.271 s |
| 72,061 tokens | On | 15,104 | **2,680.9** | **54.9** | **25.772 s** |

The larger loan-backed chunk reduced median GPU prompt work by 6.0% at 36k
tokens and 5.5% at 72k. Loading the deferred output-head and MTP resources is
a fixed boundary cost, so the 36k request was 0.9% slower by the engine's
end-to-end PP rate while the 72k request was 2.4% faster. The five-run 72k PP
ranges were 2,604.3--2,633.7 off and 2,501.1--2,696.3 on; four of five loan-on
runs were 2,673.5 PP or higher, with one retained slow outlier. This is evidence
that the general CUDA loan works beyond Turing, but also that it should remain a
manual opt-in: benefit depends on prompt length and whether the released VRAM
crosses a useful planner threshold.

## Correctness and safety

- Default behavior is unchanged; all three families are manual opt-ins.
- The Q4 attention change is prompt-only.
- `STRATA_SELECT_SIMT=1` is also prompt-only, but uses a different FP32 tiled
  scorer and summation order; it can change selected cells and reasoning text.
- Native IQ dispatch is prompt-only and only supported formats use the native
  kernel; unsupported formats fall back to MMQ/previous behavior.
- Routing IDs, expert selection, route weights, and scatter order are not
  changed by these features.
- The loan path is NVIDIA CUDA-specific but not SM75-specific. It requires
  `--serve`, `--vram-elastic`, CUDA virtual memory management, one GPU and a
  VRAM segment of at least 64 MiB.
- The loan path is rejected for batch/shared sessions, pipeline windows,
  layer-split or multi-GPU configurations, helper caches, peer-device mode and
  HIP.
- Deferred verifier reload reapplies request sampling and token history.
- Native-IQ parity checks passed on the RTX 2080. Q4 prompt-attention parity
  passed at 4,096/512 and 1,500/512. The native
  IQ3_XXS/IQ4_NL test retained a fused-vs-rounded RMS of `4.811e-05`. The
  focused validation set passed 10/10, and the Release SM75 build completed
  without an error.
- A repeat-prefill request released and restored the verifier,
  native head, MTP routed experts and draft head, then completed 512 generated
  tokens without a CUDA/runtime error.
- A bounded visible-answer capture completed normally; the matched quality
  capture for native shape selection had identical bounded reasoning and draft
  acceptance. A separate SIMT quality pair returned the same visible control
  line and normal stop, but reasoning differed, as expected from the changed
  scorer arithmetic. A full unrestricted answer-parity sweep remains future
  work.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。