Issues / #1604

#1604 `--batch-groups` disables the per-slot conversation cache even when the group size is 1 (no pad row is ever written) → alternating long conversations re-prefill from token 0 every turn

open · @113636xfh · 0 comentários · No GitHub

BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsSecurityDocumentationWindowsLinux

Descrição

Engine: 0.1.40.1 (`project(strata VERSION 0.1.40)`), Linux, CUDA 12.9.
Code references are to `src/program/generate.cpp` and `include/strata/core/verify.hpp` of this build.

Related: #1413 ("`--batch-groups G>1` + `--batch-mtp` is silently inert") — same pipelined path, different symptom. PR #1601 (a memory guard so parking checks the RAM the container really has) — §5.5 is the same gate failing on a bare-metal host, and our measured park sizes are far above the 240–750 MB quoted there. #1603 (two long-context sessions under `parallel: 2`) — likely the same root cause seen from the client side.

---

## 1. Our machine

| Item | Value |
| --- | --- |
| Board / CPU | X99-CH8, Xeon E5-2673 v4 (20C/40T @ 2.30 GHz) |
| GPUs | **3× V100-SXM2-16GB** (SM 7.0) in use — GPU0/1/2, **no NVLink**, power-capped to **150 W** each |
| PCIe (measured) | GPU0 Gen3 x16, GPU1 Gen3 x8, GPU2 Gen3 x8; all PHB; P2P works but only 2.6–3.3 GB/s |
| GPU memory bandwidth | ~818 GB/s D2D per card (≈91 % of V100 peak) |
| Host RAM | 94.2 GiB, 4-channel DDR4-2133 ECC; STREAM triad 42.5 GB/s @ 8 threads |
| Model | Qwen3.8-Flash-Next GSQ RCO **IQ3_S** (MoE, 48 layers, 12 QSA + 36 GDN) + MTP draft layer |

## 2. Our launch configuration

`strata-iq3_s.json` → `serve/server.py --engine strata --port 8100`, `"gpu": [0,1,2]`,
`"parallel": 4` (→ `--batch 4`), `"layer_split": [16,32]`:

```
--serve
--pack /strata/data/packs/iq3_s
--native  .../Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf
--ple-gguf .../Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf
--expert-profile data/expert-profile.bin
--expert-cache auto
--prefill 4096
--spec 4 --spec-min-p 0.5
--mtp /strata/data/mtp/rt
--max-context 262144
--kv int8
--kv-resident 32768
--vision
--vram-reserve-mib 1536
--vram-reserve-later-mib 1024
--ple-row-cache 16777216
--trim-stage-weights
--conversation-cache-mib 8192
--conversation-cache-slots 4
--conversation-cache-min-free-mib 6000
--batch-groups 4
--layer-split 16,32
--batch 4
```

Startup evidence that the memory side is healthy (VRAM is **not** the problem):

```
KV streaming: 32768 of 262144 cells per QSA layer in VRAM, the K/V in 1.03 GiB of pinned RAM
--batch: 4 slot sessions on CUDA0 (0.36 GiB each); 10.94 GiB free
--batch: 4 slot sessions on CUDA1 (0.36 GiB each); 11.96 GiB free
--batch: 4 slot sessions on CUDA2 (0.36 GiB each); 11.49 GiB free
--batch 4: the slot sessions take 1.46 GiB of VRAM on CUDA0 that the expert cache would otherwise hold
expert cache auto: 10.83 GiB free, 1536 MiB reserved (+0 MiB for the draft head) -> 3762 slots
expert arena: cudaHostRegister PORTABLE ok (needed 23983 2 MiB pages)
```

## 3. Exact conditions that trigger it

1. `--batch N` **and** `--batch-groups G` with **G == N**, so `GS = N / G == 1` (one slot per pipeline group).
2. Two conversations, both long (we reproduce from ~100 K tokens up), **alternating turn by turn** — the classic controller/worker or two-Codex-sessions pattern.
3. Both conversations stay in the same slots (we see slot 0 and slot 1 in the log).

Under those conditions the conversation that is *not* currently live in the main session loses its state every single turn and re-reads its whole prompt from token 0.

## 4. Symptom, verbatim from the engine log

```
conversation cache: restored 60021 tokens (checkpoint) in 136.7 ms; parked=3 bytes=5091272344
strata batch: slot 0 takes 102008 tokens (copied in 689.2 ms)
strata serve: prompt 102008 tokens = 60021 reused + 41987 read in 32687 ms (1284.5 tok/s), 1 generated
conversation cache: skip parking (physical RAM admission; need 1857 MiB plus 6000 MiB floor, or telemetry unavailable)
strata serve: prompt 45615 tokens = 44156 reused + 1459 read in 5410 ms (269.7 tok/s), 1 generated
strata batch: slot 1 takes 45616 tokens (copied in 323.2 ms)
strata serve: prompt 45616 tokens = 45615 reused + 1 read in 9 ms (109.8 tok/s), 1 generated
conversation cache: parked 45616 tokens in 113.8 ms; parked=3 bytes=4163177756 evictions=66 snapshot_bytes=1897425548 reused_kv_bytes=673992960
strata batch: slot 0 takes 103292 tokens (copied in 682.0 ms)
strata serve: prompt 103292 tokens = 0 reused + 103292 read in 68522 ms (1507.4 tok/s), 1 generated
```

and then every following turn of that same conversation:

```
prompt 104471 tokens = 0 reused + 104471 read in 70114 ms
prompt 104646 tokens = 0 reused + 104646 read in 70439 ms
prompt 105596 tokens = 0 reused + 105596 read in 70990 ms
prompt 106550 tokens = 0 reused + 106550 read in 71008 ms
prompt 107435 tokens = 0 reused + 107435 read in 71175 ms
prompt 107791 tokens = 0 reused + 107791 read in 71238 ms
prompt 108074 tokens = 0 reused + 108074 read in 71455 ms
prompt 108793 tokens = 0 reused + 108793 read in 72974 ms
```

The short conversation (45 K tokens) keeps working normally throughout; only the long one is destroyed.

Aggregate over one run (log lines 1661–9527, 1803 prompts):

| Metric | Value |
| --- | --- |
| prompts with `0 reused` | **71 / 1803 (3.9 %)** |
| total prompt-read time | 7 230 s |
| …of which spent on `0 reused` prompts | **2 923 s = 48.7 min** |
| `skip parking (physical RAM admission …)` | 23 |
| `skip parking (snapshot … exceeds available budget)` | 8 |
| `evictions` | 66 |
| **`slot N gave back … tokens of this conversation`** | **0 occurrences in the whole `--batch-groups 4` run** |
| same message in earlier runs without `--batch-groups` | 22 occurrences |

So the slot conversation cache is provably never used once `--batch-groups` is on.

## 5. Our analysis

### 5.1 The pipelined path clears `cached` unconditionally

```cpp
// generate.cpp:7619
const bool piped = o.batch > 0 && o.batch_groups > 1 && n_pipe > 1;
const int GS = piped ? o.batch / o.batch_groups : o.batch;

// generate.cpp:7672  (a slot finishing in the pipelined path)
sl.active = false;
sl.cached = false;   // the pipeline's pad rows: a pipelined slot is not reused as a cache

// generate.cpp:7705  (a group starting)
if (!sl.active) sl.cached = false;   // its pad row writes its state
```

versus the non-pipelined path at the same point:

```cpp
// generate.cpp:7575
sl.cached = o.prompt_cache > 0 && !sl.img;
```

and the resume source skips any slot with `cached == false`:

```cpp
// generate.cpp:8242
if (sl.active || !sl.cached || sl.cvec != want_cvec) continue;
```

### 5.2 The stated reason (pad rows) is unreachable when `GS == 1`

A group is only ever started if it has an active slot:

```cpp
// generate.cpp:7699
if (!pg[gi].inflight && group_active(gi)) pick = gi;
// group_active(gi) = any bs[gi*GS + t].active
```

With `GS == 1`, `group_active(gi)` **is** `bs[gi].active`, so inside

```cpp
for (int t = 0; t < GS; ++t) {
    BSlot& sl = bs[(size_t) (pick * GS + t)];
    G.tok[t] = sl.active ? sl.x : 0;
    G.pos[t] = sl.active ? sl.p : 0;
    if (!sl.active) sl.cached = false;   // never taken at GS == 1
}
```

the `!sl.active` branch **cannot execute**, and `batch_launch(pick*GS, GS, …)` launches exactly the one active slot's row. No idle slot's state is ever written. Yet `:7672` still clears `cached` for every finished slot.

### 5.3 The real constraint is the hand-off row addressing, not KV independence

`include/strata/core/verify.hpp` contrasts the two window APIs:

```cpp
/// The same over the S slots `rows` (row t is slot rows[t], any distinct slots in any order):
/// the slots not listed are not touched, so an idle slot keeps its state
/// (a finished conversation it may continue later).
bool run_slot_rows(const int* rows, int S, ...);          // non-pipelined: this is what makes the slot cache work

// ---- The stages of a layer split as a PIPELINE.  A batch window over the slot GROUP
// [base, base + S) is launched on ONE stage ...
// Rows of group `base` use hand-off rows [base, base + S), so groups never share a hand-off row.
bool batch_launch(int base, int S, ...);                  // pipelined: contiguous group only
```

So the coupling is that `batch_launch` addresses a **fixed contiguous group** (to keep hand-off rows disjoint between groups), which forces padding of idle slots inside an active group. That is an addressing limitation, not a conflict with per-conversation KV.

### 5.4 The per-slot KV is independent and already allocated — the cache is disabled for free

`qsa_state_init` (`src/core/layer.cpp:802`) allocates the authoritative K/V per QSA state whenever KV streaming is on:

```cpp
if (p.mode != 0) {
    cudaHostAlloc((void**) &h, pages * kv_block_bytes(...), cudaHostAllocMapped | cudaHostAllocPortable);
    g_kv_host_bytes += bytes;
```

and `--batch` carves a full session per slot per stage (`generate.cpp:3512`). On our box that is **1.03 GiB × 3 stages × 4 slots = 12.4 GiB of pinned host RAM, resident from startup, whether or not the slot cache is ever used.** `copy_to_slot` / `copy_from_slot` already walk every stage's slot sessions, so the pipelined path writes exactly the sessions `copy_from_slot` reads back.

**Enabling the slot cache at `GS == 1` therefore costs no additional RAM or VRAM.**

### 5.5 Why the fallback then collapses at long context

With the slot cache off, the only resume source left is the main session's `live`/`checks` plus `ConversationCache`, and both of its gates fail for long conversations:

* **Byte budget.** Measured `snapshot_bytes` fits ≈ **1.2 GiB fixed** (`live` + up to `--prompt-cache 6` checkpoints × ~118 MB) **+ ~14.4 KB per token**:

  | tokens | snapshot_bytes |
  | ---: | ---: |
  | 44 156 | 1 897 335 584 |
  | 120 323 | 2 676 582 000 |
  | 150 910 | 3 145 981 824 |
  | 210 276 | 4 348 089 744 |

  Two conversations therefore both fit in `--conversation-cache-mib 8192` only below **≈ 206 K tokens each = 78.6 % of `--max-context 262144`**. Above that, `make_room()` evicts the other conversation oldest-first. (This matches the reported "both sessions at 80 %" threshold exactly.)

  One figure relevant to PR #1601, which quotes "about 240 MB to 750 MB per park": at these context lengths our parks are **1.9 GiB at 44 K tokens and 4.3 GiB at 210 K tokens**, i.e. 3–18× that. The ~1.2 GiB fixed part (the `live` state plus up to `--prompt-cache 6` checkpoints) is what dominates below ~90 K tokens.

* **Physical-RAM admission.** `conversation_memory_admit(MemAvailable, additional, floor)` requires `MemAvailable ≥ 6 000 MiB + additional`. Our engine's RSS is **75.3 GiB** (expert arena 46.8 GiB + pinned host KV 15.5 GiB + parked conversations 4.2 GiB), so `MemAvailable` is only **7 799 MiB** → headroom **1 799 MiB** → any snapshot above ~1.8 GiB (≈ 90 K tokens) is refused outright. The floor is a host-level `MemAvailable` sample that the engine's own arenas have already consumed.

### 5.6 Two destructive-ordering defects that turn a refusal into a lost conversation

* `conversations.take(parked.index)` (`:8257`) **removes** the entry before the turn runs. If the re-park at the end of the turn is then refused by either gate, nothing puts it back → the next turn of that conversation finds no entry at all → `0 reused`. This is exactly the 103 K → 108 K sequence in §4.
* `make_room(estimate, held)` (`:6454`) evicts entries and increments `evictions_` **before** the RAM admission check at `:6464`. A park that is subsequently refused has already destroyed the oldest parked conversation. `drop_superseded()` at the top of `park_current_body` is likewise destructive before the refusal points.

## 6. Proposed fixes and their cost

**A. Minimal (what we would like upstream to accept).** At `generate.cpp:7672`, mirror the non-pipelined path when the group size is 1:

```cpp
sl.cached = (GS == 1) && o.prompt_cache > 0 && !sl.img;
```

* Cost: **zero** additional RAM/VRAM (the slot sessions already exist, §5.4); no change to the pipeline's scheduling or arithmetic; the pad-row hazard is unreachable at `GS == 1` by construction (§5.2).
* Benefit on our box: a next turn costs `copy_from_slot` — measured **309–689 ms** for 45 K–103 K tokens — instead of a **68–73 s** full re-prefill.
* Test gap to close: `tools/ba

No site

Links install, modelos, releases.