Pull requests / #789

#789 prefill: overlap shared-expert work with routed expert uploads

closed · @InB4DevOps · 0 Kommentare · Auf GitHub

BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

Beschreibung

### What

The routed-only prefill path delays expert uploads in two places:

- The CPU waits for the shared expert to finish before reading back routing IDs, although only the router needs to have finished.
- The eight-entry lookahead counts resident experts alongside streamed experts. Resident experts occupy no transfer slot, but still reduce how many uploads are queued ahead.

This changes the schedule in `src/prefill/prefill.cpp`:

1. Read back routing IDs first, then queue shared-expert computation so it can overlap CPU grouping and subsequent expert uploads.
2. Count actual streamed experts against the eight-slot lookahead.
3. Keep the existing release-event ordering before reusing a staging slot.

The ring size, buffer allocation, expert order, MMQ groups and arithmetic are unchanged. The existing timing attribution is retained.

`STRATA_PREFILL_STREAM_AHEAD=0` restores the previous schedule; the default is enabled.

**Scope:** routed-only batched prefill, normally chunks below 1,024 tokens, including short final chunks. The large-chunk full-stream schedule and output-token decoding are unchanged.

### Measured

RTX 3060 12 GB, Intel i7-12700KF, Linux, Coder IQ1_M, INT8 KV, fixed 4,160-token capacity, resident expert arena and automatic expert cache.

A/B used the same binary, with the environment switch selecting the schedule. Profiling was disabled. Each run started a fresh engine; the reported metric is batched prefill time, excluding model loading, cache refill and final-input-token decoding.

The matrix covered:

- Prompt lengths: **256, 529, 1,024, 2,048, 4,096** freshly prefilled tokens.
- Requested chunks: **128, 256, 512, 768, 1,024**.
- Synthetic ramp token IDs 100–299; one generated token.
- One discarded warmup pair and three measured A/B pairs per case, with alternating arm order and seeded case shuffling.

**25 cases, 200 retained runs:** 50 warmup and 150 measured. One initial launcher interruption was handled by rerunning its incomplete measurement pair in full; that excluded attempt is preserved locally.

| Path exercised | Cases | Median throughput gain |
|---|---:|---:|
| Routed-only | 22 | **+0.82% to +2.83%** |
| Pure 1,024-token full-stream chunks | 3 | **−0.18% to +0.20%** |

A requested chunk of 1,024 still exercises routed-only processing when the entire prompt is shorter than 1,024 tokens.

**Measurement provenance:** this matrix ran on upstream `6f32ec0` plus these scheduling changes, before the profiling and benchmark tooling was removed to produce this minimal patch. Profiling was disabled during measurement. The full matrix has not been repeated on the exact final minimal binary; that binary received the correctness checks below.

### Correctness and build checks

The final minimal patch builds with CUDA for `sm_86`.

Three fresh A/B correctness cases—six engine runs—passed:

- 529 tokens with chunk 256.
- 1,041 tokens with chunk 1,024: a full-stream chunk followed by a 17-token routed-only tail.
- 65 tokens with mmap experts and a two-slot host stager, exercising unpinned transfers and buffer reuse.

Generated tokens, GDN state hashes and sampled FP32 residual bytes matched in each pair.

The broader matrix also passed output/state, actual chunk-count and expert-placement checks. Benchmark tooling and generated results are kept outside this PR; its diff is **one file, 34 insertions and 8 deletions**.

### Limits

This is a modest scheduling improvement, not an increase in PCIe bandwidth. Larger chunks can avoid much more repeated expert traffic than this change saves.

Measurements cover one NVIDIA/Linux machine and synthetic prompts. Three pairs per case provide limited statistical certainty. Residual checks sample every 64th position rather than comparing every tensor. AMD, Windows and multi-GPU execution remain unvalidated.

Developed with an AI coding assistant; builds, tests and measurements ran locally.

Mehr auf der Site

Links zu Install, Modellen, Releases.