Pull requests / #1439

#1439 perf(prefill): cache immutable BF16-to-FP16 weight conversions

open · @agorevski · 0 comments · View on GitHub

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

Description

## Title
<!-- If Applicable, reference the GitHub issue -->
Issue: Not linked to an existing issue.

Opt-in persistent cache for immutable BF16-to-FP16 prompt weights

## Summary
<!-- Quick Summary of changes -->
Avoid repeated conversion of explicitly immutable BF16 weights on the active CUDA FP16 BF16_TC path.
This is an independent change against `d5ea7133`; it does not require capacity-aware peer placement or
independent-server CPU scheduling changes. Default behavior is unchanged.

## What changed
<!-- Specifics on files changed, and what changes were made there -->
- `include/strata/prefill/gemm.hpp`, `src/prefill/gemm.cu`: add `bf16_immutable`, explicit
  `invalidate_bf16_cache`, cache statistics/reporting, strict configuration parsing and a bounded persistent
  per-owner conversion arena. Entries are keyed by source pointer and weight shape on the owner's device;
  first-admitted entries remain resident until invalidation/destruction. Full/oversize misses use the
  previous per-call conversion path without evicting retained weights. Allocation and conversion errors
  are reported.
- `include/strata/prefill/prefill.hpp`, `src/prefill/prefill.cpp`: expose persistent reserve pricing and opt
  resident immutable model weights into caching; mutable dequantization scratch, activations and drafter
  weights are not cached.
- `src/program/generate.cpp`: price the persistent arena independently of borrowed prefill buffers in
  single-device skip-if-fits, per-device stage room, automatic/adaptive cache reserves, explicit and
  sized/per-layer native cache caps, post-writeback reserve checks and prefill-owned fit checks.
  The existing `owned_prefill_mib()` helper includes persistent per-owner memory even when buffers are
  borrowed, and includes it alongside estimated/exact owned buffers otherwise. Existing auto/explicit
  callers use that helper without adding the persistent budget again. Other independent callers of the
  helper therefore inherit the same pricing without depending on the new persistent API.
  No peer planner, peer-capacity switches, peer-ranked changes or new peer headers are included.
- `src/prefill/gemm_bf16_parity.cu`, `CMakeLists.txt`: add CPU-only configuration, conversion reuse,
  mutation bypass, shape-key separation, oversize fallback/retention, invalidate/pointer-reuse stress,
  separate stream owners, full prompt-shape parity and inactive-path reservation tests. The existing
  default cuBLAS parity test is preserved.
- `docs/DETAILS.md`: document configuration, source lifetime, invalidation, memory pricing and the measured
  throughput-neutral limitation in its own cache section.

**API, memory and default scope**

Only the exact `STRATA_BF16_TC_CACHE=1` enables caching. `STRATA_BF16_TC_CACHE_MIB` defaults to 256 MiB per
`Gemm`, accepts 1..65536 MiB and disables caching with a diagnostic for invalid values. An arena is
allocated/reserved only for the active FP16 BF16_TC path: CUDA Volta by default, Turing with
`STRATA_BF16_TC=1`; this is not an RTX 8000 name gate. Native BF16, Pascal FP32 and HIP paths reserve no
cache arena. `STRATA_BF16_TC=2` is used only by the tests to force the FP16 path on any CUDA card.

Each `Gemm` belongs to one model lifetime, device and fixed stream, with host-serialized calls. An
immutable source must stay allocated and unchanged until invalidation or destruction. Invalidate before
mutation, unload or pointer reuse; invalidation waits for the stream, drops entries and reuses arena
space without releasing the arena. Cross-stream uploads require explicit ordering. The arena is never
part of a prefill loan. It competes with expert VRAM; a larger budget can reduce expert residency.
Existing BF16_TC rounding is unchanged, not a guarantee of native-BF16 bit identity or cross-library
determinism.

## Extra Notes
<!-- Any extra notes, delete if there are none -->
**Measurements (existing session benchmark logs, not rerun for this branch)**

Physical GPU 3, RTX 8000 at 260 W, IQ3_S, speculation 4, fixed 19,000 expert slots, 8,192-token prefill;
three runs per prompt length, 128 output tokens each. Median prompt throughput:

| Prompt tokens | BF16_TC=1, no cache | 1,024 MiB conversion cache |
| ---: | ---: | ---: |
| 537 | 661 tokens/s | 654 tokens/s |
| 4,057 | 1,262 tokens/s | 1,261 tokens/s |
| 32,057 | 1,416 tokens/s | 1,413 tokens/s |

The cache retained 321 weight conversions and recorded 5,778 hits. Removing repeated conversions is
observable, but whole-model throughput was neutral in this measurement; **no speedup is claimed**.
Evidence: session files `bench-tc1-result.json`, `bench-cache1024-result.json` and
`bench-cache1024-engine.log`. Tests below are isolated-branch validation, not a new whole-model benchmark.

**Validation**

Working directory for all commands:
`/home/algore/.copilot/session-state/a7c2bcbb-4264-4eb2-879d-c71a90285b0b/files/pr-worktrees/bf16-cache`.

1. Configure with system GCC/G++ 13, CUDA 12.4.131 and architecture 75 — **passed**:

   ```sh
   mkdir -p build-scratch
   env CUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=3 TMPDIR="$PWD/build-scratch" \
     cmake -S . -B build \
       -DSTRATA_ENABLE_CUDA=ON -DSTRATA_BUILD_TESTS=ON -DCMAKE_BUILD_TYPE=Release \
       -DCMAKE_C_COMPILER=/usr/bin/gcc-13 -DCMAKE_CXX_COMPILER=/usr/bin/g++-13 \
       -DCMAKE_CUDA_HOST_COMPILER=/usr/bin/g++-13 \
       -DCMAKE_CUDA_COMPILER=/home/algore/miniconda3/bin/nvcc \
       -DCMAKE_CUDA_ARCHITECTURES=75 \
       -DSTRATA_GGML_DIR=/home/algore/GIT/strata/third_party/llama.cpp
   ```

   The first configure attempt preceded creation of `build-scratch` and failed because nvcc could not
   open its scratch output. Creating the local directory and repeating configuration passed.

2. Build only the engine and parity executable, at most two parallel jobs — **passed**:

   ```sh
   env CUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=3 TMPDIR="$PWD/build-scratch" \
     cmake --build build --target strata gemm_bf16_parity -j2
   ```

   Both targets reached 100%; the base emits compiler warnings, but no compilation/link errors occurred.
   Repeating the same build command after the composable `owned_prefill_mib()` pricing change also passed.

3. Cache-focused CTest and the default cuBLAS parity test — **passed, 5/5 tests, 7.81 seconds**:

   ```sh
   env -u STRATA_BF16_TC -u STRATA_BF16_TC_CACHE -u STRATA_BF16_TC_CACHE_MIB \
     CUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=3 \
     flock -x /home/algore/.copilot/session-state/a7c2bcbb-4264-4eb2-879d-c71a90285b0b/files/pr-gpu3.lock \
       ctest --test-dir build --output-on-failure -R '^gemm_bf16_(cache_.*|parity)$'
   ```

   - `gemm_bf16_parity`: passed, default cache-disabled behavior against cuBLAS.
   - `gemm_bf16_cache_config`: passed, CPU-only strict flag and budget parsing.
   - `gemm_bf16_cache_parity`: passed, `BF16_TC=2`, cache enabled, 1 MiB.
   - `gemm_bf16_cache_shapes`: passed, `BF16_TC=2`, cache enabled, 64 MiB; zero-difference cached vs
     uncached FP16 conversion path across prompt shapes.
   - `gemm_bf16_cache_inactive_reserve`: passed, `BF16_TC=0`, cache enabled, 65536 MiB requested; inactive
     conversion path allocated/reserved zero arena bytes.

   CTest supplied each cache test's environment from its CMake properties. All GPU execution was
   restricted to physical GPU 3 and held the exclusive lock for the complete CTest lifetime.

4. `git diff --check` — **passed**.

Validation logs are outside the worktree in session files:
`pr-bf16-cache-configure.log`, `pr-bf16-cache-build.log`, `pr-bf16-cache-final-build.log`,
`pr-bf16-cache-ctest.log`. Local nvcc scratch files were cleaned up after validation.

Cross-PR composition was also checked:

```sh
git merge-tree --write-tree perf/immutable-bf16-conversion-cache perf/capacity-aware-peer-placement
git merge-tree --write-tree perf/immutable-bf16-conversion-cache fix/independent-server-cpu-affinity
```

Both passed without conflicts. The peer planner's primary reserve calls the existing `owned_prefill_mib()`
helper, so it inherits the persistent arena price if both independent PRs are merged; it does not require
a new peer-specific cache API or duplicate reserve.

No server was started or power setting changed during isolated-branch validation.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.