Pull requests / #1017
#1017 feature: RDNA3 support / Cache-aware routing
closed · @dhoard · 0 comentarios · En GitHub
BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationLinux
Descripción
# RDNA3 support
- RDNA3 (gfx1101) support
- Swift 1.5 launcher
- Cache-aware routing
- Docker build/run
**Base:** `main` (`6f32ec0`) · **Head:** `feature/rdna3-support` (`9dbc79e`)
**Net diff:** 39 files added, 25 modified, 172 deleted (all under `bench/results/`)
Everything below is what the branch *is*, read from the tree at `9dbc79e`, not from individual commit
messages: several commits in the history add work that a later commit in the same branch removes, and those
cancellations are not part of this PR.
---
## 1. Summary
The branch turns Strata into something that runs on one AMD RDNA3 card from the repository root, and then
measures and tunes that deployment:
* a HIP/Docker build-and-run path for **gfx1101 (RX 7700 XT)** under an enforced **10 GiB VRAM contract**:
`build.sh` (builder + runtime images), `docker/` (entrypoint, model bootstrap, device probe, VRAM guard),
* **`run.sh`**, one launcher for the container (the branch's own `run2.sh`/`run3.sh` are gone, so this is
the only one): a **release axis** (`qwen|swift|coder`), per-release default quant and tuning,
release-qualified pack directories with a provenance guard, and a shipped default of **Swift 1.5 IQ3_XXS**
(measured on the reference card),
* **cache-aware routing (CAR)** in the engine: a decode window's uncached expert is replaced by the best
*resident* expert of that layer when the router's own score ratio clears a threshold. **On by default at
0.35**, with `--car-threshold 1.0` restoring byte-identical stock tokens,
* a **parallel intermediate activation quantization** phase in the CPU expert pool (PR #500),
* **gfx1101 prompt-GEMM calibration** (a 58-row hipBLASLt table plus the tooling and the record),
* two serving fixes: an unbounded/incorrect streaming tool-call parser, and a deployment-wide default
reasoning effort.
## 2. Docker / HIP build and run path for gfx1101 (new)
| File | What it does |
| --- | --- |
| `build.sh` | Builds `strata-hip-builder:<arch>` (ROCm toolchain) → compiles the engine into `./build-hip` (ccache volume) → builds `strata-hip:<arch>` (+ `-<gitsha>`, `-latest`). `--tests`, `--arch`, `--runtime-only`. |
| `docker/Dockerfile.hip-builder` | ROCm 7.2 builder image. |
| `docker/Dockerfile.hip` | Runtime image (`rocm/dev-ubuntu-24.04:7.2.1-complete`): engine, `serve/`, `tools/`, model bootstrap. The compiled binary arrives through the `engine` **named build context**, so the build context stays source-only. |
| `docker/build-engine.sh` | The in-container CMake configure/build (Ninja, MMQ prefill, ccache). |
| `docker/entrypoint-hip.sh` | Guard → device → VRAM budget → model/pack → engine config → server. All tuning by environment; the config it writes is the one the server reads. |
| `docker/bootstrap-model.sh` | Check-and-fetch the quant and the pack; `check` mode never writes. Free-space gate before any download. |
| `docker/hipinfo.py` | GPU discovery from the KFD topology (no ROCm Python bindings), with `--arch`, `--vram`, `--render-node`, `--reserve-mib`. |
| `docker/vram-guard.py` | Samples device memory against the 10 GiB budget, with `--pid` attribution to the container's engine. |
| `docker/README.md` | The image/run story and the measured facts (engine build, GPU suite, VRAM peak, first-start download). |
| `.dockerignore` | Rewritten for the HIP context: build trees, GGUFs, packs, harness state and captures excluded. |
The 10 GiB ceiling is enforced in three places in the current code, not one:
1. `src/core/hip_budget.cpp` (new) — a Linux HIP interposer over the allocation APIs, enabled from
`cmake/hip_backend.cmake` (`--export-dynamic`, `dl`), with `STRATA_VRAM_BUDGET_MIB` (default 10240),
`STRATA_VRAM_RUNTIME_RESERVE_MIB`, `STRATA_VRAM_SLACK_MIB`,
2. `include/strata/platform/device_budget.hpp` (new) — `DeviceBudget`, the serializable reservation
accounting used by callers,
3. `docker/vram-guard.py` + `run.sh`'s budget checks (`1281..10240 MiB`).
Tests: `docker/test_runtime_contract.py`, `tests/hip/device_budget.cpp` (ctest `hip_device_budget`, run with
`STRATA_VRAM_BUDGET_MIB=10240`), `src/platform/device_budget_test.cpp` (ctest `device_budget_test`).
## 3. `run.sh` — one launcher, a release axis, Swift 1.5 by default (new file)
`main` has no `run.sh`; the branch adds it and drops the branch-internal `run2.sh`/`run3.sh`.
* **Release axis:** `--release qwen|swift|coder` (or `$STRATA_RELEASE`, or implied by `$STRATA_HF_REPO`,
or by the quant). Resolution is asked of `docker/hfmodel.py --print release`, so the launcher and the
container cannot disagree, and a refusal from the resolver is the launcher's error message.
* **Per-release default quant:** swift `IQ3_XXS`, qwen `IQ3_S`, coder `IQ1_M`.
* **Tuning keyed on release+quant** (`case "$RELEASE:$MODEL"`):
| release:quant | expert cache | later allowance | prefill ring |
| --- | --- | --- | --- |
| `qwen:IQ3_XXS` | 800 (the 0.1.39 pin, unchanged) | 768 | none |
| `qwen:IQ3_S`, `swift:IQ3_XXS` | `auto` | 700 | 48 |
| anything else | 680 (conservative) | 768 | none, with a warning that the combination is unmeasured |
* **Refuses** before Docker sees data: wrong card (only `gfx1100`/`gfx1101`), `HSA_OVERRIDE_GFX_VERSION`
set, a budget outside `1281..10240`, any `--max-context` other than 131072, a filesystem too small for
the quant plus its pack plus a 20 GB floor, and an image too old to carry the allocation guard (checked
through its `io.strata.vram-budget` label).
* **$STRATA_HF_REPO now reaches the container** (it did not before), and the release-qualified
`STRATA_PACK_DIR` / `MODEL_NAME` / `HF_REPO` are forwarded for non-qwen releases while the qwen line
stays byte-identical to the pre-switch fixture.
* `docker/entrypoint-hip.sh` gives the same axis on the container side: the pack directory takes the
release tag (`packs/swift-iq3_xxs` beside `packs/iq3_xxs`), the model id comes from the release name
(`swift-1.5-iq3_xxs`), the release license is logged at every start, and the pre-release-axis defaults
are kept whenever the resolver abstains.
### Model resolution (`docker/hfmodel.py`, +381)
`FAMILIES`/`MODELS` mirror `setup.py`'s table for the three releases. Two behaviour fixes are in the
current code:
* **A shard must belong to the quant that was asked for.** The old name-blind catch-all
(`*0000{i}-of-00002.gguf`) let a request for a Swift quant that is absent in the cache resolve a
sibling quant's shard - the launcher would then report "cached" and advertise one quantization while
loading another. Offline now fails with the right fetch command.
* **Per-release PLE file.** Swift packs `per_layer_token_embd.weight` into file 1, the Qwen/Coder repos
into file 2; the resolver reports which, and the entrypoint takes it as `STRATA_PLE_FILE`.
Pinned by `docker/test_hfmodel_release.py`, `docker/test_hfmodel_ple.py`, and (for packs)
`docker/test_pack_provenance.py`: `experts.bin.src.json` shard names must belong to the selected release or
the container refuses to start (the quant vocabulary is shared, so the wrong pack is silent corruption);
a pack predating that file warns by name only.
## 4. Cache-aware routing (engine, default on)
**New:** `include/strata/core/car.hpp`, `src/core/car.cpp`, `tests/core/car_substitute_test.cpp`,
`tools/car_estimate.py`, `tools/test_car_estimate.py`.
**Changed:** `verify.hpp`/`verify.cpp` (score publication + the per-group host hook), `expert_source.hpp`/
`expert_source.cpp` (`car_apply`, `car_report`, backfill into `usage`), `elementwise.hpp`/
`elementwise.cu` (optional weights/router-score publication in `doorbell_publish_res`), `generate.cpp`
(flags, `STRATA_CAR_*`, refusals, reports).
The rule: a decode window's uncached pick `E` is replaced by the best **resident** expert `C` that the token
did not already select, when `p_C / p_E` clears the threshold. The substituted entry keeps its router weight
and the GPU computes the resident expert from a slot already in VRAM. The decision is a pure function with
its own translation unit and its own test (it builds and runs in a CPU-only configuration); the
implementation allocates nothing per token and is deterministic (best ratio first, ties by lower id).
* **Default `--car-threshold 0.35`.** `1.0` is OFF and restores the stock model's tokens byte for byte
(verified against the build that predates the feature). `--car-threshold 1.0` is the fidelity escape
hatch; `docs/DETAILS.md` states plainly that the default changes answers.
* Knobs: `--car-threshold`, `--car-warmup`, `--car-budget`, `--car-free-ratio` and `STRATA_CAR_*`
(a flag always wins over the environment). `--car-dampen 1` is **refused** rather than ignored.
* Paths it is not wired into are handled explicitly: with the feature asked for, they are an error; coming
from the default, they print "cache-aware routing is OFF for this run" and downgrade to 1.0 (no pool,
`--spec 0`, `--batch`/`--slots`, `--layer-split`, `--peer-device`, `--remote-expert-opt`,
`STRATA_VERIFY_DEVICE_PLAN=1`).
* `STRATA_DUMP_ROUTER_SCORES=<path>` writes a `STRCS1` trace (ids + the whole score row per token, before
any substitution); `tools/car_estimate.py` replays a trace offline against the same rules and can check
its copy of the rule against the engine's own golden case.
**Measured** (RX 7700 XT, gfx1101, 10 GiB share, Swift 1.5 IQ3_XXS, `./run.sh` with no flags, 4K-8K prompt,
8 trials per arm, as recorded in `docs/DETAILS.md`): decode **30.5 → 35.3 tokens/s (+15.5%)**, the window
86.0 → 75.3 ms, the CPU pool's share **-28%** (20.4 → 14.7 ms per layer-window) at **30.8% of decode misses
substituted** (mean accepted ratio 0.598); the VRAM hit kernel grows ~3x, which is where the work moved. Peak
card share 9,931 MiB off / 10,155 MiB at `--car-threshold 1.0`, both under the 10,240 MiB contract.
`tools/car_estimate.py` (offline): 30.6% of misses at 0.35 (2.7/token), 42.9% at 0.25, 18.3% at 0.5, 8.3% at
0.7; 94.5% of substitutions are ranks 5-9 of the ten.
**Limits, as documented in `docs/DETAILS.md`:** perplexity/KL against CAR off is *not* measured, and neither
are the other shipped lines (Qwen IQ3_S, Coder). The quality evidence is the coding smoke (205/205),
needles 9/9 at 1k/4k/32k, and the bounded rule above.
## 5. gfx1101 prompt-GEMM calibration
* `tools/hip/gfx1101-hipblaslt-100202.txt` — a **58-row** hipBLASLt table (RX 7700 XT, Swift 1.5 IQ3_XXS,
ROCm 7.2.1-complete, including dense T=128/512/2048). `docs/AMD_HIP.md` records it and stops claiming
gfx1101 has no table; `docs/AMD_HIP_GFX1101_TUNING.md` is the calibration/validation record.
* `tools/hip/tune_hipblaslt.cpp` gains `--warmups/--repetitions/--iterations/--max-algos` (defaults
unchanged in effect) and reports them.
* `tests/hip/prefill_hipblaslt_gemm.cpp` gains `--all`, which validates **every** row at its calibrated T
with `beta=0` and `beta=1` + output offset, plus a repeated-call descriptor-cache timing; the default
smoke keeps its four pinned cases and now prints how many rows it covered.
* `src/prefill/gemm.cu` bounds the hipBLASLt metadata (cache and fallback-shape set) at 512 entries so
arbitrary tail shapes cannot grow server metadata without limit.
## 6. CPU expert pool: the intermediate quantization as a parallel phase
`pool.hpp`/`pool.cpp` add drain mode 7: the per-(expert, token) intermediate quantization of a multi-token
layer runs across the pool and the host instead of as a single-threaded nested loop (the phase barrier is
still required; the work no longer sits on one thread). `prepare_quant_tasks` keeps a running count per
expert, so no per-task allocation happens. `STRATA_POOL_PARALLEL_QUANT=0` restores the sequential loop (the
A/B arm) and `STRATA_POOL_QUANT_MIN_TASKS` (default 8, the measured crossing) keeps it on the host below the
size whEn el sitio
Enlaces a install, modelos, releases.