Pull requests / #1486
#1486 feat: introduce Dockerfile.rocm-stable
open · @dambaev · 0 commentaires · Sur GitHub
BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
Description
## Title
<!-- If Applicable, reference the GitHub issue -->
Issue: Resolves #
## Summary
Introduces `Dockerfile.rocm-stable` based on `Dockerfile` and `setup.py`. I would prefer to use `setup.py` during the build phase, but I see no way to call `setup.py` without hardware passthrough and `docker build` does not support `--device=` option. And the existing cuda-targeted version of `Dockerfile` is using internal `setup.py` interface.
It maybe worth including of `BACKEND` variable into the scope, but `setup.py` is capable of probing during `run` phase itself...
## What changed
Allows ROCm users to build/run docker container for their hardware
## Extra Notes
runtime logs of `status` prompt on the `Strata` repo right after container boot on 2 hosts: mini pc + 9070xt oculink eGPU and asus g513qy:
<details>
<summary>1. gfx1201 9070xt, 64GB DDR5, qwen3.8 flash next coder, ctx 131k, k q8, prompt context ~70k:</summary>
```
$ docker run --rm --device=/dev/kfd --device=/dev/dri --group-add video -p 8080:8080 --ulimit memlock=-1 -v /mnt/Strata-data:/data -e MODEL=IQ1_M -e FAMILY=coder -e CONTEXT=$((128 * 1024)) -e REINSTALL=1 strata
Setting up coder-iq1_m: downloading the model (~70 GB; the engine is already in the image).
Strata - Qwen3.8-Flash-Next on a normal PC (a GPU + system RAM + CPU)
=== Step 1: checking your PC ===
Your AMD GPUs:
GPU 0: AMD Radeon RX 9070 series / AI PRO R9700 (gfx1201), 16 GB VRAM - can be used
GPU 1: AMD Radeon 780M / 760M / 740M (Ryzen 7040 / 8040, Phoenix / Hawk Point, gfx1103), 4 GB VRAM - not supported - Strata's AMD backend runs on the RX 7900 XT / XTX (gfx1100), RX 7800 XT / 7700 XT (gfx1101), RX 9060 XT (gfx1200) and RX 9070 / 9070 XT / Radeon AI PRO R9700 (gfx1201), and the RX 6800 / 6900 series (gfx1030) and RX 6700 XT (gfx1031, #524), and the RX 7600 / 7600 XT (gfx1102, one run reported, #942), all unvalidated, and the Ryzen AI Max "Strix Halo" APU (Radeon 8060S / 8050S / 8040S, gfx1151: experimental, docs/STRIX_HALO.md) only, this is gfx1103 (an integrated Radeon, not Strix Halo)
[ok] GPU: AMD Radeon RX 9070 series / AI PRO R9700 (gfx1201), 15.9 GB VRAM, gfx1201 (AMD: docs/AMD_HIP.md)
[ok] RAM: 59 GB
[ok] CPU: AMD Ryzen 7 8845HS w/ Radeon 780M Graphics (AVX-512)
=== Step 2: your choices ===
[ok] model: Qwen3.8-Flash-Next Coder
1) IQ1_M the Coder's only size: half the experts, stored like IQ3_S (3.5 bits); download 58 GB, uses ~23 GB of RAM
[ok] size: IQ1_M
[ok] context: 131072 tokens
[ok] KV cache: 8-bit
[ok] images: off
EXPERIMENTAL - speed projection: a small control vector applied while the model runs (layers 4-44).
It changes how the model answers: its package describes it as a refusal-direction projection (the
model declines far fewer requests). Off unless you choose it; when on, the web app can switch it off
per chat. Details: data/experimental-speed-projection/README.md
[ok] experimental speed projection: off
=== Step 3: Python packages ===
Installing numpy, jinja2, regex, pyyaml, tqdm, requests, cmake, ninja, pillow, psutil ...
> /opt/strata/.venv/bin/python -m pip install --quiet --disable-pip-version-check numpy==2.5.3; python_version >= "3.12" numpy==2.4.6; python_version == "3.11" numpy==2.2.6; python_version < "3.11" jinja2==3.1.6 regex==2026.9.10 pyyaml==6.0.3 tqdm==4.70.1 requests==2.34.2 cmake==4.4.3 ninja==1.13.2 pillow==12.3.0 psutil==7.2.2 markupsafe==3.0.3 certifi==2026.7.22 charset-normalizer==3.5.1 idna==3.20 urllib3==2.8.0 colorama==0.4.6; sys_platform == "win32"
[ok] numpy, jinja2, regex, pyyaml, tqdm, requests, cmake, ninja, pillow, psutil installed
=== Step 4: the Strata engine ===
[ok] llama.cpp 3cf0325 (gguf-py, ggml, mtmd)
[ok] engine already built for this PC
[ok] engine: /opt/strata/engine/strata
=== Step 5: downloading Qwen3.8-Flash-Next Coder IQ1_M ===
[ok] Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002.gguf already downloaded
[ok] Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00002-of-00002.gguf already downloaded
[ok] model files present
=== Step 6: preparing the model for Strata ===
[ok] model prepared: /data/packs/coder-iq1_m
[ok] MTP draft layer: /data/mtp/rt
=== Step 7: writing the start script ===
[ok] KV streaming on: the context's KV cache lives in RAM (1.8 GB), more experts fit in VRAM
[!] no hipBLASLt tuning table for gfx1201 with hipBLASLt 100200 (have: gfx1201-hipblaslt-100202.txt, gfx1201-hipblaslt-100500.txt): the prompt's dense matrix products use plain hipBLAS (tools/hip/tune_hipblaslt makes a table: docs/AMD_HIP.md, Tuning table)
tip: for several agents or clients at once, --conversation-cache-mib 8192 in the config's args keeps each one's conversation; without it they were measured re-reading ~90% of their prompts (#882 #440; docs/DETAILS.md, Multiple conversations)
[ok] start script: run-coder-iq1_m.sh
All set.
API (OpenAI): http://127.0.0.1:8080/v1 (any API key; model name: anything)
API (Anthropic): http://127.0.0.1:8080/v1/messages
Other devices: the server window prints this PC's address (http://<IP>:8080/) - no API key set: anyone on your network can use it
Next time: just run ./setup.sh (or run-coder-iq1_m.sh) - it starts right away
Several requests at once: left at one at a time - parallel N reduces waiting for several users but costs about 10-25% speed per request on this card (docs/BATCHING.md).
Config: /data/config/strata-coder-iq1_m.json (from MODEL IQ1_M, just set up)
Strata - Qwen3.8-Flash-Next on a normal PC (a GPU + system RAM + CPU)
[ok] GPU: AMD Radeon RX 9070 series / AI PRO R9700 (gfx1201) (16 GB, AMD)
----------------------------------------------------------------------------------------------------
Starting qwen3.8-flash-next-coder-iq1_m: it loads about 30 GB into RAM and locks part of it for the GPU.
While it does, YOUR PC CAN BE SLOW OR STOP RESPONDING FOR 1-3 MINUTES (longer the first time after a
restart). That is normal: please wait and don't close this window - the browser opens when it is ready.
Later, closing this window stops the model.
----------------------------------------------------------------------------------------------------
Settings (strata-coder-iq1_m.json): --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5
--max-context 131072 --kv int8 --kv-resident 32768; server 0.0.0.0:8080, gpu 0
loading the model (the first start takes a minute or two) ...
[strata] starting the engine: reading the model's weights ...
[strata] loading the experts into RAM (about 30 GB) and locking part of them for the GPU.
YOUR PC CAN BE SLOW OR STOP RESPONDING FOR 1-3 MINUTES NOW - this is normal.
Please wait and don't close this window; the browser opens when it is ready.
[strata] experts loaded: 23.42 GiB at 1.58 GiB/s (23 s so far)
[strata] filling the GPU's expert cache (4844 experts, 9.21 GiB of VRAM) ...
[strata] almost ready ...
ready: http://127.0.0.1:8080/v1 (OpenAI: /v1/chat/completions, Anthropic: /v1/messages, context 131072 tokens)
open http://127.0.0.1:8080/ in a browser to chat; close this window to stop the model
from other devices: http://172.17.0.2:8080/ (API: http://172.17.0.2:8080/v1)
WARNING: no API key - anyone on your network can use this model. Add "api_key": "..." to the config (clients send it as their API key; the web page asks for it)
[strata] reading the prompt: 8,192 of 77,942 tokens, 7 s so far
[strata] reading the prompt: 16,036 of 77,942 tokens, 12 s so far
[strata] reading the prompt: 24,228 of 77,942 tokens, 19 s so far
[strata] reading the prompt: 32,420 of 77,942 tokens, 24 s so far
[strata] reading the prompt: 40,612 of 77,942 tokens, 30 s so far
[strata] reading the prompt: 48,804 of 77,942 tokens, 36 s so far
[strata] reading the prompt: 56,996 of 77,942 tokens, 42 s so far
[strata] reading the prompt: 65,188 of 77,942 tokens, 50 s so far
[strata] reading the prompt: 73,380 of 77,942 tokens, 57 s so far
[strata] reading the prompt: 77,937 of 77,942 tokens, 62 s so far
[strata] thinking: 8 of max 53122 tokens, 41.7 tok/s, 63 s
[strata] thinking: 54 of max 53122 tokens, 45.0 tok/s, 64 s
[strata] thinking: 108 of max 53122 tokens, 47.6 tok/s, 65 s
[strata] thinking: 159 of max 53122 tokens, 48.4 tok/s, 66 s
[strata] thinking: 208 of max 53122 tokens, 47.8 tok/s, 67 s
[strata] writing a tool call: terminal: 260 of max 53122 tokens, 47.8 tok/s, 68 s
[strata] writing a tool call: terminal: 310 of max 53122 tokens, 48.0 tok/s, 69 s
[strata] done: 329 tokens in 70 s (48.8 tok/s) (stop, cancel=False), expert cache 76.7% hit (+2.4% of the routed experts over PCIe)
[strata] reading the prompt: 78,779 of 78,784 tokens, 4 s so far
[strata] thinking: 29 of max 52280 tokens, 46.5 tok/s, 5 s
[strata] answering: 78 of max 52280 tokens, 47.4 tok/s, 6 s
[strata] answering: 134 of max 52280 tokens, 50.2 tok/s, 7 s
[strata] answering: 187 of max 52280 tokens, 50.9 tok/s, 8 s
[strata] answering: 245 of max 52280 tokens, 52.1 tok/s, 9 s
[strata] answering: 298 of max 52280 tokens, 52.0 tok/s, 10 s
[strata] done: 310 tokens in 10 s (51.5 tok/s) (stop, cancel=False), expert cache 80.3% hit (+1.9% of the routed experts over PCIe)
```
</details>
<details>
<summary>2 gfx1031 6800m, 64GB DDR4, qwen3.8 flash next coder, ctx 131k, k q8, ~70k prompt context:</summary>
```
$ docker run --rm --device=/dev/kfd --device=/dev/dri --group-add video -p 8080:8080 --ulimit memlock=-1 -v ~/.cache/strata-data:/data -e MODEL=IQ1_M -e FAMILY=coder -e CONTEXT=131072 -e REINSTALL=1 strata
Setting up coder-iq1_m: downloading the model (~70 GB; the engine is already in the image).
Strata - Qwen3.8-Flash-Next on a normal PC (a GPU + system RAM + CPU)
=== Step 1: checking your PC ===
Your AMD GPUs:
GPU 0: AMD Radeon RX 6700 XT series (gfx1031), 12 GB VRAM - can be used
GPU 1: AMD Radeon (gfx90c), 0 GB VRAM - not supported - Strata's AMD backend runs on the RX 7900 XT / XTX (gfx1100), RX 7800 XT / 7700 XT (gfx1101), RX 9060 XT (gfx1200) and RX 9070 / 9070 XT / Radeon AI PRO R9700 (gfx1201), and the RX 6800 / 6900 series (gfx1030) and RX 6700 XT (gfx1031, #524), and the RX 7600 / 7600 XT (gfx1102, one run reported, #942), all unvalidated, and the Ryzen AI Max "Strix Halo" APU (Radeon 8060S / 8050S / 8040S, gfx1151: experimental, docs/STRIX_HALO.md) only, this is gfx90c
[ok] GPU: AMD Radeon RX 6700 XT series (gfx1031), 12.0 GB VRAM, gfx1031 (AMD: docs/AMD_HIP.md)
[ok] RAM: 62 GB
[ok] CPU: AMD Ryzen 9 5900HX with Radeon Graphics (AVX2)
=== Step 2: your choices ===
[ok] model: Qwen3.8-Flash-Next Coder
1) IQ1_M the Coder's only size: half the experts, stored like IQ3_S (3.5 bits); download 58 GB, uses ~23 GB of RAM
[ok] size: IQ1_M
[ok] context: 131072 tokens
[ok] KV cache: 8-bit
[ok] images: off
EXPERIMENTAL - speed projection: a small control vector applied while the model runs (layers 4-44).
It changes how the model answers: its package describes it as a refusal-direction projection (the
model declines far fewer requests). Off unless you choose it; when on, the web app can switch it off
per chat. Details: data/experimental-speed-projection/README.md
[ok] experimental speed projection: off
=== Step 3: Python packages ===
Installing numpy, jinja2, regex, pyyaml, tqdm, requests, cmake, ninja, pillow, psutil ...
> /opt/strata/.venv/bin/python -m pip install --quiet --disable-pip-version-check numpy==2.5.3; python_version >= "3.12" numpy==2.4.6; python_version == "3.11" numpy==2.2.6; python_version < "3.11" jinja2==3.1.6 regex==2026.9.10 pyyaml==6.0.3 tqdm==4.70.1 requests==2.34.2 cmake==4.4.3 ninja==1.13.2 pillow==12.3.0 psutil==7.2.2 markupsafe==3.0.3 certifi==2026.7.22 charset-normalizer==3.5.1 idna==3.20 urllib3==2.8.0 colorama==0.4.6; sys_platform == "win32"
[ok] numpy, jinja2, regex, pyyaml, tSur le site
Liens install, modèles, releases.