Issues / #1069
#1069 Tesla P100 (Pascal, sm_60) + RTX 2070 SUPER, Linux/Docker (CUDA 12.9), 32 GB RAM: Q2_0 low-RAM mode, about 30 tok/s decode on the P100 alone and 36-38 tok/s with the second card as a helper expert cache (benchmark.py)
closed · @tobby1968 · 1 Kommentare · Auf GitHub
BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationLinux
Beschreibung
## Summary
Strata 0.1.39 runs Qwen3.8-Flash-Next **Q2_0** on a **Tesla P100 16 GB** (compute capability 6.0, the CUDA 12 engine) on a PC with **32 GB of RAM**, through the low-RAM mode (`--mmap-experts`). Measured with your `benchmark.py` from the RTX 5090 community report (fresh prompts, no reused tokens, 256 greedy output tokens, three runs each; median [min-max]):
| Config | Prompt tokens | Prompt tok/s | Decode tok/s | TTFT (s) | Total (s) |
| --- | --- | --- | --- | --- | --- |
| P100 only | 4,096 | 284.3 [194.9-291.4] | **30.5** [24.8-31.2] | 14.5 [14.2-21.1] | 22.6 [22.5-31.4] |
| P100 only | 32,768 | 286.9 [278.4-330.5] | **30.6** [29.0-32.1] | 114.4 [99.4-117.8] | 123.2 [107.3-126.1] |
| P100 + RTX 2070 SUPER as a **helper expert cache** | 4,096 | 296.4 [202.0-299.2] | **35.8** [31.6-37.6] | 14.0 [13.7-20.4] | 21.1 [20.5-28.4] |
| P100 + RTX 2070 SUPER as a **helper expert cache** | 32,768 | 295.4 [284.8-334.6] | **38.4** [38.0-38.8] | 111.1 [98.0-115.2] | 117.7 [104.8-121.7] |
(0 reused tokens in all runs; the first 4K run of each set was slowest while the file cache was still filling.)
The docs list the P100 as "not measured" and say no Pascal card was available. This is a measurement on one, and a two-card PC (Pascal + Turing, 32 GB RAM). One machine, one model size.
## Hardware
- GPUs: Tesla P100-PCIE-16GB (sm_60, PCIe 3.0 x16, 250 W) and GeForce RTX 2070 SUPER 8 GB (sm_75, chipset x4 slot, 215 W). Strata's PCIe probe: 12.5 GB/s host->device for the P100, 3.1 GB/s for the 2070 SUPER.
- CPU: Intel Core i7-7700 (4 cores / 8 threads, no AVX-512, so the expert kernels run on AVX2).
- RAM: 32 GB (31 GiB usable), DDR4. Speed not recorded.
- Storage: WD WDS500G2B0C NVMe 500 GB (DRAM-less, PCIe 3.0 x4).
- Other workloads: during the benchmarks all other containers on the machine were stopped, so about 26 GiB of file cache was available to the experts. A database container with no load stayed up.
## Software
- Ubuntu 26.04.1 LTS (kernel 7.0.0-34, **glibc 2.43**), NVIDIA driver 580.178.04, Docker 29.8.1, NVIDIA Container Toolkit 1.20.1.
- Strata v0.1.39 (commit 6f32ec070f23ced9f50e704d854d775da52591ab), engine 0.1.39, llama.cpp 3cf0325 (as pinned).
- Built in Docker with CUDA 12.9.1. The shipped `Dockerfile` is CUDA 13 only (it refuses architectures below 75), so I made `Dockerfile.cuda12` from it (diff below). A host build with CUDA 12.9 is not possible here: glibc 2.43 clashes with the toolkit's math headers (as TROUBLESHOOTING.md says), so building inside the Ubuntu 24.04 image was the way.
## Model
ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF, `Q2_0` (shards 37.6 GB + 28.8 GB, SHA-256 checked against the repository's LFS hashes). Placed in `/data/models/Q2_0/`. Setup built the pack (`experts.bin` 34 GB, `dense.bin`) and fetched/packed the MTP layer. No vision.
## Settings
`docker run --gpus all --ulimit memlock=-1 -p 127.0.0.1:8099:8099 -v ~/strata-data:/data -e PORT=8099 -e FAMILY=qwen -e MODEL=Q2_0 -e CONTEXT=40000 -e VISION=no -e GPU=0 -e LOW_RAM=on -e STRATA_CUDA=12 -e STRATA_EXPERIMENTAL_SM60=1 -e REINSTALL=1 strata-cuda12:v0.1.39` (no API key, loopback only, so that `benchmark.py` can call it unchanged).
Resulting config (`bench-gpu0/strata-q2_0.json`): `--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --max-context 40000 --kv int8 --mmap-experts`. Expert cache on the P100: 7,719 slots (9.94 GiB).
**Helper-card run** (`bench-gpu0-helper/strata-q2_0.json`): the same config, `"gpu": 0` unchanged, plus `"env": {"CUDA_VISIBLE_DEVICES": "0,1", "CUDA_DEVICE_ORDER": "PCI_BUS_ID"}` and the arguments `--expert-cache-device1 4500 --remote-expert-opt`. The log then says `CUDA1: 4500 additional experts, 5.79 GiB; results return through pinned host rows`. The decode expert-cache hit rate was about 98% (the P100-only run: about 62-67% on the shorter ad-hoc requests), and CUDA1 computed about 50,000 expert entries per 256-token answer.
`benchmark.py` was run inside the container with `--root /opt/strata --pack /data/packs/q2_0 --url http://127.0.0.1:8099 --targets 4096,32768 --runs 3`.
## Differences from the community protocol
- Context limit 40,000 (not 131,072), so **128,000 was not run**; the `needle_bench.py` recall checks were not run either.
- One warm-up request excluded, three runs per length, serial, increasing length, same loaded engine, no restart; loading time excluded.
- Per-run JSON: `bench-gpu0/` and `bench-gpu0-helper/` (`results.json`, `summary.json`, raw timings `tokens-*-raw.json`, `engine.log`).
## Extra ad-hoc runs (single requests, rough)
Streaming `/v1/chat/completions`, thinking off, token counts from `usage`; decode = completion tokens / (total - TTFT).
| Run | Prompt / output tokens | P100 only, other services stopped | P100 only, other services up (idle) | P100 + 2070 SUPER (layer split), services stopped |
| --- | --- | --- | --- | --- |
| short | 107 / 30 | 22.6 tok/s | 11.4 tok/s (first call after start) | 16.2 tok/s |
| long prompt | 1,742 / 19-26 | 19.8 tok/s; prompt 112 tok/s | 16.7 tok/s; 104 tok/s | 16.4 tok/s; 106 tok/s |
| long generation | 34 / 588-600 | 27.9 tok/s | 29.2 tok/s | 32.8 tok/s |
With the layer split the 2070 SUPER's 8 GB was almost full (7.4 GB), the log warned `prompt chunk 4864 tokens, not 8192: CUDA1 ... has 1910 expert-cache slots`, and the results were less steady than with the helper cache.
## Observations and questions
1. Nothing failed on Pascal: build, start, MTP, the server and the benchmark all worked; `cuobjdump` shows sm_60 and sm_75 code in `engine-cuda12/strata`. I did not compare outputs numerically with another GPU (the answers read fluently in Japanese).
2. **A second GPU as a helper cache works, but it took some digging on a 2-GPU PC.** With `"gpu": [0, 1]` the server adds `--layer-split auto` itself; `--expert-cache-device1 4500 --remote-expert-opt` then stops the engine with `--expert-cache-remote with a layer split needs a GPU that runs no stage (2 visible, 2 used by the split)`, and an explicit `--layer-split 48` was overridden. With `"gpu": 0` the second card is hidden (`CUDA1 experts: CUDA device is not visible`). What worked was keeping `"gpu": 0` and making the second card visible through the config's `"env"` (`CUDA_VISIBLE_DEVICES=0,1`). Is there a supported way to do this (for example a `--helper-gpus 1` option in setup)? The layer-split log line suggests `--expert-cache-device1 N`, but docs/SECOND_GPU.md describes a patch and `--gpu1-experts`, so a short two-GPU recipe in the docs would help.
3. On the 32 GB PC the low-RAM mode works well only while most of the experts fit the OS file cache (other services stopped). The first start wrote a 34 GB `experts.bin` even though the 0.1.31 notes say a native pack can read straight from the GGUF files when there is no `experts.bin`. Is that expected for this path?
4. The Docker entry point cannot pass `--gguf-dir` (or extra engine arguments); I put the downloaded shards into `/data/models/Q2_0/` and edited the saved config for the helper run. The folder name is case-sensitive (`q2_0` made setup start downloading again).
## Dockerfile.cuda12 (diff against the shipped Dockerfile)
```
FROM nvidia/cuda:13.0.0-devel-ubuntu24.04 -> FROM nvidia/cuda:12.9.1-devel-ubuntu24.04
ARG CUDA_ARCHITECTURES=75;80;86;89;120 -> ARG CUDA_ARCHITECTURES=60;75
ARG BUILD_VISION=1 -> ARG BUILD_VISION=0
cmake_build(... "build", ...) -> cmake_build(... "build-cuda12", ... "-DSTRATA_EXPERIMENTAL_SM60=ON" ...)
eng = setup.ROOT / "engine" -> eng = setup.ROOT / setup.ENGINE12_DIR # engine-cuda12/
BUILD.json meta -> + "toolkit": 12
RUN rm -rf build build-vision -> RUN rm -rf build-cuda12 build-vision-cuda12
```
Run with `-e STRATA_CUDA=12 -e STRATA_EXPERIMENTAL_SM60=1`. If it helps, I can send this as a pull request (an `ARG`-driven variant of the Dockerfile) and add a row for the P100 to the support matrix in `docs/OLDER_GPUS.md`. I can also run `needle_bench.py` or other tests on this PC (Pascal + Turing, 32 GB RAM).
Attachments:
- P100 only (benchmark.py): [summary.json](https://github.com/user-attachments/files/33097768/summary.json), [engine.log](https://github.com/user-attachments/files/33097775/engine.log)
- P100 + RTX 2070 SUPER as a helper expert cache: [summary.json](https://github.com/user-attachments/files/33097786/summary.json), [engine.log](https://github.com/user-attachments/files/33097793/engine.log)
- Two-GPU layer split (engine log): [engine-layer-split-2gpu.log](https://github.com/user-attachments/files/33097724/engine-layer-split-2gpu.log)
- Dockerfile.cuda12: [Dockerfile.cuda12.txt](https://github.com/user-attachments/files/33097831/Dockerfile.cuda12.txt)
Credentials removed; the benchmark prompts are the synthetic ones generated by `benchmark.py`.Mehr auf der Site
Links zu Install, Modellen, Releases.