Issues / #1203
#1203 AMD Radeon Pro V620 (gfx1030) - Dockerfile ???
open · @luckylinux · 2 コメント · GitHub で見る
Setup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationLinux
本文
I am trying to come up with a `Dockerfile` / `Containerfile` for the AMD Radeon Pro V620.
This is based on:
- The `Containerfile` I have working for building `llama.cpp`
- The existing `Dockerfile` of this Project
- Some suggestions by Deepseek AI
- The [Older GPUs](https://github.com/Niko1221/Strata/blob/main/docs/OLDER_GPUS.md) Documentation Page
I tried both the Python `setup` Approach and the direct approach which seems to be what is recommended for [Older GPUs](https://github.com/Niko1221/Strata/blob/main/docs/OLDER_GPUS.md).
The Python's `setup` Approach does NOT work (apparently `strata` is trying to force some Static build onto `llama.cpp` which is not supported).
The Manual `cmake` builds apparently only the CPU target (not the GPU) and only `strata`, not `llama.cpp`.
Not sure if I should use the `llama.cpp` Image (well, the builder Image at least in the `builder` stage) as a Base for building `strata`, which is what the Documentation seems to suggest for the [AMD Radeon Instinct Mi50](https://github.com/Niko1221/Strata/blob/main/docs/AMD_HIP.md#gfx906-instinct-mi50--mi60-radeon-vii-wave64-built-from-source).
`Containerfile` for `llama.cpp` (working):
```
ARG LLAMA_VERSION
# Build Environment
FROM docker.MYDOMAIN.TLD/docker.io/library/debian:trixie-slim AS builder
ARG LLAMA_VERSION
ARG DEBIAN_FRONTEND=noninteractive
RUN --mount=type=cache,mode=0777,target=/var/cache/apk,sharing=locked \
--mount=type=cache,mode=0777,target=/var/lib/apk,sharing=locked \
apt-get update && apt-get install -y \
wget gnupg2 ca-certificates git build-essential cmake ninja-build
WORKDIR /build
RUN git clone https://github.com/ggml-org/llama.cpp . && \
git fetch --tags && \
git checkout ${LLAMA_VERSION}
# Register AMD's official ROCm repository for Debian 13
RUN mkdir --parents --mode=0755 /etc/apt/keyrings && \
wget -qO- https://stable.repo.amd.com/rocm/gpg/packages.gpg | gpg --dearmor > /etc/apt/keyrings/amdrocm.gpg && \
tee /etc/apt/sources.list.d/amdrocm-stable.sources <<EOF
X-Repo-Id: amdrocm-stable
Types: deb
URIs: https://stable.repo.amd.com/rocm/core/packages/debian13/
Suites: stable
Components: main
Architectures: amd64
Signed-By: /etc/apt/keyrings/amdrocm.gpg
Enabled: yes
EOF
# Install AMD's official SDK developer bundle for your specific GPU architecture (gfx1030)
RUN --mount=type=cache,mode=0777,target=/var/cache/apt,sharing=locked \
--mount=type=cache,mode=0777,target=/var/lib/apt,sharing=locked \
apt-get update && apt-get install -y amdrocm-core-dev10.0-gfx1030
# Fix for the Builder Stage asset bundling step
ENV HIPCXX=/opt/rocm/bin/hipcc
# Add this line right before running CMake to allow asset compilation binaries to link
ENV LD_LIBRARY_PATH=/opt/rocm/lib:/usr/local/lib
# Set the HIP compiler to clang (adjust version if your path differs)
# Compile standard GPU package structure safely
# -DGGML_HIPBLAS=ON \
# -DCMAKE_HIP_ARCHITECTURES="gfx1030" \
# -DAMDGPU_TARGETS="gfx1030" \
RUN export HIPCXX=/opt/rocm/bin/hipcc && \
cmake -B build -G Ninja \
-DGGML_HIP=ON \
-DGPU_TARGETS="gfx1030" \
-DCMAKE_BUILD_TYPE=Release \
-DGGML_NATIVE=OFF \
-DBUILD_SHARED_LIBS=OFF \
-DCMAKE_C_COMPILER=/opt/rocm/bin/hipcc \
-DCMAKE_CXX_COMPILER=/opt/rocm/bin/hipcc \
&& cmake --build build --config Release --target llama-server llama-cli -j 8
# Runtime Environment
FROM docker.MYDOMAIN.TLD/docker.io/library/debian:trixie-slim AS runtime
ARG DEBIAN_FRONTEND=noninteractive
RUN --mount=type=cache,mode=0777,target=/var/cache/apk,sharing=locked \
--mount=type=cache,mode=0777,target=/var/lib/apk,sharing=locked \
apt-get update && \
apt-get install -y wget gnupg2 ca-certificates libgomp1 && \
mkdir -p /etc/apt/keyrings
# Re-register the AMD mirror to grab the slim production shared objects (.so entries)
RUN mkdir --parents --mode=0755 /etc/apt/keyrings && \
wget -qO- https://stable.repo.amd.com/rocm/gpg/packages.gpg | gpg --dearmor > /etc/apt/keyrings/amdrocm.gpg && \
tee /etc/apt/sources.list.d/amdrocm-stable.sources <<EOF
X-Repo-Id: amdrocm-stable
Types: deb
URIs: https://stable.repo.amd.com/rocm/core/packages/debian13/
Suites: stable
Components: main
Architectures: amd64
Signed-By: /etc/apt/keyrings/amdrocm.gpg
Enabled: yes
EOF
# Install AMD's official basic base runtime layer targeting the V620 architecture
RUN --mount=type=cache,mode=0777,target=/var/cache/apt,sharing=locked \
--mount=type=cache,mode=0777,target=/var/lib/apt,sharing=locked \
apt-get update && apt-get install -y amdrocm10.0-gfx1030
WORKDIR /app
# CRUCIAL ARCHITECTURAL ALIGNMENT:
# Copy the executables AND ALL compiled .so files into the exact same folder (/app).
# This satisfies llama.cpp's dynamic plugin loader lookup logic, preventing status 139 segfaults.
COPY --from=builder /build/build/bin/* /app/
# Map the system library cache directly to our execution path wrapper
# ENV LD_LIBRARY_PATH=/app:/usr/local/lib
# RUN ldconfig
# Set hardware mask states natively inside the wrapper image
# ENV HSA_OVERRIDE_GFX_VERSION=9.0.6
# ENV HSA_ENABLE_SDMA=0
# Establish environmental execution configurations for RDNA 2 compute handling
ENV HSA_ENABLE_SDMA=0
ENV LD_LIBRARY_PATH=/opt/rocm/lib:/usr/local/lib
# Refresh the system's dynamic linker cache so it recognizes the new libraries
RUN ldconfig
# Run Server
EXPOSE 8080
ENTRYPOINT ["/app/llama-server"]
CMD ["--help"]
```
`Containerfile` for `strata`:
```
ARG STRATA_VERSION
# Build Environment
FROM docker.MYDOMAIN.TLD/docker.io/library/debian:trixie-slim AS builder
# STRATA_EXECV=1: setup.py replaces itself with the server, so the server is PID 1
# and docker stop's SIGTERM reaches it (see setup.start). Normal Linux starts, which
# don't set it, keep spawning the server as a child.
ENV DEBIAN_FRONTEND=noninteractive \
PYTHONUNBUFFERED=1 \
LANG=C.UTF-8 \
ROCM_PATH=/opt/rocm \
HIP_PATH=/opt/rocm \
HIPCXX=/opt/rocm/bin/hipcc \
STRATA_EXECV=1
RUN --mount=type=cache,mode=0777,target=/var/cache/apk,sharing=locked \
--mount=type=cache,mode=0777,target=/var/lib/apk,sharing=locked \
apt-get update && apt-get install -y --no-install-recommends \
build-essential ca-certificates curl wget gnupg2 git build-essential cmake ninja-build libatomic1 libgomp1 \
python3 python3-pip python3-venv unzip
WORKDIR /build
RUN git clone https://github.com/Niko1221/Strata.git . && \
git fetch --tags && \
git checkout ${STRATA_VERSION}
# Register AMD's official ROCm repository for Debian 13
RUN mkdir --parents --mode=0755 /etc/apt/keyrings && \
wget -qO- https://stable.repo.amd.com/rocm/gpg/packages.gpg | gpg --dearmor > /etc/apt/keyrings/amdrocm.gpg && \
tee /etc/apt/sources.list.d/amdrocm-stable.sources <<EOF
X-Repo-Id: amdrocm-stable
Types: deb
URIs: https://stable.repo.amd.com/rocm/core/packages/debian13/
Suites: stable
Components: main
Architectures: amd64
Signed-By: /etc/apt/keyrings/amdrocm.gpg
Enabled: yes
EOF
# Install AMD's official SDK developer bundle for your specific GPU architecture (gfx1030)
RUN --mount=type=cache,mode=0777,target=/var/cache/apt,sharing=locked \
--mount=type=cache,mode=0777,target=/var/lib/apt,sharing=locked \
apt-get update && apt-get install -y amdrocm-core-dev10.0-gfx1030
# Fix for the Builder Stage asset bundling step
ENV HIPCXX=/opt/rocm/bin/hipcc
# Add this line right before running CMake to allow asset compilation binaries to link
ENV LD_LIBRARY_PATH=/opt/rocm/lib:/usr/local/lib
# Parameters
ARG CUDA_ARCHITECTURES=
ARG AMDGPU_TARGETS=gfx1030
ARG BUILD_VISION=1
RUN python3 -m venv .venv \
&& .venv/bin/pip install --no-cache-dir --upgrade pip \
&& .venv/bin/pip install --no-cache-dir -r requirements.txt \
&& chmod +x setup.sh docker-entrypoint.sh
# llama.cpp at the pinned commit, then the engine and the image encoder, built
# exactly the way setup.py builds them. BUILD.json is what setup.py reads to
# decide whether an engine is current: source=local with a matching src hash
# means the first start reuses it instead of recompiling.
#RUN .venv/bin/python - <<'PYEOF'
#import json, os, pathlib, shutil
#import setup
#
#llama = setup.get_llama_cpp()
#nvcc, _ = setup.find_nvcc()
#arch = os.environ.get("AMDGPU_TARGETS", "gfx1030").strip().strip('"')
#vision = "gpu" if os.environ.get("BUILD_VISION", "1") == "1" else "none"
#
#setup.cmake_build(setup.ROOT, setup.ROOT / "build", "strata",
# ["-DGGML_HIP=ON",
# "-DGPU_TARGETS=gfx1030",
# "-DCMAKE_BUILD_TYPE=Release",
# "-DGGML_NATIVE=OFF",
# "-DBUILD_SHARED_LIBS=OFF",
# "-DGGML_STATIC=OFF",
# "-DCMAKE_C_COMPILER=/opt/rocm/bin/hipcc",
# "-DCMAKE_CXX_COMPILER=/opt/rocm/bin/hipcc",
# "-DSTRATA_ENABLE_CUDA=OFF",
# "-DSTRATA_BUILD_TESTS=OFF",
# f"-DCMAKE_HIP_ARCHITECTURES={arch}",
# f"-DCMAKE_HIP_COMPILER=/opt/rocm/llvm/bin/clang++",
# f"-DSTRATA_GGML_DIR={llama}",
# ],
# None,
# "build-strata.sh"
# )
#
#if vision != "none":
# setup.cmake_build(setup.ROOT / "tools" / "vision", setup.ROOT / "build-vision", "strata-vision",
# ["-DGGML_HIP=ON",
# "-DSTRATA_VISION_CUDA=OFF",
# "-DGPU_TARGETS=gfx1030",
# "-DCMAKE_BUILD_TYPE=Release",
# "-DGGML_NATIVE=OFF",
# "-DBUILD_SHARED_LIBS=OFF",
# "-DGGML_STATIC=OFF",
# "-DCMAKE_C_COMPILER=/opt/rocm/bin/hipcc",
# "-DCMAKE_CXX_COMPILER=/opt/rocm/bin/hipcc",
# "-DSTRATA_ENABLE_CUDA=OFF",
# "-DSTRATA_BUILD_TESTS=OFF",
# f"-DCMAKE_HIP_ARCHITECTURES={arch}",
# f"-DCMAKE_HIP_COMPILER=/opt/rocm/llvm/bin/clang++",
# f"-DLLAMA_DIR={llama}",
# ],
# None,
# "build-vision.sh"
# )
#
#eng = setup.ROOT / "engine"
#eng.mkdir(exist_ok=True)
#shutil.copy2(setup.ROOT / "build" / setup.sh, eng / setup.sh)
#if vision != "none":
# shutil.copy2(setup.ROOT / "build-vision" / "bin" / setup.sh, eng / setup.sh)
#bindir = pathlib.Path(nvcc).parent
#meta = {"source": "local", "version": setup.source_version(),
# "archs": [int(a.split("-")[0]) for a in arch.split(";") if a.split("-")[0].isdigit()], "vision": vision,
# "vision": vision,
# "backend": "hip",
# "src": setup.source_hash(setup.ENGINE_SOURCES),
# "vision_src": setup.source_hash(setup.VISION_SOURCES) if vision != "none" else None}
#(eng / "BUILD.json").write_text(json.dumps(meta, indent=1))
#PYEOF
# Patch
RUN sed -i 's|set(GGML_STATIC ON CACHE BOOL "" FORCE)|set(GGML_STATIC OFF CACHE BOOL "" FORCE)|g' CMakeLists.txt && \
sed -i 's|set(GGML_STATIC ON CACHE BOOL "" FORCE)|set(GGML_STATIC OFF CACHE BOOL "" FORCE)|g' sycl/CMakeLists.txt && \
grep -ri GGML_STATIC
# These Flags are not used anyways
# -DLLAMA_DIR=/build/llama.cpp \
# -DSTRATA_VISION_CUDA=OFF \
# sed -i 's/if (GGML_STATIC)/if (FALSE)/' \
# /build/build/_deps/strata_llamacpp-src/ggml/src/ggml-hip/CMakeLists.txt \
# Build
RUN export HIPCXX=/opt/rocm/bin/hipcc && \
cmake -B build -G Ninja \
-DGGML_HIP=ON \
-DSTRATA_ENABLE_CUDA=OFF \
-DGPU_TARGETS="gfx1030" \
-DCMAKE_BUILD_TYPE=Release \
-DGGML_NATIVE=OFF \
-DBUILD_SHARED_LIBS=OFF \
-DGGML_STATIC=OFF \
-DCMAKE_C_COMPILER=/opt/rocm/bin/hipcc \
-DCMAKE_CXX_COMPILER=/opt/rocm/bin/hipcc \
-DSTRATA_BUILD_TESTS=OFF \
&& cmake --build build --config Release -j 8
# the cmake trees are build-time only; the engine itself is what the container needs
# RUN rm -rf build関連リンク
インストール・モデル・リリースへの站内リンク。