Issues / #1203

#1203 AMD Radeon Pro V620 (gfx1030) - Dockerfile ???

open · @luckylinux · 2 comments · View on GitHub

Setup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationLinux

Description

I am trying to come up with a `Dockerfile` / `Containerfile` for the AMD Radeon Pro V620.

This is based on:
- The `Containerfile` I have working for building `llama.cpp`
- The existing `Dockerfile` of this Project
- Some suggestions by Deepseek AI
- The [Older GPUs](https://github.com/Niko1221/Strata/blob/main/docs/OLDER_GPUS.md) Documentation Page

I tried both the Python `setup` Approach and the direct approach which seems to be what is recommended for [Older GPUs](https://github.com/Niko1221/Strata/blob/main/docs/OLDER_GPUS.md).

The Python's `setup` Approach does NOT work (apparently `strata` is trying to force some Static build onto `llama.cpp` which is not supported).

The Manual `cmake` builds apparently only the CPU target (not the GPU) and only `strata`, not `llama.cpp`.

Not sure if I should use the `llama.cpp` Image (well, the builder Image at least in the `builder` stage) as a Base for building `strata`, which is what the Documentation seems to suggest for the [AMD Radeon Instinct Mi50](https://github.com/Niko1221/Strata/blob/main/docs/AMD_HIP.md#gfx906-instinct-mi50--mi60-radeon-vii-wave64-built-from-source).

`Containerfile` for `llama.cpp` (working):
```
ARG LLAMA_VERSION

# Build Environment
FROM docker.MYDOMAIN.TLD/docker.io/library/debian:trixie-slim AS builder

ARG LLAMA_VERSION
ARG DEBIAN_FRONTEND=noninteractive

RUN --mount=type=cache,mode=0777,target=/var/cache/apk,sharing=locked \
    --mount=type=cache,mode=0777,target=/var/lib/apk,sharing=locked \
    apt-get update && apt-get install -y \
    wget gnupg2 ca-certificates git build-essential cmake ninja-build

WORKDIR /build
RUN git clone https://github.com/ggml-org/llama.cpp . && \
    git fetch --tags && \
    git checkout ${LLAMA_VERSION}

# Register AMD's official ROCm repository for Debian 13
RUN mkdir --parents --mode=0755 /etc/apt/keyrings && \
    wget -qO- https://stable.repo.amd.com/rocm/gpg/packages.gpg | gpg --dearmor > /etc/apt/keyrings/amdrocm.gpg && \
    tee /etc/apt/sources.list.d/amdrocm-stable.sources <<EOF
X-Repo-Id: amdrocm-stable
Types: deb
URIs: https://stable.repo.amd.com/rocm/core/packages/debian13/
Suites: stable
Components: main
Architectures: amd64
Signed-By: /etc/apt/keyrings/amdrocm.gpg
Enabled: yes
EOF

# Install AMD's official SDK developer bundle for your specific GPU architecture (gfx1030)
RUN --mount=type=cache,mode=0777,target=/var/cache/apt,sharing=locked \
    --mount=type=cache,mode=0777,target=/var/lib/apt,sharing=locked \
    apt-get update && apt-get install -y amdrocm-core-dev10.0-gfx1030

# Fix for the Builder Stage asset bundling step
ENV HIPCXX=/opt/rocm/bin/hipcc

# Add this line right before running CMake to allow asset compilation binaries to link
ENV LD_LIBRARY_PATH=/opt/rocm/lib:/usr/local/lib

# Set the HIP compiler to clang (adjust version if your path differs)
# Compile standard GPU package structure safely
#    -DGGML_HIPBLAS=ON \
#    -DCMAKE_HIP_ARCHITECTURES="gfx1030" \ 
#    -DAMDGPU_TARGETS="gfx1030" \ 
RUN export HIPCXX=/opt/rocm/bin/hipcc && \
    cmake -B build -G Ninja \
    -DGGML_HIP=ON \
    -DGPU_TARGETS="gfx1030" \ 
    -DCMAKE_BUILD_TYPE=Release \
    -DGGML_NATIVE=OFF \
    -DBUILD_SHARED_LIBS=OFF \
    -DCMAKE_C_COMPILER=/opt/rocm/bin/hipcc \
    -DCMAKE_CXX_COMPILER=/opt/rocm/bin/hipcc \
    && cmake --build build --config Release --target llama-server llama-cli -j 8

# Runtime Environment
FROM docker.MYDOMAIN.TLD/docker.io/library/debian:trixie-slim AS runtime

ARG DEBIAN_FRONTEND=noninteractive

RUN --mount=type=cache,mode=0777,target=/var/cache/apk,sharing=locked \
    --mount=type=cache,mode=0777,target=/var/lib/apk,sharing=locked \
    apt-get update && \
    apt-get install -y wget gnupg2 ca-certificates libgomp1 && \
    mkdir -p /etc/apt/keyrings

# Re-register the AMD mirror to grab the slim production shared objects (.so entries)
RUN mkdir --parents --mode=0755 /etc/apt/keyrings && \
    wget -qO- https://stable.repo.amd.com/rocm/gpg/packages.gpg | gpg --dearmor > /etc/apt/keyrings/amdrocm.gpg && \
    tee /etc/apt/sources.list.d/amdrocm-stable.sources <<EOF
X-Repo-Id: amdrocm-stable
Types: deb
URIs: https://stable.repo.amd.com/rocm/core/packages/debian13/
Suites: stable
Components: main
Architectures: amd64
Signed-By: /etc/apt/keyrings/amdrocm.gpg
Enabled: yes
EOF

# Install AMD's official basic base runtime layer targeting the V620 architecture
RUN --mount=type=cache,mode=0777,target=/var/cache/apt,sharing=locked \
    --mount=type=cache,mode=0777,target=/var/lib/apt,sharing=locked \
    apt-get update && apt-get install -y amdrocm10.0-gfx1030

WORKDIR /app

# CRUCIAL ARCHITECTURAL ALIGNMENT: 
# Copy the executables AND ALL compiled .so files into the exact same folder (/app).
# This satisfies llama.cpp's dynamic plugin loader lookup logic, preventing status 139 segfaults.
COPY --from=builder /build/build/bin/* /app/

# Map the system library cache directly to our execution path wrapper
# ENV LD_LIBRARY_PATH=/app:/usr/local/lib
# RUN ldconfig

# Set hardware mask states natively inside the wrapper image
# ENV HSA_OVERRIDE_GFX_VERSION=9.0.6
# ENV HSA_ENABLE_SDMA=0

# Establish environmental execution configurations for RDNA 2 compute handling
ENV HSA_ENABLE_SDMA=0
ENV LD_LIBRARY_PATH=/opt/rocm/lib:/usr/local/lib

# Refresh the system's dynamic linker cache so it recognizes the new libraries
RUN ldconfig

# Run Server
EXPOSE 8080
ENTRYPOINT ["/app/llama-server"]
CMD ["--help"]

```

`Containerfile` for `strata`:
```
ARG STRATA_VERSION

# Build Environment
FROM docker.MYDOMAIN.TLD/docker.io/library/debian:trixie-slim AS builder

# STRATA_EXECV=1: setup.py replaces itself with the server, so the server is PID 1
# and docker stop's SIGTERM reaches it (see setup.start). Normal Linux starts, which
# don't set it, keep spawning the server as a child.
ENV DEBIAN_FRONTEND=noninteractive \
    PYTHONUNBUFFERED=1 \
    LANG=C.UTF-8 \
    ROCM_PATH=/opt/rocm \
    HIP_PATH=/opt/rocm \
    HIPCXX=/opt/rocm/bin/hipcc \
    STRATA_EXECV=1

RUN --mount=type=cache,mode=0777,target=/var/cache/apk,sharing=locked \
    --mount=type=cache,mode=0777,target=/var/lib/apk,sharing=locked \
    apt-get update && apt-get install -y --no-install-recommends \
        build-essential ca-certificates curl wget gnupg2 git build-essential cmake ninja-build libatomic1 libgomp1 \
        python3 python3-pip python3-venv unzip

WORKDIR /build

RUN git clone https://github.com/Niko1221/Strata.git . && \
    git fetch --tags && \
    git checkout ${STRATA_VERSION}

# Register AMD's official ROCm repository for Debian 13
RUN mkdir --parents --mode=0755 /etc/apt/keyrings && \
    wget -qO- https://stable.repo.amd.com/rocm/gpg/packages.gpg | gpg --dearmor > /etc/apt/keyrings/amdrocm.gpg && \
    tee /etc/apt/sources.list.d/amdrocm-stable.sources <<EOF
X-Repo-Id: amdrocm-stable
Types: deb
URIs: https://stable.repo.amd.com/rocm/core/packages/debian13/
Suites: stable
Components: main
Architectures: amd64
Signed-By: /etc/apt/keyrings/amdrocm.gpg
Enabled: yes
EOF

# Install AMD's official SDK developer bundle for your specific GPU architecture (gfx1030)
RUN --mount=type=cache,mode=0777,target=/var/cache/apt,sharing=locked \
    --mount=type=cache,mode=0777,target=/var/lib/apt,sharing=locked \
    apt-get update && apt-get install -y amdrocm-core-dev10.0-gfx1030

# Fix for the Builder Stage asset bundling step
ENV HIPCXX=/opt/rocm/bin/hipcc

# Add this line right before running CMake to allow asset compilation binaries to link
ENV LD_LIBRARY_PATH=/opt/rocm/lib:/usr/local/lib

# Parameters
ARG CUDA_ARCHITECTURES=
ARG AMDGPU_TARGETS=gfx1030
ARG BUILD_VISION=1

RUN python3 -m venv .venv \
    && .venv/bin/pip install --no-cache-dir --upgrade pip \
    && .venv/bin/pip install --no-cache-dir -r requirements.txt \
    && chmod +x setup.sh docker-entrypoint.sh

# llama.cpp at the pinned commit, then the engine and the image encoder, built
# exactly the way setup.py builds them. BUILD.json is what setup.py reads to
# decide whether an engine is current: source=local with a matching src hash
# means the first start reuses it instead of recompiling.
#RUN .venv/bin/python - <<'PYEOF'
#import json, os, pathlib, shutil
#import setup
#
#llama = setup.get_llama_cpp()
#nvcc, _ = setup.find_nvcc()
#arch = os.environ.get("AMDGPU_TARGETS", "gfx1030").strip().strip('"')
#vision = "gpu" if os.environ.get("BUILD_VISION", "1") == "1" else "none"
#
#setup.cmake_build(setup.ROOT, setup.ROOT / "build", "strata",
#                  ["-DGGML_HIP=ON",
#                   "-DGPU_TARGETS=gfx1030",
#                   "-DCMAKE_BUILD_TYPE=Release",
#                   "-DGGML_NATIVE=OFF",
#                   "-DBUILD_SHARED_LIBS=OFF",
#                   "-DGGML_STATIC=OFF",
#                   "-DCMAKE_C_COMPILER=/opt/rocm/bin/hipcc",
#                   "-DCMAKE_CXX_COMPILER=/opt/rocm/bin/hipcc",
#                   "-DSTRATA_ENABLE_CUDA=OFF",
#                   "-DSTRATA_BUILD_TESTS=OFF",
#                   f"-DCMAKE_HIP_ARCHITECTURES={arch}",
#                   f"-DCMAKE_HIP_COMPILER=/opt/rocm/llvm/bin/clang++",
#                   f"-DSTRATA_GGML_DIR={llama}",
#                  ],
#                  None,
#                  "build-strata.sh"
#                 )
#
#if vision != "none":
#    setup.cmake_build(setup.ROOT / "tools" / "vision", setup.ROOT / "build-vision", "strata-vision",
#                      ["-DGGML_HIP=ON",
#                       "-DSTRATA_VISION_CUDA=OFF",
#                       "-DGPU_TARGETS=gfx1030",
#                       "-DCMAKE_BUILD_TYPE=Release",
#                       "-DGGML_NATIVE=OFF",
#                       "-DBUILD_SHARED_LIBS=OFF",
#                       "-DGGML_STATIC=OFF",
#                       "-DCMAKE_C_COMPILER=/opt/rocm/bin/hipcc",
#                       "-DCMAKE_CXX_COMPILER=/opt/rocm/bin/hipcc",
#                       "-DSTRATA_ENABLE_CUDA=OFF",
#                       "-DSTRATA_BUILD_TESTS=OFF",
#                       f"-DCMAKE_HIP_ARCHITECTURES={arch}",
#                       f"-DCMAKE_HIP_COMPILER=/opt/rocm/llvm/bin/clang++",
#                       f"-DLLAMA_DIR={llama}",
#                       ],
#                       None,
#                       "build-vision.sh"
#                     )
#
#eng = setup.ROOT / "engine"
#eng.mkdir(exist_ok=True)
#shutil.copy2(setup.ROOT / "build" / setup.sh, eng / setup.sh)
#if vision != "none":
#    shutil.copy2(setup.ROOT / "build-vision" / "bin" / setup.sh, eng / setup.sh)
#bindir = pathlib.Path(nvcc).parent
#meta = {"source": "local", "version": setup.source_version(),
#        "archs": [int(a.split("-")[0]) for a in arch.split(";") if a.split("-")[0].isdigit()], "vision": vision,
#        "vision": vision,
#        "backend": "hip",
#        "src": setup.source_hash(setup.ENGINE_SOURCES),
#        "vision_src": setup.source_hash(setup.VISION_SOURCES) if vision != "none" else None}
#(eng / "BUILD.json").write_text(json.dumps(meta, indent=1))
#PYEOF

# Patch
RUN sed -i 's|set(GGML_STATIC ON CACHE BOOL "" FORCE)|set(GGML_STATIC OFF CACHE BOOL "" FORCE)|g' CMakeLists.txt && \
    sed -i 's|set(GGML_STATIC ON CACHE BOOL "" FORCE)|set(GGML_STATIC OFF CACHE BOOL "" FORCE)|g' sycl/CMakeLists.txt && \
    grep -ri GGML_STATIC

# These Flags are not used anyways
#    -DLLAMA_DIR=/build/llama.cpp \
#    -DSTRATA_VISION_CUDA=OFF \

#    sed -i 's/if (GGML_STATIC)/if (FALSE)/' \
#    /build/build/_deps/strata_llamacpp-src/ggml/src/ggml-hip/CMakeLists.txt \

# Build
RUN export HIPCXX=/opt/rocm/bin/hipcc && \
    cmake -B build -G Ninja \
    -DGGML_HIP=ON \
    -DSTRATA_ENABLE_CUDA=OFF \
    -DGPU_TARGETS="gfx1030" \
    -DCMAKE_BUILD_TYPE=Release \
    -DGGML_NATIVE=OFF \
    -DBUILD_SHARED_LIBS=OFF \
    -DGGML_STATIC=OFF \
    -DCMAKE_C_COMPILER=/opt/rocm/bin/hipcc \
    -DCMAKE_CXX_COMPILER=/opt/rocm/bin/hipcc \
    -DSTRATA_BUILD_TESTS=OFF \
    && cmake --build build --config Release -j 8

# the cmake trees are build-time only; the engine itself is what the container needs
# RUN rm -rf build

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.