diff --git a/README.md b/README.md index 6029a190..7e096d57 100644 --- a/README.md +++ b/README.md @@ -4,6 +4,9 @@ MAD is a platform that consists of curated list of AI models that allow us to run on various GPU architectures seamlessly while tracking performance and generating dashboards for insights. +**Documentation:** start at [docs/README.md](docs/README.md). It takes you from running your first +model to running, configuring and extending the multinode inference workloads. + ## Blueprints This repository provides state-of-the-art deep learning recipes for training, inference and easy deployment on AMD Instinct GPUs. diff --git a/benchmark/kimi_k3/README.md b/benchmark/kimi_k3/README.md index f68f1c76..458788f8 100644 --- a/benchmark/kimi_k3/README.md +++ b/benchmark/kimi_k3/README.md @@ -12,6 +12,11 @@ MAD supports Kimi-K3 day-0 inference across **three** serving frameworks on AMD | **SGLang** | `pyt_sglang_kimi-k3` | `lmsysorg/sglang-rocm:rocm720-mi35x-k3-20260727` | [SGLang cookbook](https://docs.sglang.io/cookbook/autoregressive/Moonshotai/Kimi-K3) | | **ATOM** | `pyt_atom_kimi-k3` | `rocm/atom-dev:rocm7.2.4_ubuntu24.04_py3.12_pytorch2.10.0_20260727_kimi_k3` | ROCm ATOM | +> **On MI300X (gfx942)?** All four models above carry `skip_gpu_arch: gfx942` — +> the ~1.5 TB checkpoint does not fit a single 8-GPU MI300X node, so K3 needs +> multi-node sharding there. See [`mi300x/`](mi300x/README.md) for the 2-node +> colocated and 4-node prefill/decode-disaggregated recipes. + ## Hardware requirements - **8x MI350X or MI355X** (TP8) diff --git a/benchmark/kimi_k3/mi300x/README.md b/benchmark/kimi_k3/mi300x/README.md new file mode 100644 index 00000000..340e37ba --- /dev/null +++ b/benchmark/kimi_k3/mi300x/README.md @@ -0,0 +1,5 @@ +# Kimi-K3 inference on AMD Instinct MI300X (gfx942) + +This guide moved to [docs/kimi-k3.md](../../../docs/kimi-k3.md), which covers Kimi-K3 on MI300X and +MI355X, colocated and disaggregated: the cards, the image, site settings, known failure modes and +results. diff --git a/docker/sglang_disagg_inference.ubuntu.amd.Dockerfile b/docker/sglang_disagg_inference.ubuntu.amd.Dockerfile index 47875e20..a92a7b43 100644 --- a/docker/sglang_disagg_inference.ubuntu.amd.Dockerfile +++ b/docker/sglang_disagg_inference.ubuntu.amd.Dockerfile @@ -36,6 +36,10 @@ WORKDIR /sgl-workspace RUN pip install --upgrade sglang-router +# Runtime deps of the disagg launcher and proxy. Baked in so the launcher scripts +# do not have to install them on every node of every run. +RUN pip install py-spy flask pyyaml + WORKDIR /sgl-workspace/mori ARG MORI_COMMIT="158c7e8335a0b19b3f1f422ff134d7869252135e" diff --git a/docker/vllm_disagg_inference.ubuntu.amd.Dockerfile b/docker/vllm_disagg_inference.ubuntu.amd.Dockerfile index 06dd18aa..74eb5388 100644 --- a/docker/vllm_disagg_inference.ubuntu.amd.Dockerfile +++ b/docker/vllm_disagg_inference.ubuntu.amd.Dockerfile @@ -225,6 +225,13 @@ ENV _ROCM_DIR=/opt/rocm \ _RIXL_BRANCH=f33a5599 \ _RIXL_INSTALL_DIR=/usr/local/RIXL/install \ _NIXLBENCH_INSTALL_DIR=/usr/local/RIXL +# rocSHMEM is pinned like every other source here. It tracked rocm-systems `develop` +# until rocm-systems 16dc5673 (2026-09-25, "replace hip atomic builtins with scoped +# counterparts") made projects/rocshmem/src/atomic.hpp use scoped atomics on float and +# double, which this image's ROCm 7.2 compiler rejects ("address argument to atomic +# operation must be a pointer to integer or pointer"); every build of this image failed +# from then on. This is that commit's parent, the rocSHMEM the image last built with. +ARG ROCSHMEM_REF=9c43b23229cf2f835b838f72c0d1af3d4876167c RUN if [ "${WITH_NIXL}" != "1" ]; then \ echo "WITH_NIXL=${WITH_NIXL}: skipping UCX/RIXL/rocSHMEM/DeepEP (MoRI-EP + base DeepEP only)"; \ else set -e && \ @@ -257,7 +264,7 @@ RUN if [ "${WITH_NIXL}" != "1" ]; then \ --config-settings=setup-args="-Ddisable_gds_backend=true" . && \ # rocSHMEM (DeepEP dep) cd /tmp && git clone --no-checkout --filter=blob:none https://github.com/ROCm/rocm-systems.git && \ - cd rocm-systems && git sparse-checkout set --cone projects/rocshmem && git checkout develop && \ + cd rocm-systems && git sparse-checkout set --cone projects/rocshmem && git checkout "${ROCSHMEM_REF}" && \ mkdir -p /tmp/rocshmem-build && cd /tmp/rocshmem-build && \ /tmp/rocm-systems/projects/rocshmem/scripts/build_configs/all_backends \ -DUSE_EXTERNAL_MPI=OFF -DGPU_TARGETS="${GFX_COMPILATION_ARCH}" && \ diff --git a/docker/vllm_kimi_k3.ubuntu.amd.Dockerfile b/docker/vllm_kimi_k3.ubuntu.amd.Dockerfile new file mode 100644 index 00000000..dcf93395 --- /dev/null +++ b/docker/vllm_kimi_k3.ubuntu.amd.Dockerfile @@ -0,0 +1,343 @@ +# CONTEXT {'gpu_vendor': 'AMD', 'guest_os': 'UBUNTU'} +############################################################################### +# +# MIT License +# +# Copyright (c) 2026 Advanced Micro Devices, Inc. +# +# Permission is hereby granted, free of charge, to any person obtaining a copy +# of this software and associated documentation files (the "Software"), to deal +# in the Software without restriction, including without limitation the rights +# to use, copy, modify, merge, publish, distribute, sublicense, and/or sell +# copies of the Software, and to permit persons to whom the Software is +# furnished to do so, subject to the following conditions: +# +# The above copyright notice and this permission notice shall be included in all +# copies or substantial portions of the Software. +# +# THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR +# IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, +# FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE +# AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER +# LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, +# OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE +# SOFTWARE. +# +################################################################################# +# ============================================================================= +# vllm_kimi_k3.ubuntu.amd.Dockerfile +# +# The one Kimi-K3 image for every Kimi-K3 vLLM card, on MI300X (gfx942) and +# MI355X (gfx950), colocated multinode and disaggregated alike. It replaces +# pyt_vllm_kimi_k3_mi300x, pyt_vllm_kimi_k3_mi355x and vllm_disagg_inference.kimik3, +# which were the same stack apart from the GPU arch and a few pins +# (vLLM, MoRI and AITER were identical in all three). +# +# GPU ARCH. Built for exactly one arch, MAD_SYSTEM_GPU_ARCHITECTURE (gfx942 or +# gfx950): madengine passes it, and the MAD CI pipeline sets it from the compute +# nodes it probed. Everything arch-specific follows it: MoRI's JIT target, the +# vLLM compile, and, with WITH_NIXL=1, rocSHMEM and DeepEP. Runtime differences +# between the arches (AITER MLA gfx950-only, int4 MoE requant on gfx942) are +# recipe knobs in scripts/*/models.yaml, not build steps. +# +# AITER. Built from source at the commit the Kimi-K3 release images ship: +# ROCm/aiter 68e42f5f (0.1.17.dev395, "k3-for-amd"). Both earlier donor images, +# amdsiloai/vllm:kimi-k3-mi325x-release-v2 and vllm/vllm-openai-rocm:kimi-k3, +# carry that same layer (digest e2b3951e36ca), with its Kimi-K3 tuned MoE +# configs, aiter.ops.triton.conv, and flydsl 0.2.4. Building it here replaces +# copying their site-packages over a separately installed AITER, so the image +# has one base and no second image to pull. +# +# Build (the image copies nothing from the build context, so any context works): +# docker build -f docker/vllm_kimi_k3.ubuntu.amd.Dockerfile \ +# --build-arg MAD_SYSTEM_GPU_ARCHITECTURE=gfx942 -t /vllm-kimi-k3:gfx942 docker +# +# WITH_NIXL=1 (default) adds UCX/RIXL/rocSHMEM/DeepEP for the rixl connector; +# the Kimi recipes run moriio, so =0 is a faster build with the same serving path. +# +# STATUS +# gfx942 (MI300X): the stack the Kimi-K3 cards run today; single-needle NIAH passed +# 10K-900K on 2 prefill + 2 decode nodes with these pins (PR #193), with AITER +# supplied by grafting the same 68e42f5f build this file now compiles. +# gfx950 (MI355X): BRING-UP, NOT VALIDATED. Nothing shows MoRIIO disagg with +# Kimi-K3 has run on gfx950: the vLLM ref is the fork's gfx942 branch. The same +# stack has run MoRIIO + MoRI-EP on gfx950 for GLM, which needed EP16 +# startup-deadlock gates and a MoRI combine fix that are NOT in this vLLM ref, and +# ionic NICs needed a patched MoRI. Expect the first gfx950 runs to find those. +# +# NO RUNTIME PATCHERS +# All Kimi-K3 MoRIIO connector fixes (4-KV-cache-group block routing, multi-chunk +# compute-progress prefill gate, KDA gather sync-free) are committed in the vLLM +# source this builds (VLLM_REF below). Nothing patches site-packages at start. +# +# PINNING +# Every source is pinned to an immutable commit SHA, not a branch name: these are +# personal forks whose branches can be force-pushed or deleted, and MAD needs the +# image to be rebuildable to the same bits a year from now. The human-readable +# branch each SHA came from is in the comment above it. +# ============================================================================= + +# Image ARGs consumed by FROM must be declared before the first FROM (buildkit +# global scope); declaring them later scopes them to a single stage and the second +# FROM resolves blank. +# +# Open ROCm vLLM CI base - same one the shared disagg image builds on. +ARG BASE_IMAGE=rocm/vllm-dev:ci_base-0fcd9b99cc9d63202da4c858d8ebc6582c9e2491 + +FROM ${BASE_IMAGE} + +ENTRYPOINT [] +WORKDIR /app + +# Pin the *build toolchain*, not just the sources. +# +# Every source below is pinned to an immutable commit SHA, which makes the build +# look reproducible -- but pip builds wheels in an isolated environment and +# resolves build dependencies (setuptools, wheel) fresh from PyPI at build time. +# So a newer setuptools published after this file was last exercised can break a +# build whose sources have not moved at all. That is exactly what happened: +# setuptools >= 80 added +# assert isinstance(self.compiler, CCompiler) +# to distutils' build_ext.build_extension, which MoRI's legacy +# Cython.Distutils.build_ext path violates, failing the amd_mori wheel with +# AssertionError: run() must precede build_extension() +# while every pinned SHA was still correct. +# +# PIP_CONSTRAINT reaches inside pip's isolated build environments, which a plain +# `pip install setuptools==X` in the image does not. Set globally so later stages +# (AITER, vLLM, router) cannot regress the same way. +ARG SETUPTOOLS_CONSTRAINT="setuptools<80" +RUN printf '%s\n' "${SETUPTOOLS_CONSTRAINT}" > /etc/pip-constraints.txt +ENV PIP_CONSTRAINT=/etc/pip-constraints.txt + +# The GPU this image is built for. No default: an image built for the wrong arch +# fails at runtime, so the build refuses to guess. K3_GFX_ARCH carries it to each +# build step (named so it is not read as a fixed arch by madengine's Dockerfile +# arch check, which parses GFX_COMPILATION_ARCH / PYTORCH_ROCM_ARCH / GPU_ARCHS). +ARG MAD_SYSTEM_GPU_ARCHITECTURE +ENV K3_GFX_ARCH=${MAD_SYSTEM_GPU_ARCHITECTURE} +RUN case "${K3_GFX_ARCH}" in \ + gfx942|gfx950) echo "Kimi-K3 image for ${K3_GFX_ARCH}" ;; \ + *) echo "MAD_SYSTEM_GPU_ARCHITECTURE must be gfx942 or gfx950, got '${K3_GFX_ARCH}'" >&2; exit 1 ;; \ + esac && mkdir -p /app && echo "GPU_ARCH=${K3_GFX_ARCH}" >> /app/versions.txt +ARG MAX_JOBS=32 +ARG NVCC_THREADS=8 +# UCX/RIXL/rocSHMEM/DeepEP for the rixl connector. On by default so the image serves +# every connector; the Kimi recipes use moriio, and =0 builds faster without it. +ARG WITH_NIXL=1 +ARG NIC_COMPILATION_ARCH="cx7" + +# ----------------------------------------------------------------------------- +# 1. MoRI v1.2.2 (the K3 recipe's pin; the shared disagg image pins v1.2.1). +# JIT-built, so this swaps the sources the EP kernels compile from at runtime. +# Do NOT pass USE_IONIC=OFF / USE_BNXT=OFF - disabling NIC backends produced a +# MoRI that deadlocked at the cross-node EP all-to-all init. +# ----------------------------------------------------------------------------- +ARG MORI_REPO=https://github.com/ROCm/mori.git +# tag v1.2.2 +ARG MORI_REF=fe12a11a7d6c6acd0771b772366ed9ed5e0d3d44 +ENV MORI_GPU_ARCHS=${K3_GFX_ARCH} +# UMBP needs gRPC headers absent from this base and is unrelated to EP dispatch. +ENV BUILD_UMBP=OFF BUILD_UMBP_SPDK=OFF +RUN sed -i 's|http://|https://|g' /etc/apt/sources.list 2>/dev/null || true && \ + sed -i 's|http://|https://|g' /etc/apt/sources.list.d/*.list 2>/dev/null || true && \ + apt-get update && apt-get install -y --no-install-recommends \ + git build-essential cmake ninja-build ccache libssl-dev pkg-config curl ca-certificates && \ + pip install meson==0.64.0 "pybind11[global]" tqdm prettytable && \ + pip uninstall -y amd_mori amd-mori amd-mori-nightly mori 2>/dev/null || true && \ + rm -rf /tmp/mori-src && \ + git clone --recursive "${MORI_REPO}" /tmp/mori-src && \ + cd /tmp/mori-src && git checkout "${MORI_REF}" && git submodule update --init --recursive && \ + BUILD_UMBP=OFF pip install . && \ + python3 -c "import mori, mori.io, mori.ops; print('MoRI OK at', mori.__path__[0])" && \ + mkdir -p /app && echo "MORI_REF=${MORI_REF}" >> /app/versions.txt && \ + rm -rf /tmp/mori-src + +# ----------------------------------------------------------------------------- +# 2. AITER from source at the Kimi-K3 release commit, and flydsl 0.2.4. +# 68e42f5f carries the Kimi-K3 tuned MoE configs (kimik3_{a8w4,fp4}_tuned_fmoe, +# without which K3's MoE profiling shape falls back to a heuristic FlyDSL kernel +# that aborts LLVM), aiter.ops.triton.conv (K3's vision tower), and the #3658 +# top_k_top_p fix DP-EP disagg needs. K3's int4 SiTUv2 path on gfx942 +# (_setup_kernel_k3_situ_gfx942 -> compile_moe_gemm1) needs flydsl >= 0.2.4. +# Same build as the GLM-5.1 image's AITER step; kernels JIT for the GPU at runtime. +# ----------------------------------------------------------------------------- +ARG AITER_REPO=https://github.com/ROCm/aiter.git +# The release images built it from branch k3-for-amd of a fork that is not public. The +# commit is in ROCm/aiter's fork network but on none of its branches, so a clone does not +# have it ("reference is not a tree"); fetching it by SHA does. --tags brings the release +# tags its version is derived from (the images report 0.1.17.dev395+g68e42f5f4; tags +# added upstream since then change the version number, not the code). +ARG AITER_REF=68e42f5f461556596ae294200f1a3f13378c8582 +RUN rm -rf /tmp/aiter-src && mkdir -p /tmp/aiter-src && cd /tmp/aiter-src && \ + git init -q && git remote add origin "${AITER_REPO}" && \ + git fetch -q --tags origin "${AITER_REF}" && git checkout -q FETCH_HEAD && \ + git submodule update --init --recursive && \ + (pip uninstall -y amd_aiter amd-aiter aiter 2>/dev/null || true) && \ + GPU_ARCHS="${K3_GFX_ARCH}" pip install --no-build-isolation --no-deps -v . && \ + pip install --no-deps --force-reinstall "flydsl==0.2.4" && \ + python3 -c "import importlib.metadata as m; v=m.version('amd-aiter'); \ +assert '68e42f5f' in v, v; assert m.version('flydsl')=='0.2.4'; \ +f=[str(x) for x in m.files('amd-aiter') if str(x).startswith('aiter/ops/triton/conv/')]; \ +assert f, 'aiter.ops.triton.conv missing'; print('AITER OK', v, '+ flydsl 0.2.4')" && \ + echo "AITER_REF=${AITER_REF}" >> /app/versions.txt && \ + rm -rf /tmp/aiter-src /opt/vllm_cache/aiter_jit /root/.aiter + +# ----------------------------------------------------------------------------- +# 3. vLLM: full source compile of the K3 + MoRIIO branch. All K3 connector fixes +# are committed in this tree - there is no runtime patcher: +# - 4-KV-cache-group block routing (K3 allocates 3 KDA/mamba groups + 1 MLA; +# the stock connector hardcoded 2-group indices and sent MLA KV to mamba +# block ids, so decode read empty blocks and generated without context) +# - multi-chunk prefill transfer gated on compute progress, not block count +# (the block-count gate fired after chunk 1 for prompts fitting one padded +# block, so only max_num_batched_tokens of KV ever crossed) +# - KDA gather made sync-free (a per-layer per-chunk device->CPU sync that +# turned >500K-token prefills into an apparent hang) +# Repo is public; no credentials are needed or accepted here (the upstream +# recipe took a GH_TOKEN build-arg, which would bake the token into image +# metadata - removed). +# ----------------------------------------------------------------------------- +ARG VLLM_REPO=https://github.com/raviguptaamd/vllm.git +# branch kimi-k3-wideep-disagg-fullsource-v3 @ 2026-08-11 +ARG VLLM_REF=862bfd8ca4db78b9cbcbcf9ec6013638e3ae6543 +ENV VLLM_TARGET_DEVICE=rocm \ + PYTORCH_ROCM_ARCH=${K3_GFX_ARCH} \ + MAX_JOBS=${MAX_JOBS} +# MAX_JOBS/NVCC_THREADS are set INLINE on the pip line: under the legacy builder +# the ENV above does not reach the pip build subprocess, and vLLM's setup.py +# compute_num_jobs then dies on `int("")`. +RUN rm -rf /tmp/vllm-src && \ + git clone "${VLLM_REPO}" /tmp/vllm-src && \ + cd /tmp/vllm-src && git checkout "${VLLM_REF}" && \ + echo "VLLM_REF=${VLLM_REF}" >> /app/versions.txt && \ + pip uninstall -y vllm 2>/dev/null || true && \ + MAX_JOBS="${MAX_JOBS:-32}" NVCC_THREADS="${NVCC_THREADS:-8}" \ + pip install --no-deps --no-build-isolation -v . && \ + python3 -c "import vllm; print('vLLM', vllm.__version__, 'from', vllm.__file__)" && \ + rm -rf /tmp/vllm-src + +# Cross-check MoRI + AITER survived the vLLM install (no silent downgrade). +RUN python3 - <<'PYEOF' +from importlib.metadata import version as v, PackageNotFoundError +def get(names): + for n in names: + try: return v(n) + except PackageNotFoundError: pass + return None +av = get(("amd-aiter", "amd_aiter", "aiter")) +assert av and "68e42f5f" in av, f"AITER replaced by the vLLM install: {av!r}" +import mori, mori.io, mori.ops +print("Post-vLLM check OK: AITER", av, "+ MoRI importable") +PYEOF + +# ----------------------------------------------------------------------------- +# 4. vllm-router: DP-rank round-robin + the 2P2D KV-notify fix +# (remote_dp_rank_override + remote_dp_size). Without the notify fix a 2P2D +# EP16 run reproducibly wedges on "remote blocks never arrived" deferred-write +# expiries, because decode's notify targets the wrong DP rank. +# Same source the shared disagg image uses; pinned to its SHA here. +# ----------------------------------------------------------------------------- +ARG ROUTER_REPO=https://github.com/raviguptaamd/router.git +# branch ravgupta/discovery-dp-rank-roundrobin @ 2026-08-24 (the 2P2D KV-notify fix) +ARG ROUTER_REF=82dc9811af17412e6e24b5942a5486bc502df23a +ARG RUST_TOOLCHAIN=1.88.0 +RUN if ! command -v cargo >/dev/null 2>&1; then \ + curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y --default-toolchain "${RUST_TOOLCHAIN}"; \ + fi && \ + export PATH="/root/.cargo/bin:${PATH}" && \ + rm -rf /tmp/vllm-router-src && \ + git clone --filter=blob:none "${ROUTER_REPO}" /tmp/vllm-router-src && \ + cd /tmp/vllm-router-src && git checkout "${ROUTER_REF}" && \ + cargo build --release && \ + install -m 755 target/release/vllm-router /usr/local/bin/vllm-router && \ + vllm-router --help 2>&1 | grep -q moriio && \ + echo "VLLM_ROUTER_REF=${ROUTER_REF}" >> /app/versions.txt && \ + rm -rf /tmp/vllm-router-src + +# ----------------------------------------------------------------------------- +# 4b. NIXL/RIXL transport (WITH_NIXL=1, the default). Identical to the shared disagg +# image, with rocSHMEM pinned as it pins it. +# ----------------------------------------------------------------------------- +ENV _ROCM_DIR=/opt/rocm \ + _UCX_SOURCE=https://github.com/ROCm/ucx.git \ + _UCX_BRANCH=da3fac2a \ + _UCX_INSTALL_DIR=/usr/local/ucx/ \ + _RIXL_SOURCE=https://github.com/ROCm/RIXL.git \ + _RIXL_BRANCH=f33a5599 \ + _RIXL_INSTALL_DIR=/usr/local/RIXL/install \ + _NIXLBENCH_INSTALL_DIR=/usr/local/RIXL +ARG ROCSHMEM_REF=9c43b23229cf2f835b838f72c0d1af3d4876167c +RUN if [ "${WITH_NIXL}" != "1" ]; then \ + echo "WITH_NIXL=${WITH_NIXL}: skipping UCX/RIXL/rocSHMEM/DeepEP (MoRI-EP only)"; \ + else set -e && \ + echo "WITH_NIXL=1: building UCX + RIXL + rocSHMEM + DeepEP" && \ + apt-get update && apt-get install -y \ + autoconf automake libtool autogen pkg-config m4 gcc make \ + librdmacm-dev rdmacm-utils infiniband-diags ibverbs-utils perftest ethtool \ + libibverbs-dev rdma-core strace libgflags-dev \ + libaio-dev liburing-dev libcpprest-dev libgrpc-dev libgrpc++-dev \ + libprotobuf-dev protobuf-compiler-grpc wget && \ + pip install meson==0.64.0 "pybind11[global]" pyyaml && \ + cd /tmp && git clone "${_UCX_SOURCE}" && cd ucx && git checkout "${_UCX_BRANCH}" && \ + ./autogen.sh && mkdir -p build && cd build && \ + ../configure --prefix="${_UCX_INSTALL_DIR}" --with-rocm="${_ROCM_DIR}" \ + --disable-go --disable-java --disable-assertions --enable-mt && \ + make -j && make install && \ + cd /tmp && wget -q https://github.com/google/googletest/archive/refs/tags/v1.14.0.tar.gz && \ + tar -xzf v1.14.0.tar.gz && cd googletest-1.14.0 && mkdir -p build && cd build && \ + cmake -DBUILD_SHARED_LIBS=on .. && make -j && make install && \ + cd /tmp && git clone "${_RIXL_SOURCE}" && cd RIXL && git checkout "${_RIXL_BRANCH}" && \ + meson setup build/ --prefix="${_RIXL_INSTALL_DIR}" -Ducx_path="${_UCX_INSTALL_DIR}" \ + -Ddisable_gds_backend=true -Dcudapath_inc="${_ROCM_DIR}/include" -Dcudapath_lib="${_ROCM_DIR}/lib" && \ + cd build && ninja && ninja install && cd /tmp/RIXL && \ + pip install --config-settings=setup-args="-Dcudapath_inc=${_ROCM_DIR}/include" \ + --config-settings=setup-args="-Dcudapath_lib=${_ROCM_DIR}/lib" \ + --config-settings=setup-args="-Ducx_path=${_UCX_INSTALL_DIR}" \ + --config-settings=setup-args="-Ddisable_gds_backend=true" . && \ + cd /tmp && git clone --no-checkout --filter=blob:none https://github.com/ROCm/rocm-systems.git && \ + cd rocm-systems && git sparse-checkout set --cone projects/rocshmem && git checkout "${ROCSHMEM_REF}" && \ + mkdir -p /tmp/rocshmem-build && cd /tmp/rocshmem-build && \ + /tmp/rocm-systems/projects/rocshmem/scripts/build_configs/all_backends \ + -DUSE_EXTERNAL_MPI=OFF -DGPU_TARGETS="${K3_GFX_ARCH}" && \ + cd /tmp && git clone https://github.com/ROCm/DeepEP.git && cd DeepEP && \ + PYTORCH_ROCM_ARCH="${K3_GFX_ARCH}" CFLAGS="-O3 -fPIC" \ + CXXFLAGS="-O3 -fPIC --offload-arch=${K3_GFX_ARCH}" HIP_CXX_FLAGS="-O3 -fPIC" \ + python3 setup.py --variant rocm --nic "${NIC_COMPILATION_ARCH}" build develop && \ + echo "WITH_NIXL build complete" >> /app/versions.txt && \ + rm -rf /tmp/ucx /tmp/googletest-1.14.0 /tmp/v1.14.0.tar.gz /tmp/rocm-systems /tmp/rocshmem-build; \ + fi +ENV LD_LIBRARY_PATH="/usr/local/ucx/lib:/usr/local/lib:/usr/local/RIXL/install/lib:${LD_LIBRARY_PATH}" \ + PATH="/usr/local/ucx/bin:${PATH}" + +# ----------------------------------------------------------------------------- +# 5. Cache locations (structural: WHERE the JIT/compile caches live in the image). +# Mount target for the launcher's persistent host JIT cache. +# +# This image ships NO runtime recipe / tuning / platform ENV, matching the +# shared disagg image. Everything run-tunable is applied at launch so the same +# image serves any cluster without a rebuild: +# - K3 serving recipe (VLLM_ROCM_USE_AITER_MLA=0, AITER_SITUV2_A8W4, +# KV_CACHE_MEMORY_BYTES, *_CUDAGRAPH_MODE, *_MORI_BACKEND, ...) +# -> scripts/vllm_dissag/models.yaml, entry "Kimi-K3" +# - ROCm-7.2.3 GPU-RDMA platform env + MoRI fabric tuning +# -> scripts/vllm_dissag/connectors/moriio.env +# The slurm launcher forwards both via `docker -e` (platform env must reach +# PID 1 - PyTorch reads alloc-conf at import). +# ----------------------------------------------------------------------------- +ENV AITER_JIT_DIR=/opt/vllm_cache/aiter_jit \ + VLLM_CACHE_ROOT=/opt/vllm_cache/vllm \ + TRITON_CACHE_DIR=/opt/vllm_cache/triton \ + COMGR_CACHE_DIR=/opt/vllm_cache/comgr + +# ----------------------------------------------------------------------------- +# 6. CRITICAL: scrub build-time MoRI JIT state. The `import mori` verifications +# above compile/lock MoRI EP kernels under /root/.mori/jit on the BUILD host, +# leaving stale .hsaco.lock files. At runtime MoriAll2AllManager finds those +# locks, waits on a build-in-progress whose owner PID is long gone, and +# DEADLOCKS at ep:0 init. A clean image ships /root/.mori empty. +# ----------------------------------------------------------------------------- +RUN rm -rf /root/.mori /tmp/mori_jit_* && mkdir -p /root/.mori && \ + echo "JIT_SCRUBBED: /root/.mori + /tmp/mori_jit_* cleared at build end" >> /app/versions.txt + +RUN cat /app/versions.txt 2>/dev/null | tail -20 || true diff --git a/docs/README.md b/docs/README.md new file mode 100644 index 00000000..991cd956 --- /dev/null +++ b/docs/README.md @@ -0,0 +1,123 @@ +# MAD documentation + +MAD (Model Automation and Dashboarding) is a curated list of AI models that run on various GPU +architectures while tracking performance and generating dashboards for insights. It provides deep +learning recipes for training, inference and deployment on AMD Instinct GPUs. + +You do not run MAD directly. You run it with **madengine**, a command-line tool from +[ROCm/madengine](https://github.com/ROCm/madengine). madengine reads the model definitions in this +repository, builds a Docker image for each model, runs the model's script inside a container (or +submits it to a SLURM cluster), and writes the results to a CSV file. + +This folder is a guide from first contact to writing your own workloads. Start at step 1 of the +learning path and stop when you have what you need. + +## Who should read what + +| You want to | Read | +|---|---| +| Run one model on one GPU host | [Getting started](getting-started.md) | +| Add a model, a Dockerfile or a script | [Adding a model](adding-a-model.md) | +| Understand multinode inference (disaggregated prefill/decode, colocated) | [Multinode overview](multinode-overview.md) | +| Run a multinode card on a SLURM cluster | [Running multinode workloads](multinode-running.md) | +| Change a setting and know which layer wins | [Configuration](configuration.md) | +| Read or compare benchmark results | [Benchmarks and results](benchmarks-and-results.md) | +| Look up every knob of one launcher or one model | [vLLM disaggregated](vllm-disagg.md), [SGLang disaggregated](sglang-disagg.md), [Kimi-K3](kimi-k3.md) | + +## Learning path + +1. **Run a single-node model.** Install madengine, run a model with one command, and learn what + happens during a run and where the results go. + [getting-started.md](getting-started.md) +2. **Add a model.** Write a model card, a Dockerfile and a run script, and report performance in + the format madengine expects. [adding-a-model.md](adding-a-model.md) +3. **Learn the multinode concepts.** Disaggregated prefill/decode versus colocated multinode, + launchers, KV connectors, EP backends and node topology. + [multinode-overview.md](multinode-overview.md) +4. **Run a multinode workload.** Prerequisites, running a card through madengine, `sbatch` or + `salloc`, reading logs and results, and fixing failures. + [multinode-running.md](multinode-running.md) +5. **Configure.** Every configuration layer, its precedence, and where to change a given setting. + [configuration.md](configuration.md) +6. **Benchmark and read results.** Throughput sweeps, needle-in-a-haystack (NIAH), agentic replay, + and the `perf.csv` schema. [benchmarks-and-results.md](benchmarks-and-results.md) +7. **Use the references.** Full per-launcher and per-model references: + [vllm-disagg.md](vllm-disagg.md), [sglang-disagg.md](sglang-disagg.md), + [kimi-k3.md](kimi-k3.md). + +## Page index + +| Page | What it covers | +|---|---| +| [README.md](README.md) | This page: what MAD is, the learning path, the glossary and the index. | +| [getting-started.md](getting-started.md) | Prerequisites, installation, running models, tags, timeouts, debugging, model discovery, build versus run, and where results go. | +| [adding-a-model.md](adding-a-model.md) | Every model-card field, Dockerfile resolution, the GPU architecture build argument, run scripts, the performance reporting contract, and multinode cards. | +| [multinode-overview.md](multinode-overview.md) | Concepts: disaggregated prefill/decode versus colocated multinode, launchers, KV connectors, EP backends, topology, architecture diagrams. | +| [multinode-running.md](multinode-running.md) | Running a multinode card through madengine, `sbatch` or `salloc`; logs, results, failure modes, troubleshooting and offline checks. | +| [configuration.md](configuration.md) | Every configuration layer and its precedence, `models.yaml` recipes, connector environment, `cluster.sh`, madengine presets and `--additional-context` keys, `mad-config.yaml`. | +| [benchmarks-and-results.md](benchmarks-and-results.md) | Throughput sweep, NIAH, agentic replay, the `perf.csv` schema and status semantics. | +| [vllm-disagg.md](vllm-disagg.md) | Full reference for [`scripts/vllm_dissag`](../scripts/vllm_dissag). | +| [sglang-disagg.md](sglang-disagg.md) | Full reference for [`scripts/sglang_disagg`](../scripts/sglang_disagg). | +| [kimi-k3.md](kimi-k3.md) | Kimi-K3 on MI300X and MI355X, colocated and disaggregated. | + +## Blueprints + +These are the supported model families, with the documentation that ships next to each one. + +| Blueprint | Description | Models | +|-----------|-------------|--------| +| [Kimi-K3 inference (vLLM / SGLang / ATOM)](../benchmark/kimi_k3/README.md) | Kimi-K3 (2.8T) day-0 inference on MI350X/MI355X across three frameworks. See also [kimi-k3.md](kimi-k3.md). | [moonshotai/Kimi-K3](https://huggingface.co/moonshotai/Kimi-K3) | +| [xDiT diffusion inference](../benchmark/xdit/README.md) | Diffusion Transformer inference using xDiT | FLUX.1, FLUX.1 Kontext, FLUX.2, FLUX.2 Klein, HunyuanVideo, HunyuanVideo 1.5, LTX-2, Stable Diffusion 3.5, Wan 2.1, Wan 2.2, Z-Image Turbo | +| [JAX MaxText training](../benchmark/jax_maxtext/README.md) | Train LLMs on AMD Instinct GPUs using JAX MaxText | Llama 2 7B/70B, Llama 3/3.1 8B/70B, Llama 3.1 405B, Llama 3.3 70B, DeepSeek-V2-lite 16B, Mixtral-8x7B | +| [vLLM inference](../benchmark/vllm/README.md) | LLM inference with vLLM on AMD Instinct GPUs | DeepSeek-R1, gpt-oss-20b/120b, Kimi-K3, Llama-2-70b, Llama-3.1-8b/405b, Llama-3.3-70b, Llama-4-Scout/Maverick, Mixtral-8x7b/8x22b, Phi-4, Qwen3-8b/32b/30b-a3b/235b-a22b | +| [SGLang inference](../benchmark/sglang/README.md) | LLM inference with SGLang on AMD Instinct GPUs | DeepSeek-R1-Distill-Qwen-32B, Kimi-K3 | +| PyTorch training | Train LLMs on AMD Instinct GPUs using AMD's Primus. Primus notes are in [benchmark/primus/README.md](../benchmark/primus/README.md). | Llama 2/3/3.1/3.2/3.3/4, GPT-OSS 20B/120B, Qwen2/2.5/3, Flux, SDXL, DLRM, and others | +| [PyTorch inference](../benchmark/pytorch_inference/README.md) | Inference recipes for multimodal, video and vision transformer models | Mochi video, Chai-1, CLIP (ViT-B-32), Wan2.1, Janus-Pro-7B, HunyuanVideo | +| Megatron-LM training | Train LLMs on AMD Instinct GPUs using ROCm Megatron-LM | Llama 2 7B/70B, Llama 3/3.1 8B/70B, Llama 3.3 70B, DeepSeek-V2-lite, DeepSeek-V3, Mixtral 8x7B/8x22B, Qwen 2.5 7B/72B | +| MPT-30B training (llm-foundry) | LLM training for Mosaic Pretrained Transformer (MPT) models using llm-foundry | MPT-30B | +| PyTorch PEFT/FSDP fine-tuning | Fine-tuning a Hugging Face model with the LoRA approach and the FSDP strategy | Llama-2-70b-chat-hf | +| [Large EP microbenchmark](../scripts/large-ep-benchmark/README.md) | MoE large expert parallelism with MoRI-EP and DeepEP communication microbenchmarks | No specific models | +| [vLLM disaggregated P/D inference](../scripts/vllm_dissag/README.MD) | Distributed inference with prefill/decode disaggregation in vLLM (Default, MoRI EP, DeepEP). See also [vllm-disagg.md](vllm-disagg.md). | DeepSeek-R1, DeepSeek-V3, DeepSeek-V3-5layer, amd-Llama-3.3-70B-Instruct-FP8-KV, Llama-3.1-405B-Instruct-FP8-KV, gpt-oss-120b | +| [SGLang disaggregated P/D inference](../scripts/sglang_disagg/README.MD) | Distributed inference with prefill/decode disaggregation in SGLang (MoRI IO, Mooncake). See also [sglang-disagg.md](sglang-disagg.md). | Llama-3.1-8B, Qwen3-32B, Llama-3.3-70B-FP8, Llama-3.1-405B-FP8, Mixtral-8x7B, DeepSeek-V3, DeepSeek-R1 | +| [SGLang disaggregated P/D inference with WideEP/LargeEP](../scripts/sglang_disagg/README.MD) | Distributed inference with prefill/decode disaggregation in SGLang with WideEP/LargeEP | DeepSeek-V3, DeepSeek-R1 | +| [KVCache Transfer Bench](../scripts/kvcache_transfer_bench/README.md) | Inter-node transfer benchmark | No specific models | + +The root [README](../README.md) links the PyTorch training, Megatron-LM, MPT-30B and PEFT/FSDP +blueprints to files that are not in this checkout, so those rows have no link here. + +## Glossary + +Terms are listed in the order you meet them. + +| Term | Meaning | +|---|---| +| **madengine** | The CLI from [ROCm/madengine](https://github.com/ROCm/madengine) that discovers, builds, runs and reports MAD models. `pip install -r requirements.txt` in this repository installs it. Its main commands are `discover`, `build`, `run`, `report` and `database`. | +| **Model card** (or **card**) | One JSON object in a `models.json` file that describes one workload: its name, Dockerfile, script, GPU count, tags, arguments and, for multinode work, its launcher, node count and environment. Cards live in `scripts//models.json`. A directory can also generate cards in Python with `get_models_json.py`. See [adding-a-model.md](adding-a-model.md). | +| **Tag** | A label in a card's `tags` list, such as `pyt`, `vllm` or `inference`. `madengine run --tags X` selects every card whose name or tags match `X`. See [getting-started.md](getting-started.md#selecting-models-with-tags). | +| **Recipe** | The per-model serving flags and environment for a multinode workload, kept apart from the card. For vLLM disaggregated serving it is the model's entry in [`scripts/vllm_dissag/models.yaml`](../scripts/vllm_dissag/models.yaml); SGLang has [`scripts/sglang_disagg/models.yaml`](../scripts/sglang_disagg/models.yaml). See [configuration.md](configuration.md). | +| **Launcher** | Two related meanings. (1) The batch script a multinode card runs, for example `run_xPyD_models.slurm` or `run_multinode.slurm`. (2) The value of a card's `distributed.launcher`, which tells madengine how to start the workload (`torchrun`, `vllm`, `sglang`, `slurm_multi` and others). | +| **slurm_multi** | The madengine launcher used by every MAD multinode inference card. madengine writes a wrapper SBATCH script that exports the card's `env_vars`, then runs the card's own `.slurm` script on the head node. That script starts the per-node Docker containers itself with `srun`. The hyphenated `slurm-multi` is accepted as an alias. | +| **Connector** | In disaggregated serving, the component that moves the KV cache from the prefill server to the decode server. vLLM uses `rixl` (NixlConnector) or `moriio` (MoRIIOConnector). SGLang uses MoRI IO or Mooncake as its transfer backend. See [multinode-overview.md](multinode-overview.md). | +| **EP backend** | The all-to-all communication library used for wide expert parallelism (wideEP) in mixture-of-experts models: `mori` (MoRI-EP) or `deepep` (DeepEP). In vLLM each connector pairs with its own backend: `moriio` with `mori`, `rixl` with `deepep`. | +| **xP/yD** | The shape of a disaggregated run: `xP` prefill nodes and `yD` decode nodes. The job needs `xP + yD` nodes. For example `1P/1D` is 2 nodes. | +| **`--additional-context`** | A JSON string passed to `madengine build` or `madengine run` that adds or overrides configuration: `gpu_vendor`, `guest_os`, `docker_env_vars`, `docker_build_arg`, `env_vars`, a `slurm` block, and more. `--additional-context-file` reads the same JSON from a file, and `--additional-context` is merged over it key by key. | +| **Build manifest** | `build_manifest.json`, written by `madengine build` (and by the build half of `madengine run`). It records each built image, the card it was built for, the registry image if one was pushed, and the context. `madengine run --manifest-file build_manifest.json` runs from it without building again. | +| **perf.csv** | The results table. `madengine run` appends one row per model (or one row per result for cards with `multiple_results`), including failed runs as `FAILURE` rows. Change the file name with `-o`. See [benchmarks-and-results.md](benchmarks-and-results.md). | +| **skip_gpu_arch** | A card field listing GPU architectures (comma-separated, for example `gfx942`) the card must not run on. madengine checks it before running, locally or on SLURM, and writes a `SKIPPED` row instead. `--disable-skip-gpu-arch` turns the check off. | +| **MAD_SYSTEM_GPU_ARCHITECTURE** | The host GPU architecture, for example `gfx942` (MI300X) or `gfx950` (MI355X). madengine sets it as an environment variable inside the container, and passes it as a Docker build argument so a Dockerfile can build for one architecture. See [adding-a-model.md](adding-a-model.md#the-mad_system_gpu_architecture-build-argument). | + +## madengine documentation + +madengine has its own documentation in the +[`docs/` folder of ROCm/madengine](https://github.com/ROCm/madengine/tree/main/docs). These pages +are the ones MAD users need most. + +| Page | What it covers | +|---|---| +| [installation.md](https://github.com/ROCm/madengine/blob/main/docs/installation.md) | Installing madengine, setting up the MAD package, testing Docker GPU access for ROCm and CUDA, and fixing import, permission and ROCm-path problems. | +| [usage.md](https://github.com/ROCm/madengine/blob/main/docs/usage.md) | The five commands, the three model discovery methods, the build and run workflows, timeouts, debugging, profiling, reporting and MongoDB upload. | +| [cli-reference.md](https://github.com/ROCm/madengine/blob/main/docs/cli-reference.md) | Every option of `discover`, `build`, `run`, `report` and `database`, with defaults, examples and exit codes. | +| [configuration.md](https://github.com/ROCm/madengine/blob/main/docs/configuration.md) | Every `--additional-context` key: defaults, log error pattern scan, pinned image digests, Docker environment, build arguments, mounts, timeouts, Kubernetes and SLURM blocks, profiling, pre/post scripts, credentials and configuration priority. | +| [launchers.md](https://github.com/ROCm/madengine/blob/main/docs/launchers.md) | Each distributed launcher (`torchrun`, DeepSpeed, Megatron-LM, TorchTitan, Primus, vLLM, SGLang, SGLang disaggregated, `slurm_multi`), with a comparison matrix and troubleshooting. | +| [distributed-config.md](https://github.com/ROCm/madengine/blob/main/docs/distributed-config.md) | How distributed inference configuration splits by scope (the run, the site, the model, the measurement) across madengine `--config`, `cluster.sh`, `models.yaml`, `configs/*.yaml` and `mad-config.yaml`. | +| [deployment.md](https://github.com/ROCm/madengine/blob/main/docs/deployment.md) | Deploying workloads to Kubernetes and SLURM: workflow, configuration examples and troubleshooting. | diff --git a/docs/adding-a-model.md b/docs/adding-a-model.md new file mode 100644 index 00000000..4845962b --- /dev/null +++ b/docs/adding-a-model.md @@ -0,0 +1,678 @@ +# Adding a model + +This page shows how to add a workload to MAD. A workload is three things: a **model card** that +describes it, a **Dockerfile** that builds its image, and a **script** that runs it and prints its +performance. The page explains every card field, how madengine turns a card into a Docker build, +how your script must report results, and how to add a multinode card. + +Read [Getting started](getting-started.md) first. It explains discovery, tags and what happens +during a run. + +## Contents + +- [Where cards live](#where-cards-live) +- [Step 1: choose a name](#step-1-choose-a-name) +- [Step 2: write the card](#step-2-write-the-card) +- [Card field reference](#card-field-reference) +- [Step 3: write the Dockerfile](#step-3-write-the-dockerfile) +- [How the Dockerfile is found](#how-the-dockerfile-is-found) +- [The MAD_SYSTEM_GPU_ARCHITECTURE build argument](#the-mad_system_gpu_architecture-build-argument) +- [Step 4: write the script](#step-4-write-the-script) +- [Report performance](#report-performance) +- [Generating cards with get_models_json.py](#generating-cards-with-get_models_jsonpy) +- [Restricting a card to some GPUs](#restricting-a-card-to-some-gpus) +- [Adding a multinode card](#adding-a-multinode-card) +- [Checklist](#checklist) + +## Where cards live + +Cards used to live in the root `models.json`. They now live next to their scripts, in +`scripts//models.json`. The root `models.json` is an empty list, kept because madengine +requires the file to exist. [`tools/migrate_to_dir_models.py`](../tools/migrate_to_dir_models.py) +did the move; its path rules show how a root card maps to a directory card: + +| Field | Root `models.json` | `scripts//models.json` | +|---|---|---| +| `dockerfile` | `docker/X` | `../../docker/X` | +| `scripts` | `scripts/` or `scripts//` | `.` | +| `scripts` | `scripts//f.sh` | `f.sh` | +| `scripts` | `scripts//...` | `../../scripts//...` | +| `dockercontext` | kept as is, relative to the repository root | kept as is; madengine does not rewrite it | + +Put a new card in the `models.json` of the directory that holds its script. Create the directory +if it is a new family. Remember the discovery rules from +[Getting started](getting-started.md#how-madengine-discovers-models): the card's name gets the +directory as a prefix, its paths are relative to the directory, and a directory cannot have both +`models.json` and `get_models_json.py`. + +## Step 1: choose a name + +Name the workload `{framework}_{project}_{workload}`. Examples: + +- `tf2_huggingface_gpt2` +- `pyt_torchvision_resnet50` +- `ort_onnx_bert` + +MAD's current cards follow the same shape, for example `pyt_vllm_deepseek-r1`, +`pyt_sglang_disagg_mori_io_llama-3.1-8b` and `pyt_ncf_training`. The name must be unique. + +## Step 2: write the card + +The root README gives this example of a card, written for the root `models.json`: + +```json +{ + "name": "tf2_bert_large", + "url": "https://github.com/ROCmSoftwarePlatform/bert", + "dockerfile": "docker/tf2_bert_large", + "scripts": "scripts/tf2_bert_large", + "n_gpus": "4", + "owner": "john.doe@amd.com", + "training_precision": "fp32", + "tags": [ + "per_commit", + "tf2", + "bert", + "fp32" + ], + "args": "" +} +``` + +The same card in `scripts/tf2_bert_large/models.json` uses paths relative to that directory: + +```json +[ + { + "name": "tf2_bert_large", + "url": "https://github.com/ROCmSoftwarePlatform/bert", + "dockerfile": "../../docker/tf2_bert_large", + "scripts": "run.sh", + "n_gpus": "4", + "owner": "john.doe@amd.com", + "training_precision": "fp32", + "tags": ["per_commit", "tf2", "bert", "fp32"], + "args": "" + } +] +``` + +A real single-node card, from [`scripts/ncf/models.json`](../scripts/ncf/models.json): + +```json +{ + "name": "pyt_ncf_training", + "dockerfile": "../../docker/pyt_ncf_training", + "scripts": "run.sh", + "url": "https://github.com/ROCm/DeepLearningExamples", + "data": "", + "n_gpus": "1", + "owner": "mad.support@amd.com", + "training_precision": "fp32", + "multiple_results": "results_ncf.csv", + "tags": ["pyt", "training", "recommendation", "ncf"], + "timeout": -1, + "args": "" +} +``` + +Check the card with `madengine discover --tags --verbose`. It prints the card as madengine +resolved it, with the directory prefix on the name and the paths rewritten. + +## Card field reference + +These are all the fields that appear in MAD's cards, plus the fields madengine accepts that MAD +cards do not use yet. "Required" follows the root README; madengine itself fills defaults for most +fields. + +### Core fields + +| Field | Required | Type | Description | +|---|---|---|---| +| `name` | Yes | string | Unique model identifier. For a directory card madengine prefixes it with the directory, for example `vllm/pyt_vllm_deepseek-r1`. It sets the image name, the log file names and `MAD_MODEL_NAME`. | +| `url` | Yes | string | Repository to clone into the container before the script runs. madengine runs `git clone `, updates its submodules, and records the commit. The model directory is named after the last part of the URL, which should contain only letters, digits, `-` and `_`. Many MAD cards set `""` or omit it; the script then runs in an empty `run_directory`. | +| `dockerfile` | Yes | string | Path prefix of the Dockerfile, without the `.Dockerfile` suffix. See [How the Dockerfile is found](#how-the-dockerfile-is-found). | +| `scripts` | Yes | string | What to run. A `.sh` or `.slurm` file runs with `bash`, a `.py` file with `python3`, and a directory runs its `run.sh`. The file's directory is copied into the model directory before the run. | +| `n_gpus` | Yes | string | Number of GPUs. `"-1"` means all available; madengine converts it to the system GPU count. MAD cards use `"-1"`, `"1"` and `"8"`. | +| `owner` | Yes | string | Contact email. MAD cards use `mad.support@amd.com`. | +| `training_precision` | Yes | string | Precision label such as `fp16` or `fp32`. Inference cards leave it `""`. | +| `tags` | Yes | list of strings | Labels for selection with `--tags`. Include the framework (`pyt`), the family (`vllm`, `sglang_disagg`) and the kind (`inference`, `training`). | +| `args` | No | string | Arguments appended to the script. Extra arguments from a tag (`--tags :key=value`) are appended after these. Example: `"--model_repo deepseek-ai/DeepSeek-R1-0528 --config configs/default.yaml"`. | +| `data` | No | string | Name of a data provider entry in the root [`data.json`](../data.json). MAD defines one entry, `huggingface`, which most inference cards use. `""` means no data step. | +| `timeout` | No | integer | Timeout in seconds for this card, overriding the 7200 s default. `--timeout` on the command line overrides it. `0` or any negative value means no timeout; most MAD cards set `-1`. | +| `multiple_results` | No | string | Name of a CSV file the script writes with one row per result. See [Report performance](#report-performance). | +| `skip_gpu_arch` | No | string | Comma-separated GPU architectures the card must not run on, for example `"gfx942"`. See [Restricting a card to some GPUs](#restricting-a-card-to-some-gpus). | + +### Build fields + +| Field | Description | +|---|---| +| `dockercontext` | Docker build context. The default is `./docker`, so a Dockerfile can only `COPY` files from `docker/`. madengine uses `.` (the repository root) when the `dockerfile` path contains `primus`. Set `dockercontext` to `.` when the Dockerfile copies from elsewhere in the repository. The JAX and Primus card generators set it. madengine does not rewrite it relative to the card's directory. | +| `docker_build_arg` | A JSON object of Docker build arguments for this card only, for example pins like `{"VLLM_REF": ""}`. `docker_build_arg` in `--additional-context` wins for the same key. No MAD card uses it yet. | +| `cred` | Name of a credential entry in `credential.json`. madengine uses it to clone a private `url` and passes its values as build arguments. No MAD card uses it. | + +### Multinode fields + +Only multinode cards use these. See [Adding a multinode card](#adding-a-multinode-card). + +| Field | Description | +|---|---| +| `distributed.launcher` | How madengine starts the workload. Every MAD multinode card uses `slurm_multi`. | +| `distributed.nnodes` | Number of nodes. Must equal `slurm.nodes`. | +| `slurm.nodes` | Number of nodes madengine allocates (`#SBATCH --nodes`). | +| `slurm.gpus_per_node` | GPUs per node. MAD cards use `8`. | +| `slurm.time` | Wall-clock limit. MAD cards use `"24:00:00"`; override it for your partition. | +| `env_vars` | Environment the launcher script starts with: the image, the model, the topology and the benchmark knobs. madengine exports each one in the job script it generates. | + +## Step 3: write the Dockerfile + +Create the Dockerfile in `docker/`. The root README's example: + +```dockerfile +# CONTEXT {'gpu_vendor': 'AMD', 'guest_os': 'UBUNTU'} +FROM rocm/tensorflow:latest + +# Install system dependencies +RUN apt update && apt install -y \ + wget \ + unzip \ + && rm -rf /var/lib/apt/lists/* + +# Install Python dependencies +RUN pip install --no-cache-dir \ + pandas \ + numpy + +# Download model data +RUN URL=https://example.com/model-data.zip && \ + wget --directory-prefix=/data -c $URL && \ + ZIP_NAME=$(basename $URL) && \ + unzip /data/$ZIP_NAME -d /data && \ + rm /data/$ZIP_NAME + +# Set working directory +WORKDIR /workspace +``` + +Conventions MAD's Dockerfiles follow: + +- **Name it `.ubuntu.amd.Dockerfile`**, for example `docker/pyt_vllm.ubuntu.amd.Dockerfile`, + and point the card's `dockerfile` at `docker/` (or `../../docker/` from a + directory card). +- **Start with a `# CONTEXT` line** in the first five lines. madengine reads it to decide whether + the Dockerfile fits the run. See below. +- **Take the base image from `BASE_DOCKER`**, so users can change it without editing the file: + + ```dockerfile + # CONTEXT {'gpu_vendor': 'AMD', 'guest_os': 'UBUNTU'} + ARG BASE_DOCKER=rocm/pytorch + FROM $BASE_DOCKER + ``` + + That is the whole of [`docker/dummy.ubuntu.amd.Dockerfile`](../docker/dummy.ubuntu.amd.Dockerfile). + To build on a different base: + + ```bash + madengine build --tags \ + --additional-context '{"docker_build_arg": {"BASE_DOCKER": "rocm/pytorch:rocm6.1_ubuntu22.04_py3.10"}}' + ``` + +- **Pin sources to commits.** [`docker/vllm_kimi_k3.ubuntu.amd.Dockerfile`](../docker/vllm_kimi_k3.ubuntu.amd.Dockerfile) + pins every source to an immutable commit SHA so the image can be rebuilt to the same bits later. + It also sets `PIP_CONSTRAINT` globally, because pip resolves build dependencies such as + setuptools from PyPI in an isolated environment, and a plain `pip install setuptools==X` does + not reach that environment. A newer setuptools (80 and later) broke the MoRI wheel build that + way while every pinned source was unchanged. + +## How the Dockerfile is found + +The card's `dockerfile` value is a **prefix**, not a file name. madengine lists the files that +start with it and keeps only these two forms: + +- `.Dockerfile` +- `...Dockerfile`, where `` and `` are single words of letters, + digits, `_` or `-` + +The match is exact. Other files that merely share the prefix are ignored. For example, with +`"dockerfile": "../../docker/vllm_disagg_inference"`: + +| File | Built for this card? | +|---|---| +| `docker/vllm_disagg_inference.ubuntu.amd.Dockerfile` | Yes | +| `docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile` | No: `glmv5.1.ubuntu.amd` is not `.` | + +The GLM-5.1 cards select that second file with their own prefix, +`../../docker/vllm_disagg_inference.glmv5.1`. To give a card its own image, give the Dockerfile +its own prefix and point the card at it. + +madengine then reads the `# CONTEXT` line from the first five lines of each file and keeps the +files whose context matches the run. A key in the header must equal the same key in the run +context (`gpu_vendor` defaults to `AMD`, `guest_os` to `UBUNTU`). A Dockerfile without the header +cannot be read this way, and madengine finds no Dockerfile for the card. Each kept file is built +as its own image. + +The image is named `ci-_`, lowercased, with `/` +replaced by `_`. The build log is `_.build.live.log`. + +## The MAD_SYSTEM_GPU_ARCHITECTURE build argument + +`MAD_SYSTEM_GPU_ARCHITECTURE` is the GPU architecture, such as `gfx942` (MI300X) or `gfx950` +(MI355X). madengine sets it in two places: + +- **Inside the container**, as an environment variable, so scripts can adapt. For example + [`scripts/huggingface_bert/run.sh`](../scripts/huggingface_bert/run.sh) picks a batch size + from it. +- **At build time**, as a Docker build argument, so a Dockerfile can compile for one + architecture. A Dockerfile opts in by declaring it: + + ```dockerfile + ARG MAD_SYSTEM_GPU_ARCHITECTURE + ENV HIP_ARCHITECTURES=${MAD_SYSTEM_GPU_ARCHITECTURE} + ``` + + [`docker/pyt_clip_inference.ubuntu.amd.Dockerfile`](../docker/pyt_clip_inference.ubuntu.amd.Dockerfile), + `pyt_janus_pro_inference` and `pyt_wan2.1_inference` do this. + +Where the build value comes from, first match wins: + +1. `docker_build_arg.MAD_SYSTEM_GPU_ARCHITECTURE` in `--additional-context`. +2. `--target-archs` on `madengine build`, which builds one image per listed architecture and + passes each one in turn. The image name gets an `_` suffix. +3. Local detection. `madengine run --tags ...` (build and run together) detects the local GPU and + passes its architecture. A standalone `madengine build` does not detect, because a build host + may have no GPU. + +**A declaration with no default makes the argument required.** When a Dockerfile declares +`ARG MAD_SYSTEM_GPU_ARCHITECTURE` with no value (or an empty one), and none of the three sources +gives one, madengine prints a warning that names the cards and Dockerfiles affected and suggests: + +```bash +--additional-context '{"docker_build_arg": {"MAD_SYSTEM_GPU_ARCHITECTURE": "gfx942"}}' +``` + +A declaration with a non-empty default (`ARG MAD_SYSTEM_GPU_ARCHITECTURE=gfx942`) does not need +it. + +**Example: an image built for exactly one architecture.** +[`docker/vllm_kimi_k3.ubuntu.amd.Dockerfile`](../docker/vllm_kimi_k3.ubuntu.amd.Dockerfile) is the +image for the Kimi-K3 vLLM multinode cards, on MI300X (`gfx942`) and MI355X (`gfx950`), colocated +and disaggregated. It declares the argument with no default, on purpose: an image built for the +wrong architecture fails at runtime, so the build refuses to guess. + +```dockerfile +ARG MAD_SYSTEM_GPU_ARCHITECTURE +ENV K3_GFX_ARCH=${MAD_SYSTEM_GPU_ARCHITECTURE} +RUN case "${K3_GFX_ARCH}" in \ + gfx942|gfx950) echo "Kimi-K3 image for ${K3_GFX_ARCH}" ;; \ + *) echo "MAD_SYSTEM_GPU_ARCHITECTURE must be gfx942 or gfx950, got '${K3_GFX_ARCH}'" >&2; exit 1 ;; \ + esac +``` + +Everything architecture-specific in the build follows that value: MoRI's JIT target, the vLLM +compile, and, with `WITH_NIXL=1`, rocSHMEM and DeepEP. The value is copied into `K3_GFX_ARCH` so +that madengine's Dockerfile architecture check, which parses `GFX_COMPILATION_ARCH`, +`PYTORCH_ROCM_ARCH` and `GPU_ARCHS`, does not read it as a fixed architecture. Build it through +madengine with the architecture of the nodes it will run on: + +```bash +madengine build --tags pyt_vllm_kimi-k3_mi300x_pp2xtp8 --registry \ + --additional-context '{"docker_build_arg": {"MAD_SYSTEM_GPU_ARCHITECTURE": "gfx942"}}' +``` + +or by hand, from the repository root: + +```bash +docker build -f docker/vllm_kimi_k3.ubuntu.amd.Dockerfile \ + --build-arg MAD_SYSTEM_GPU_ARCHITECTURE=gfx942 -t /vllm-kimi-k3:gfx942 . +``` + +Details of that image are in [kimi-k3.md](kimi-k3.md). + +## Step 4: write the script + +Create the script in your `scripts//`. The root README's example `run.sh`: + +```bash +#!/bin/bash +set -e + +# Model configuration +MODEL_CONFIG_DIR=/data/model_config +BATCH_SIZE=2 +SEQUENCE_LENGTH=512 +TRAIN_STEPS=100 +WARMUP_STEPS=10 +LEARNING_RATE=1e-4 + +# Prepare data +echo "Preparing training data..." +python3 prepare_data.py \ + --config_dir=$MODEL_CONFIG_DIR \ + --batch_size=$BATCH_SIZE \ + --seq_length=$SEQUENCE_LENGTH + +# Train model +echo "Starting model training..." +python3 train_model.py \ + --config_dir=$MODEL_CONFIG_DIR \ + --batch_size=$BATCH_SIZE \ + --max_seq_length=$SEQUENCE_LENGTH \ + --num_train_steps=$TRAIN_STEPS \ + --num_warmup_steps=$WARMUP_STEPS \ + --learning_rate=$LEARNING_RATE \ + 2>&1 | tee training.log + +# Report performance +echo "Generating performance metrics..." +python3 report_metrics.py +``` + +What the script can rely on at run time: + +- **Working directory.** The script runs inside the model directory: the cloned repository if the + card has a `url`, or `run_directory` otherwise. The files from the card's script directory are + copied into it, so helpers next to `run.sh` are in the current directory. +- **Paths.** The MAD checkout is mounted at `/myworkspace` and is the container's working + directory, so the model directory is `/myworkspace/` and `../` is the MAD root. +- **Arguments.** `$@` holds the card's `args` plus any extra arguments from the tag. +- **Environment.** `MAD_SYSTEM_GPU_ARCHITECTURE`, `MAD_RUNTIME_NGPUS`, `MAD_SYSTEM_NGPUS`, + `MAD_MODEL_NAME`, `MAD_OUTPUT_CSV` and anything passed with `docker_env_vars`. See + [Getting started](getting-started.md#environment-variables-inside-the-container). + +## Report performance + +madengine does not know what your model measures. Your script tells it, in one of two ways. + +### Single result + +Print one line of this form to standard output: + +```python +print(f"performance: {throughput} examples/sec") +``` + +madengine searches the run log for `performance:`, then a number, then a metric name. The number +can be an integer, a decimal or scientific notation (`1.23e+4`). A unit suffix such as `/s` and a +comma may sit between the number and the metric, in either order, so all of these parse: + +```text +performance: 14164 samples_per_second +performance: 14164/s, samples_per_second +performance: 14164, /s samples_per_second +``` + +If there is no such line, madengine falls back to the Hugging Face Trainer's +`train_samples_per_second` value, with the metric `samples_per_second`. If neither is found, the +run has no performance and its status is `FAILURE`. +[`scripts/huggingface_bert/run.sh`](../scripts/huggingface_bert/run.sh) ends this way: + +```bash +set +x +echo "performance: $performance samples_per_second" +``` + +The script turns off shell tracing (`set +x`) before it prints. With tracing on, bash also logs +the `echo` command itself (`+ echo 'performance: ...'`), and madengine takes the first match in +the log. + +### Multiple results + +When one run produces several results, for example one per batch size or per precision, write a +CSV file and name it in the card's `multiple_results` field. madengine also passes the name to the +container as `MAD_OUTPUT_CSV`. + +The CSV must have the columns `model`, `performance` and `metric`: + +```csv +model,performance,metric +model_1,156.7,examples/sec +model_2,89.3,tokens/sec +``` + +The root README names the first column `models`. madengine's results code requires `model`, and +MAD's scripts write `model` (see [`scripts/dummy/run_multi.sh`](../scripts/dummy/run_multi.sh) and +[`scripts/ncf/get_ncf_model_metrics.py`](../scripts/ncf/get_ncf_model_metrics.py)). Use `model`. + +How madengine reads the file: + +- It looks for the file in the MAD root first, then in the model directory, and copies it out of + the model directory when it is there. +- A missing `performance` column, or a `performance` column empty in every row, means the run has + no performance. +- Each row becomes one row in `perf.csv`. The row's model name is `_`. +- Extra columns in your CSV (for example batch size, input length, tensor-parallel size) are + carried into `perf.csv`. +- Each row gets its own status: `SUCCESS` when its `performance` is set, `FAILURE` when it is + empty. + +The dummy card is the smallest working example. Its script writes four rows and copies the file +to the MAD root: + +```bash +echo "model,performance,metric +1,$RANDOM,samples_per_sec +2,$RANDOM,samples_per_sec +3,$RANDOM,samples_per_sec +4,$RANDOM,samples_per_sec" >>perf_dummy.csv + +cp perf_dummy.csv ../ +``` + +The full `perf.csv` schema and status rules are in +[Benchmarks and results](benchmarks-and-results.md). + +## Generating cards with get_models_json.py + +When a family has many similar cards, generate them in Python instead of listing them. Put a +`get_models_json.py` in `scripts//` (and no `models.json` in the same directory). It must +define `list_models()`, which returns a list of `CustomModel` objects: + +```python +from madengine.utils.discover_models import CustomModel + +def list_models(): + return [ + CustomModel( + name="default", + dockerfile="../../docker/primus", + dockercontext=".", + scripts="run.sh", + n_gpus="-1", + owner="mad.support@amd.com", + tags=["training", "primus", "megatron", "pretrain"], + args="", + ) + ] +``` + +`CustomModel` has the core fields (`name`, `dockerfile`, `dockercontext`, `scripts`, `url`, +`cred`, `owner`, `data`, `n_gpus`, `timeout`, `training_precision`, `tags`, `args`, +`multiple_results`, `skip_gpu_arch`). Its defaults are `n_gpus="-1"` and `timeout=7200`. As with +`models.json`, madengine prefixes the name with the directory and resolves `dockerfile` and +`scripts` relative to it. + +MAD has three generators. [`scripts/primus_train/get_models_json.py`](../scripts/primus_train/get_models_json.py) +makes one card per Primus example config and passes the config path in `args`. +`scripts/jax-maxtext/get_models_json.py` and `scripts/jax-maxdiffusion/get_models_json.py` do the +same for JAX. + +## Restricting a card to some GPUs + +`skip_gpu_arch` lists the architectures a card must not run on, separated by commas +(`"gfx942"`, or `"gfx942,gfx950"`). madengine applies it before running: + +- **Locally**, it compares the list with the host GPU. +- **On SLURM**, it probes the compute nodes' architecture before submitting, and drops the card + from the job. If the probe cannot tell, madengine submits anyway with a warning; set + `slurm.gpu_arch` in `--additional-context` to enforce the check. + +A skipped card gets a `SKIPPED` row in `perf.csv`. `--disable-skip-gpu-arch` turns the check off. + +Multinode cards can also declare `GPU_ARCHS`, the architectures they support. The launcher reads it +on the allocation and refuses the wrong nodes. For `scripts/vllm_dissag` it lives in +`models.yaml` under the card's `MODEL_NAME`; elsewhere it is in the card's `env_vars`, and a card's +own value wins. The two fields must agree, or one path skips a card that the other runs. +[`scripts/common/check_gpu_arch_declarations.py`](../scripts/common/check_gpu_arch_declarations.py) +checks, for every card that declares `GPU_ARCHS`, that: + +1. no architecture it supports is also in `skip_gpu_arch`; +2. every known architecture (`gfx942`, `gfx950`) it does not support is in `skip_gpu_arch`, so + madengine skips it without spending an allocation to find out. + +```bash +python3 scripts/common/check_gpu_arch_declarations.py +``` + +## Adding a multinode card + +A multinode card runs one workload across several nodes of a SLURM cluster. Read +[Multinode overview](multinode-overview.md) first for the concepts, and +[Running multinode workloads](multinode-running.md) for how to run one. + +### What is different + +Every MAD multinode card uses the `slurm_multi` launcher. With it, madengine does not start a +container itself. Instead it: + +1. writes a wrapper SBATCH script (`slurm_results/madengine_.sh`) that exports the card's + `env_vars`; +2. pulls the image on every node in parallel when it is a registry image; +3. submits the wrapper with `sbatch`, or runs it with `bash` when you are already inside a + `salloc` allocation; +4. runs the card's own `.slurm` script on the head node, with the card's `args`. That script + starts the per-node containers itself with `srun`; +5. collects the `perf.csv` the script writes. + +The card's `.slurm` script is the **launcher**. The card's `env_vars` are the launcher's contract: +the same variables work when you export them and submit the launcher with `sbatch` yourself. + +### An example card + +From [`scripts/vllm_dissag/models.json`](../scripts/vllm_dissag/models.json): + +```json +{ + "name": "pyt_vllm_disagg_nixl_deepseek-v3", + "dockerfile": "../../docker/vllm_disagg_inference", + "scripts": "run_xPyD_models.slurm", + "url": "", + "data": "huggingface", + "n_gpus": "-1", + "owner": "mad.support@amd.com", + "training_precision": "", + "tags": ["pyt", "vllm", "vllm_disagg", "nixl", "inference"], + "timeout": -1, + "distributed": { + "launcher": "slurm_multi", + "nnodes": 2 + }, + "env_vars": { + "DOCKER_IMAGE_NAME": "", + "MODEL_NAME": "DeepSeek-V3", + "xP": "1", + "yD": "1", + "WIDE_EP": "1", + "RUN_MORI": "0", + "RUN_DEEPEP": "0", + "BENCHMARK_COMBINATIONS": "1024/1024" + }, + "args": "-N 2 -n 2", + "slurm": { + "nodes": 2, + "gpus_per_node": 8, + "time": "24:00:00" + } +} +``` + +### Rules for a multinode card + +- **`slurm.nodes` and `distributed.nnodes` must agree.** madengine sizes the allocation from + `slurm.nodes` only (it becomes `#SBATCH --nodes`, default 1), and the launcher then reads + `SLURM_NNODES` to find its nodes. `distributed.nnodes` carries the same number for launcher + detection but does not size the allocation. If they differ, the job gets the wrong number of + nodes. `--additional-context '{"slurm": {"nodes": N}}'` overrides the card. +- **Put the node count in `args` too.** MAD cards set `"args": "-N -n "`, the same + count you give `sbatch` when you submit the launcher by hand. +- **For disaggregated cards, nodes = `xP + yD`.** `xP` is the number of prefill nodes and `yD` the + number of decode nodes, so `xP=1, yD=1` needs 2 nodes and `xP=2, yD=2` needs 4. +- **Use `"n_gpus": "-1"`** and set GPUs per node in `slurm.gpus_per_node`. +- **Leave allocation defaults out of the card.** Partition, GPUs per node and exclusivity come from + madengine's SLURM presets: partition `amd-rccl`, 8 GPUs per node, exclusive. Users override them + in `--additional-context`. The card's `slurm.time` of `24:00:00` exceeds many partitions' + limits, so users usually pass their own `slurm.time`. +- **Declare GPU support** with `skip_gpu_arch`, and with `GPU_ARCHS` if the launcher should refuse + wrong nodes. See [Restricting a card to some GPUs](#restricting-a-card-to-some-gpus). + +### The `DOCKER_IMAGE_NAME` placeholder + +A multinode launcher runs the image named by `DOCKER_IMAGE_NAME` on every node, so each node must be +able to pull it. MAD cannot know your registry, so every multinode card ships with: + +```json +"DOCKER_IMAGE_NAME": "" +``` + +This is a fill-me-in marker, not an image. Supply the real image in one of these ways: + +- `madengine build --tags --registry ` builds the image, pushes it, and records + the pushed name for the run. +- `madengine build --tags --use-image /:` uses an image you already + pushed. +- Under plain `sbatch`, `export DOCKER_IMAGE_NAME=/:` before submitting. + +A `slurm_multi` build with neither `--registry` nor `--use-image` (nor `--build-on-compute`) takes +the card's `DOCKER_IMAGE_NAME` as the image, and would take the placeholder. When no card image is +set, it stops with `slurm_multi launcher requires --registry or --use-image`. madengine's SLURM +deployment does not accept a name that starts with `<` as an image, and it refuses a local +`ci-...` image for a run of more than one node, because the other nodes cannot pull it. + +Keep the placeholder in cards you contribute. Do not commit a private image name. + +### Recipes and environment + +Keep the card small. Model-specific serve flags belong in the **recipe**, not in the card: + +- vLLM disaggregated: an entry in [`scripts/vllm_dissag/models.yaml`](../scripts/vllm_dissag/models.yaml) + keyed by `MODEL_NAME`, plus the name in `VALID_MODELS` in `run_xPyD_models.slurm`. See + [vllm-disagg.md](vllm-disagg.md). +- SGLang disaggregated: [`scripts/sglang_disagg/models.yaml`](../scripts/sglang_disagg/models.yaml). + See [sglang-disagg.md](sglang-disagg.md). +- vLLM colocated multinode: the card's `env_vars` (for example `TP_SIZE`, `PP_SIZE`, + `COLOCATED_EXTRA_ARGS`) and [`scripts/vllm_multinode/mad-config.kimi-k3.yaml`](../scripts/vllm_multinode/mad-config.kimi-k3.yaml). + See [kimi-k3.md](kimi-k3.md). + +The launcher's environment is layered, weakest first: `scripts/common/cluster.sh` defaults, the +recipe in `models.yaml`, the card's `env_vars`, then what the user sets (madengine `env_vars` in +`--additional-context`, or the exported environment under `sbatch`). The full rules and every knob +are in [configuration.md](configuration.md). + +### Offline checks + +These need no GPUs. Run them before you open a pull request: + +```bash +bash scripts/vllm_dissag/tests/argv_assert.sh # the serve argv per connector and mode +python3 scripts/common/check_srun_quotes.py # no apostrophe truncates an srun body +python3 scripts/common/check_gpu_arch_declarations.py # cards and recipes agree on GPUs +``` + +A launcher can also be previewed without starting servers: `DRY_RUN=1` prints each node's server +command. See [Running multinode workloads](multinode-running.md). + +## Checklist + +Before you open a pull request: + +1. The card is in `scripts//models.json`, with a unique name and useful tags. +2. `madengine discover --tags --verbose` shows the card with the paths you expect. +3. The Dockerfile is `docker/.ubuntu.amd.Dockerfile`, starts with a `# CONTEXT` line, and + is found for the card (only `.Dockerfile` and `...Dockerfile` + match). +4. If the Dockerfile declares `ARG MAD_SYSTEM_GPU_ARCHITECTURE` with no default, you built it with + the argument set. +5. The script prints `performance: `, or writes the `multiple_results` CSV with + `model`, `performance` and `metric` columns. +6. `madengine run --tags --live-output` finishes with `SUCCESS` and a row in `perf.csv`. +7. For a multinode card: `slurm.nodes` equals `distributed.nnodes`, `DOCKER_IMAGE_NAME` is + ``, `skip_gpu_arch` and `GPU_ARCHS` agree, and the offline checks pass. diff --git a/docs/benchmarks-and-results.md b/docs/benchmarks-and-results.md new file mode 100644 index 00000000..1481b062 --- /dev/null +++ b/docs/benchmarks-and-results.md @@ -0,0 +1,578 @@ +# Benchmarks and results + +A multinode job starts its servers, then runs one benchmark against them from rank 0 +and writes `perf.csv`. This page explains each benchmark, every knob, the files they +write, and how to read a result. For how to launch the job, see +[multinode-running.md](multinode-running.md). + +## Contents + +- [Choosing a benchmark](#choosing-a-benchmark) +- [Throughput sweep](#throughput-sweep) +- [Long context](#long-context) +- [NIAH: long-context retrieval](#niah-long-context-retrieval) +- [Agentic replay](#agentic-replay) +- [perf.csv](#perfcsv) +- [The CONCURRENCY log and CSV](#the-concurrency-log-and-csv) +- [How madengine collects the CSV](#how-madengine-collects-the-csv) +- [Reading a result](#reading-a-result) +- [What makes a run pass](#what-makes-a-run-pass) + +## Choosing a benchmark + +`BENCHMARK_SCRIPT` picks the benchmark. The launcher maps it to a script in +[`scripts/vllm_dissag/`](../scripts/vllm_dissag/) and runs it on rank 0 once the +servers and the router are ready. + +| `BENCHMARK_SCRIPT` | Script | Measures | vLLM disagg | vLLM colocated | SGLang disagg | +|---|---|---|---|---|---| +| `sweep` (default) | `benchmark_xPyD.sh` | Throughput per (ISL, OSL, concurrency) cell | yes | yes | yes (its own `benchmark_xPyD.sh`) | +| `long_context` | `benchmark_long_context.sh` | Steady-state throughput and latency with per-shape warmup, concurrency 1 first | yes | yes | no | +| `niah` | `benchmark_niah.sh` | Retrieval accuracy: needles found out of 10 per context size | yes | yes | no | +| `agentic` | `benchmark_agentic.sh` | Replay of agentic coding traces; throughput, latency and prefix-cache hit rate | yes | no | yes | + +Any other value stops the launcher with an error. The benchmark talks to the router or +proxy on `127.0.0.1:$BENCHMARK_PORT` (the colocated head's `SERVE_PORT`), so requests +take the same path a client's would. + +The Kimi-K3 cards default to `niah`. The agentic cards set `agentic`. Every other card +runs the sweep. + +## Throughput sweep + +[`benchmark_xPyD.sh`](../scripts/vllm_dissag/benchmark_xPyD.sh) runs +`vllm bench serve` with the `random` dataset at every combination of input/output +length and concurrency. + +### What it does + +1. Waits 10 seconds. +2. Runs one **global warmup**: `WARMUP_PROMPTS` prompts at concurrency `WARMUP_CON`, + ISL `WARMUP_ISL`, OSL `WARMUP_OSL`. The warmup is written to the log before the + `Running the benchserving script for iter: 1` marker, so it never becomes a + result row. +3. For each iteration (`BENCHMARK_ITR`), for each `ISL/OSL` pair in + `BENCHMARK_COMBINATIONS`: + - If `SHAPE_WARMUP=1`, runs a **per-shape warmup** at the real ISL/OSL and low + concurrency, logged to `..._SHAPEWARMUP.log`. + - For each concurrency in `BENCHMARK_CON`, runs one **cell**: `2 x concurrency` + prompts (at least 16), request rate `inf`, `--ignore-eos`, `--max-concurrency` + set to the concurrency. It prints `[RUNNING] prompts isl osl con + (timeout s)` first, then waits 10 seconds after the cell. +4. Parses the log into `..._CONCURRENCY.csv` and `perf.csv` with `parse_to_csv.py`. + +### Per-cell timeout and `[STALL]` + +Each cell runs under `timeout`. The limit scales with the tokens per request: + +``` +timeout = max(STEP_TIMEOUT, STEP_TIMEOUT * (isl + osl) / 2048) +``` + +With the default `STEP_TIMEOUT=1800`, a 1024/1024 cell gets 1800 s and an 8192/1024 +cell gets 8100 s. A cell that hits the limit (exit code 124) prints no result block. +The script writes instead: + +``` +[STALL] isl= osl= con= timed out after s +``` + +to both `..._CONCURRENCY.log` and `..._STALLS.log`, and moves on to the next cell. The +parser turns that line into a `FAILURE` row, so a sweep that stopped answering cannot +pass by omission. + +### Knobs + +| Variable | Default | Meaning | +|---|---|---| +| `BENCHMARK_ITR` | `1` | Iterations of the whole sweep. The CSV keeps the maximum throughput per cell across iterations. | +| `BENCHMARK_CON` | `8 16 32 64 128 256 512` | Concurrency levels, space-separated. | +| `BENCHMARK_COMBINATIONS` | `1024/1024 8192/1024 1024/8192` | `ISL/OSL` pairs, space-separated. Most disagg cards set `1024/1024`. | +| `STEP_TIMEOUT` | `1800` | Base per-cell timeout in seconds, scaled as above. | +| `WARMUP_CON` | `1` | Global warmup concurrency. | +| `WARMUP_PROMPTS` | `16` | Global warmup prompts. | +| `WARMUP_ISL` / `WARMUP_OSL` | `32` / `32` | Global warmup lengths. | +| `SHAPE_WARMUP` | `0` | `1` adds a warmup at each real shape before its cells. | +| `SHAPE_WARMUP_CON` | `4` | Per-shape warmup concurrency. | +| `SHAPE_WARMUP_PROMPTS` | `8` | Per-shape warmup prompts. | +| `SHAPE_WARMUP_TIMEOUT` | `2400` | Per-shape warmup timeout in seconds. | +| `BENCHMARK_PORT` | set by the launcher | Router or proxy port. | + +Why per-shape warmup is off by default: the global warmup is ISL=OSL=32 at concurrency +1, so it never exercises a shape's prefill path, its Triton or AITER kernel variants, or +the decode cudagraph batch sizes. But every model shares this sweep, and an A/B at +1024/1024 concurrency 8 measured it neutral, so recipes validated without it are not +shifted for no gain. GLM-5.1 opts in through its `models.yaml` `env:` block, because +its published latency numbers were measured that way. + +Example: a short sweep of two concurrencies at one shape: + +```bash +export BENCHMARK_CON="8 64" BENCHMARK_COMBINATIONS="1024/1024" BENCHMARK_ITR=1 +``` + +## Long context + +[`benchmark_long_context.sh`](../scripts/vllm_dissag/benchmark_long_context.sh) is a +steady-state serving benchmark. It differs from the sweep in three ways: + +- **Per-shape warmup on every cell.** Each (ISL/OSL, concurrency) cell sends + `--num-warmups` requests of the same shape first and discards them, so the measured + result is not contaminated by first-hit JIT, cudagraph, kernel-autotune or + connector-handshake costs. +- **Concurrency 1 first.** Concurrency 1 is the primary latency metric, so the default + list starts with it, and it is measured cleanly warmed. The order follows + `BENCHMARK_CON` as given. +- **Metrics:** total throughput per GPU, TTFT, ITL and TPOT. + +| Variable | Default | Meaning | +|---|---|---| +| `BENCHMARK_CON` | `1 4 8` | Concurrency list. | +| `BENCHMARK_COMBINATIONS` | `1024/1024` | `ISL/OSL` list. | +| `WARMUPS` | `2` | `--num-warmups` per cell. | +| `NUM_PROMPTS_FACTOR` | `4` | Measured prompts per cell = factor x concurrency (at least 16). The sweep uses 2 x concurrency. | +| `STEP_TIMEOUT` | `2400` | Base per-cell timeout, scaled by tokens as in the sweep. Timeouts print `[STALL]`. | +| `GPUS_TOTAL` | `max(xP, yD) x GPUS_PER_NODE` | GPU count used in the log header. | + +It writes `benchmark_long_context__