From 56f83c729868ebce530b452153e8612ec394d99f Mon Sep 17 00:00:00 2001 From: raviguptaamd Date: Wed, 8 Jul 2026 06:42:53 +0000 Subject: [PATCH 01/41] vllm_dissag: GLM-5.1-FP8 (MLA+DSA) MoRI-EP WideEP disaggregated enablement Adds GLM-5.1-FP8 (GlmMoeDsaForCausalLM = MLA + DeepSeek Sparse Attention) to the MoRI-EP WideEP disaggregated serving path, stacked on the #171 unified launcher. Fully isolated from DeepSeek-V3/R1: GLM gets its own image + a MODEL_NAME-gated runtime path, so existing models are byte-identical to develop. Defects fixed (validated 1P/1D EP8 + 2P/2D EP16, NIAH 2k-35k = 10/10, no crash): - Long-context accuracy collapse: vLLM #47766 cache-key fix keeps the persistent sparse-MLA kernel ON (keys metadata on per-request context+query len). - 8k disagg prefill crash: DSA adds a 2nd (indexer) KV cache per layer that the single-geometry MoRIIO connector never transferred; paired + shipped prefill-> decode. Plus DSA invalid-token kernel fix (#45324) and shik-latest DP-notify. Changes: - docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile: NEW per-model image (raviguptaamd/vllm glm5.1-dsa-wideEP_on_shik_latest + aiter e03fa6040 + mori 42e895472b08 + router). The base vllm_disagg_inference Dockerfile (DSV3/R1) is left untouched. Future models add their own Dockerfile the same way. - models.json: card pyt_vllm_disagg_mori_glm-5.1-fp8 (GLM_SKIP_PATCHERS=1: image carries the DSA fixes in-source). - models.yaml: GLM-5.1-FP8 recipe (block=1, AITER MLA on, eager, mori backends). DeepSeek-V3 dp: caps (--max-num-seqs 64 --max-model-len 32768) to bound the newer base's decode logits workspace (isolated to the DSV3 entry). - connectors/moriio.sh: MODEL_NAME-gated GLM DSA runtime patchers (pure no-op for other models); GLM_SKIP_PATCHERS switch for baked-fix images. - 9 idempotent, anchor-based, self-skipping GLM DSA patcher scripts. KNOWN OPEN DEFECT (future work): 4P/4D EP32 emits corrupted tokens at all context lengths (suspect moriep all-to-all combine at scale); use 1P/1D and 2P/2D. Co-Authored-By: Claude --- ...gg_inference.glmv5.1.ubuntu.amd.Dockerfile | 335 ++++++++++++++++++ models.json | 32 ++ .../apply_glm_aiter_sampling_oob_fix.py | 148 ++++++++ .../apply_glm_dsa_indexer_warmup_fix.py | 251 +++++++++++++ .../vllm_dissag/apply_glm_dsa_kernel_fix.py | 85 +++++ .../apply_glm_dsa_moriio_dualkv_fix.py | 176 +++++++++ .../apply_glm_dsa_moriio_engine_fix.py | 116 ++++++ .../apply_glm_dsa_moriio_gate_fix.py | 133 +++++++ .../apply_glm_dsa_moriio_instrument.py | 92 +++++ ...pply_glm_dsa_persistent_kernel_gate_fix.py | 129 +++++++ .../apply_glm_moriio_abort_guard_fix.py | 98 +++++ scripts/vllm_dissag/connectors/moriio.sh | 100 +++++- scripts/vllm_dissag/keepalive_bench.sh | 18 + scripts/vllm_dissag/models.yaml | 102 +++++- scripts/vllm_dissag/run_xPyD_models.slurm | 32 +- scripts/vllm_dissag/vllm_disagg.sh | 28 +- 16 files changed, 1857 insertions(+), 18 deletions(-) create mode 100644 docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile create mode 100644 scripts/vllm_dissag/apply_glm_aiter_sampling_oob_fix.py create mode 100755 scripts/vllm_dissag/apply_glm_dsa_indexer_warmup_fix.py create mode 100755 scripts/vllm_dissag/apply_glm_dsa_kernel_fix.py create mode 100755 scripts/vllm_dissag/apply_glm_dsa_moriio_dualkv_fix.py create mode 100755 scripts/vllm_dissag/apply_glm_dsa_moriio_engine_fix.py create mode 100755 scripts/vllm_dissag/apply_glm_dsa_moriio_gate_fix.py create mode 100755 scripts/vllm_dissag/apply_glm_dsa_moriio_instrument.py create mode 100644 scripts/vllm_dissag/apply_glm_dsa_persistent_kernel_gate_fix.py create mode 100644 scripts/vllm_dissag/apply_glm_moriio_abort_guard_fix.py create mode 100755 scripts/vllm_dissag/keepalive_bench.sh diff --git a/docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile b/docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile new file mode 100644 index 00000000..f61a8819 --- /dev/null +++ b/docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile @@ -0,0 +1,335 @@ +# CONTEXT {'gpu_vendor': 'AMD', 'guest_os': 'UBUNTU'} +############################################################################### +# +# MIT License +# +# Copyright (c) 2025 Advanced Micro Devices, Inc. +# +# Permission is hereby granted, free of charge, to any person obtaining a copy +# of this software and associated documentation files (the "Software"), to deal +# in the Software without restriction, including without limitation the rights +# to use, copy, modify, merge, publish, distribute, sublicense, and/or sell +# copies of the Software, and to permit persons to whom the Software is +# furnished to do so, subject to the following conditions: +# +# The above copyright notice and this permission notice shall be included in all +# copies or substantial portions of the Software. +# +# THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR +# IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, +# FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE +# AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER +# LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, +# OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE +# SOFTWARE. +# +################################################################################# +# ============================================================================= +# vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile +# GLM-5.1-FP8 (MLA + DeepSeek Sparse Attention) MoRI-EP WideEP disagg image. +# PER-MODEL image, isolated from the base vllm_disagg_inference Dockerfile +# (which stays pinned to the DeepSeek-V3 / R1 stack). This split lets each model +# pin its own vLLM/AITER/MoRI without disturbing the others -- add a new +# vllm_disagg_inference..ubuntu.amd.Dockerfile per future model +# (e.g. Kimi-2.6) rather than repinning the shared DSV3 image. +# +# ALL connectors in one image: moriio (TP + MoRI-EP wideEP) + rixl (NIXL TP + +# DeepEP wideEP). = the fullsource MoRI stack, plus a UCX/RIXL/rocSHMEM/DeepEP +# transport layer gated by --build-arg WITH_NIXL (default 1 = everything). +# +# docker build -f docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile \ +# -t /vllm-disagg:glmv5.1 . +# export DOCKER_IMAGE_NAME=/vllm-disagg:glmv5.1 +# +# WITH_NIXL=1 (default) => builds UCX + RIXL(+nixlbench) + rocSHMEM + DeepEP from +# source, so all four connector combos (moriio TP/wideEP, rixl NIXL TP, DeepEP +# wideEP) are present (~+30-45 min build vs WITH_NIXL=0). +# WITH_NIXL=0 => MoRI-EP only (moriio TP/wideEP + deepep-from-base); lean, faster. +# +# STATUS (GLM-5.1-FP8 on this stack): 1P/1D EP8 + 2P/2D EP16 NIAH 2k-35k = 10/10, +# no crash; long-context accuracy fixed via vLLM #47766 (persistent sparse-MLA kept +# ON). 4P/4D EP32 is a KNOWN OPEN DEFECT: token corruption at ALL context lengths +# (garbage output even at 2k), distinct from the long-context bug; prime suspect is +# the moriep all-to-all combine at EP32 scale -> deferred to future work. Use 1P/1D +# and 2P/2D only. (BASE_IMAGE is a gated nightly; override --build-arg BASE_IMAGE=...) +# ============================================================================= +# Reconstructs the validated v1.2.1 (mori121) runtime stack by applying the recipe's +# component pins ON TOP of the open ROCm vLLM ci_base, cloning each source from +# public Git (no local build-contexts). Mirrors dist-inf-cookbook +# Dockerfile.vllm.mori121_shareable: +# +# - BASE: rocm/vllm-dev:ci_base-0fcd9b99... (open ROCm 7.2 / cp312 CI base). +# - MoRI -> built from ROCm/MoRI @ v1.2.1 (BUILD_UMBP=OFF). +# - AITER -> STOCK ROCm/aiter @ e03fa6040 compiled from source + flydsl 0.1.7-0.1.9; +# stale JIT wiped. (#47766 keeps persistent MLA ON -> aiter native gqa64 fold.) +# - vLLM -> COMPILED from shikamd123/vllm @ +# vllm_2p2d_wide-ep_write_shikpate_test_06_29_customer (Wide-EP multi-pod PD, the +# connector/router reference for the 2P2D DP=EP=16 topology). Full compile: it is +# a different commit than the base's, so a .py-only overlay would be ABI-mismatched. +# - RDMA fix (expandable_segments:False x2 + HSA_ENABLE_IPC_MODE_LEGACY=0) is NOT baked +# here — it lives in scripts/vllm_dissag/connectors/.env and the launcher +# forwards it via docker -e. ROCm 7.2.3 cannot dmabuf-export VMM memory, else MoRI +# RegisterRdmaMemoryRegion EFAULTs (errno 14) on the first disagg WRITE. +# - vllm-router (vllm-project/router PR#181 = DP-rank round-robin + 2P2D KV-notify +# dpfix) built in -> no external router binary needed. +# - validated recipe knobs baked as ENV. The MoRIIO disagg fixes (#39276 notify, +# #41751 LL split, DP-rank hash-failsafe) are native in this vLLM (no runtime patcher). +# +# Build context = repo root: +# docker build -f docker/vllm_disagg_inference.ubuntu.amd.Dockerfile -t / . +# +# BASE_IMAGE is the open rocm/vllm-dev ci_base pinned by the validated recipe +# (dist-inf-cookbook Dockerfile.vllm.mori121_shareable). Override --build-arg +# BASE_IMAGE=... to build on a different ROCm base. vLLM compile is long (~30-60 min). +# ============================================================================= + +ARG BASE_IMAGE=rocm/vllm-dev:ci_base-0fcd9b99cc9d63202da4c858d8ebc6582c9e2491 +FROM ${BASE_IMAGE} + +ENTRYPOINT [] +WORKDIR /app + +ARG GFX_COMPILATION_ARCH="gfx942" +ARG PYTORCH_ROCM_ARCH="gfx942" +ARG MAX_JOBS=32 +# NIXL/RIXL transport for the rixl connector. Default 1 => all connectors built +# (UCX/RIXL/rocSHMEM/DeepEP). Set --build-arg WITH_NIXL=0 for a lean MoRI-EP-only image. +ARG WITH_NIXL=1 +ARG NIC_COMPILATION_ARCH="cx7" + +# ----------------------------------------------------------------------------- +# 1. MoRI: replace the base's bundled MoRI with the validated ROCm/MoRI @ v1.2.1 +# (the version for the 06_29 mori121 image, dist-inf-cookbook +# Dockerfile.vllm.mori121_shareable). v1.2.1 carries the EP/RDMA correctness fixes +# plus the ROCm-7.2.3 dmabuf registration path used by the connector .env +# (expandable_segments:False). MoRI is JIT-built, so this swaps the JIT sources the +# kernels compile from at runtime. +# BUILD CONFIG: match the cookbook build — MORI_GPU_ARCHS=gfx942, BUILD_UMBP=OFF, +# DEFAULT NIC backends. Do NOT pass USE_IONIC=OFF / USE_BNXT=OFF: disabling NIC +# backends produced a MoRI that deadlocked at the cross-node EP all-to-all init. +# ----------------------------------------------------------------------------- +ARG MORI_REPO=https://github.com/ROCm/mori.git +# 42e895472b08: MoRI main tip past v1.2.1, validated by MAD-private #338 for GLM-5.1 +# DSA WideEP disagg (v1.2.1 large-transfer notify path was insufficient at high EP). +ARG MORI_REF=42e895472b08 +ENV MORI_GPU_ARCHS=gfx942 +# Newer MoRI added the UMBP subsystem which requires gRPC (grpcpp/grpcpp.h) not +# present in this base; UMBP is unrelated to the EP dispatch/combine kernels, so +# disable it to avoid pulling in a gRPC build dependency. +ENV BUILD_UMBP=OFF BUILD_UMBP_SPDK=OFF +# Build/install matches dist-inf-cookbook Dockerfile.vllm.mori121_shareable for v1.2.1: +# `BUILD_UMBP=OFF pip install .` (default build isolation). apt/pip build tooling kept +# for bases that lack it; harmless where already present. +RUN sed -i 's|http://|https://|g' /etc/apt/sources.list 2>/dev/null || true && \ + sed -i 's|http://|https://|g' /etc/apt/sources.list.d/*.list 2>/dev/null || true && \ + apt-get update && apt-get install -y --no-install-recommends \ + git build-essential cmake ninja-build ccache libssl-dev pkg-config curl ca-certificates && \ + pip install meson==0.64.0 "pybind11[global]" tqdm prettytable && \ + pip uninstall -y amd_mori amd-mori amd-mori-nightly mori 2>/dev/null || true && \ + rm -rf /tmp/mori-src && \ + git clone --recursive "${MORI_REPO}" /tmp/mori-src && \ + cd /tmp/mori-src && git checkout "${MORI_REF}" && git submodule update --init --recursive && \ + BUILD_UMBP=OFF pip install . && \ + python3 -c "import mori, mori.io, mori.ops; print('MoRI OK at', mori.__path__[0])" && \ + mkdir -p /app && echo "MORI_REF=${MORI_REF}@$(git -C /tmp/mori-src rev-parse HEAD)" >> /app/versions.txt && \ + rm -rf /tmp/mori-src + +# ----------------------------------------------------------------------------- +# 2. AITER: build STOCK upstream ROCm/aiter @ e03fa6040 from source (NO fork, +# NO gqa64-fold patch). Under vLLM #47766 the sparse-MLA persistent path stays +# ON, so GLM's gqa=64 decode hits aiter's PRE-EXISTING persistent gqa64->16 fold +# (aiter/mla.py: `nhead in range(32,128+1,16) and persistent_mode`); the fork's +# extra non-persistent fold is never exercised, so stock is sufficient. +# Validated by MAD-private #338: 1P/1D EP8 + 2P/2D EP16 NIAH PASS on this exact +# aiter tip under #47766. Pin the exact commit (the one tested), not the release +# wheel. Then invalidate the stale prewarmed JIT cache compiled against the old .so. +# ----------------------------------------------------------------------------- +ARG AITER_REPO=https://github.com/ROCm/aiter.git +ARG AITER_REF=e03fa6040 +RUN echo "Compiling STOCK AITER (no fork) from ${AITER_REPO}@${AITER_REF}" && \ + rm -rf /tmp/aiter-src && \ + git clone --recursive "${AITER_REPO}" /tmp/aiter-src && \ + cd /tmp/aiter-src && git checkout "${AITER_REF}" && \ + git submodule update --init --recursive && \ + (pip uninstall -y amd_aiter amd-aiter aiter 2>/dev/null || true) && \ + pip install --no-build-isolation --no-deps -v . && \ + pip install --no-deps -U "flydsl>=0.1.7,<0.1.9" && \ + echo "AITER_REF=${AITER_REF}@$(git rev-parse HEAD) (stock ROCm/aiter, no fork)" >> /app/versions.txt && \ + rm -rf /tmp/aiter-src && \ + python3 - <<'PYEOF' +# Verify aiter/mla.py installed + has the persistent gqa64 fold, WITHOUT importing +# aiter/torch (torch->amdsmi->libamd_smi.so is not loadable at build: no GPU in sandbox). +import glob, pathlib +cands = glob.glob("/usr/local/lib/python*/dist-packages/aiter/mla.py") + \ + glob.glob("/usr/lib/python*/dist-packages/aiter/mla.py") +assert cands, "aiter/mla.py not found in site-packages after install" +src = pathlib.Path(cands[0]).read_text() +assert "persistent_mode" in src, f"AITER persistent fold path MISSING in {cands[0]}" +print("STOCK AITER OK (persistent gqa64 fold path present):", cands[0]) +PYEOF +RUN rm -rf /opt/vllm_cache/aiter_jit /root/.aiter && echo "cleared stale AITER JIT cache" && \ + echo "AITER_REF=${AITER_REF} (stock)" >> /app/versions.txt + +# ----------------------------------------------------------------------------- +# 3. vLLM: compile from source at the 06_29 validated Wide-EP WRITE-mode branch +# (matches the published dist-inf-cookbook mori121 image). Full source compile +# (the base ships a different commit). The MoRIIO disagg fixes (#39276 notify, +# #41751 LL split, DP-rank hash-failsafe) are native in this branch, so no runtime +# patcher is needed. Override VLLM_REF to rebuild a different commit; build only +# committed commits (no working-tree edits). +# ----------------------------------------------------------------------------- +# VLLM_REPO/REF are a PUBLIC GitHub repo + branch (the Wide-EP WRITE-mode vLLM the +# dist-inf-cookbook mori121 image builds from). Override to your own vLLM fork/branch. +ARG VLLM_REPO=https://github.com/raviguptaamd/vllm.git +ARG VLLM_REF=glm5.1-dsa-wideEP_on_shik_latest +ENV VLLM_TARGET_DEVICE=rocm \ + PYTORCH_ROCM_ARCH=${PYTORCH_ROCM_ARCH} \ + MAX_JOBS=${MAX_JOBS} +RUN rm -rf /tmp/vllm-src && \ + git clone "${VLLM_REPO}" /tmp/vllm-src && \ + cd /tmp/vllm-src && git checkout "${VLLM_REF}" && \ + echo "VLLM_REF=${VLLM_REF}@$(git rev-parse HEAD)" >> /app/versions.txt && \ + pip uninstall -y vllm 2>/dev/null || true && \ + pip install --no-deps --no-build-isolation -v . && \ + python3 -c "import vllm; print('vLLM', vllm.__version__, 'from', vllm.__file__)" && \ + rm -rf /tmp/vllm-src + +# Cross-check MoRI + AITER survived the vLLM install (no silent downgrade). +RUN python3 - <<'PYEOF' +from importlib.metadata import version as v, PackageNotFoundError +def get(names): + for n in names: + try: return v(n) + except PackageNotFoundError: pass + return None +av = get(("amd-aiter", "amd_aiter", "aiter")) +# Stock source build of ROCm/aiter@e03fa6040 reports 0.1.17.dev195+ge03fa6040. +# Verify the aiter install survived the vLLM install (present + carries the e03fa6040 +# commit tag) rather than pinning a release version string. +assert av and "e03fa6040" in av, f"AITER missing/downgraded (want e03fa6040 build): {av!r}" +import mori, mori.io, mori.ops +print("Post-vLLM check OK: AITER", av, "+ MoRI importable") +PYEOF + +# ----------------------------------------------------------------------------- +# 4. vllm-router (DP-rank round-robin + MoRIIO connector) — built in, so NO +# external vllm-router binary is needed (leave ROUTER_BINARY unset). +# Source = vllm-project/router PR #181 branch, which now carries BOTH the +# round-robin DP-rank fix (11841c0d) AND the 2P2D KV-notify fix (6409ac1: +# remote_dp_rank_override + remote_dp_size). The KV-notify fix is REQUIRED: +# without it the 2P2D EP=16 run reproducibly wedges with "remote blocks never +# arrived" deferred-write expiries (decode notify targets the wrong DP rank). +# This is the exact source of the validated vllm-router-2p2d-dpfix binary. +# Pinned Rust toolchain (>=1.88: router deps time/home require rustc 1.88). +# ----------------------------------------------------------------------------- +ARG ROUTER_REPO=https://github.com/raviguptaamd/router.git +ARG ROUTER_REF=ravgupta/discovery-dp-rank-roundrobin +ARG RUST_TOOLCHAIN=1.88.0 +RUN if ! command -v cargo >/dev/null 2>&1; then \ + curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y --default-toolchain "${RUST_TOOLCHAIN}"; \ + fi && \ + export PATH="/root/.cargo/bin:${PATH}" && \ + rm -rf /tmp/vllm-router-src && \ + git clone --filter=blob:none "${ROUTER_REPO}" /tmp/vllm-router-src && \ + cd /tmp/vllm-router-src && git checkout "${ROUTER_REF}" && \ + cargo build --release && \ + install -m 755 target/release/vllm-router /usr/local/bin/vllm-router && \ + vllm-router --help 2>&1 | grep -q moriio && \ + echo "VLLM_ROUTER_REF=${ROUTER_REPO}@${ROUTER_REF}@$(git -C /tmp/vllm-router-src rev-parse HEAD)" >> /app/versions.txt && \ + rm -rf /tmp/vllm-router-src + +# ----------------------------------------------------------------------------- +# 4b. WITH_NIXL=1 (default): UCX + RIXL(+nixlbench) + rocSHMEM + DeepEP from source, +# so the rixl connector (NIXL TP + DeepEP wideEP) is present. Single guarded RUN so +# WITH_NIXL=0 skips it entirely (no layers, no cost). Build-verified on ci_base. +# ----------------------------------------------------------------------------- +ENV _ROCM_DIR=/opt/rocm \ + _UCX_SOURCE=https://github.com/ROCm/ucx.git \ + _UCX_BRANCH=da3fac2a \ + _UCX_INSTALL_DIR=/usr/local/ucx/ \ + _RIXL_SOURCE=https://github.com/ROCm/RIXL.git \ + _RIXL_BRANCH=f33a5599 \ + _RIXL_INSTALL_DIR=/usr/local/RIXL/install \ + _NIXLBENCH_INSTALL_DIR=/usr/local/RIXL +RUN if [ "${WITH_NIXL}" != "1" ]; then \ + echo "WITH_NIXL=${WITH_NIXL}: skipping UCX/RIXL/rocSHMEM/DeepEP (MoRI-EP + base DeepEP only)"; \ + else set -e && \ + echo "WITH_NIXL=1: building UCX + RIXL + rocSHMEM + DeepEP" && \ + apt-get update && apt-get install -y \ + autoconf automake libtool autogen pkg-config m4 gcc make \ + librdmacm-dev rdmacm-utils infiniband-diags ibverbs-utils perftest ethtool \ + libibverbs-dev rdma-core strace libgflags-dev \ + libaio-dev liburing-dev libcpprest-dev libgrpc-dev libgrpc++-dev \ + libprotobuf-dev protobuf-compiler-grpc wget && \ + pip install meson==0.64.0 "pybind11[global]" pyyaml && \ + # UCX + cd /tmp && git clone "${_UCX_SOURCE}" && cd ucx && git checkout "${_UCX_BRANCH}" && \ + ./autogen.sh && mkdir -p build && cd build && \ + ../configure --prefix="${_UCX_INSTALL_DIR}" --with-rocm="${_ROCM_DIR}" \ + --disable-go --disable-java --disable-assertions --enable-mt && \ + make -j && make install && \ + # googletest (RIXL dep) + cd /tmp && wget -q https://github.com/google/googletest/archive/refs/tags/v1.14.0.tar.gz && \ + tar -xzf v1.14.0.tar.gz && cd googletest-1.14.0 && mkdir -p build && cd build && \ + cmake -DBUILD_SHARED_LIBS=on .. && make -j && make install && \ + # RIXL + python bindings + cd /tmp && git clone "${_RIXL_SOURCE}" && cd RIXL && git checkout "${_RIXL_BRANCH}" && \ + meson setup build/ --prefix="${_RIXL_INSTALL_DIR}" -Ducx_path="${_UCX_INSTALL_DIR}" \ + -Ddisable_gds_backend=true -Dcudapath_inc="${_ROCM_DIR}/include" -Dcudapath_lib="${_ROCM_DIR}/lib" && \ + cd build && ninja && ninja install && cd /tmp/RIXL && \ + pip install --config-settings=setup-args="-Dcudapath_inc=${_ROCM_DIR}/include" \ + --config-settings=setup-args="-Dcudapath_lib=${_ROCM_DIR}/lib" \ + --config-settings=setup-args="-Ducx_path=${_UCX_INSTALL_DIR}" \ + --config-settings=setup-args="-Ddisable_gds_backend=true" . && \ + # rocSHMEM (DeepEP dep) + cd /tmp && git clone --no-checkout --filter=blob:none https://github.com/ROCm/rocm-systems.git && \ + cd rocm-systems && git sparse-checkout set --cone projects/rocshmem && git checkout develop && \ + mkdir -p /tmp/rocshmem-build && cd /tmp/rocshmem-build && \ + /tmp/rocm-systems/projects/rocshmem/scripts/build_configs/all_backends \ + -DUSE_EXTERNAL_MPI=OFF -DGPU_TARGETS="${GFX_COMPILATION_ARCH}" && \ + # DeepEP (build develop against the installed vLLM/torch) + cd /tmp && git clone https://github.com/ROCm/DeepEP.git && cd DeepEP && \ + PYTORCH_ROCM_ARCH="${GFX_COMPILATION_ARCH}" CFLAGS="-O3 -fPIC" \ + CXXFLAGS="-O3 -fPIC --offload-arch=${GFX_COMPILATION_ARCH}" HIP_CXX_FLAGS="-O3 -fPIC" \ + python3 setup.py --variant rocm --nic "${NIC_COMPILATION_ARCH}" build develop && \ + echo "WITH_NIXL build complete" >> /app/versions.txt && \ + rm -rf /tmp/ucx /tmp/googletest-1.14.0 /tmp/v1.14.0.tar.gz /tmp/rocm-systems /tmp/rocshmem-build; \ + fi +ENV LD_LIBRARY_PATH="/usr/local/ucx/lib:/usr/local/lib:/usr/local/RIXL/install/lib:${LD_LIBRARY_PATH}" \ + PATH="/usr/local/ucx/bin:${PATH}" + +# ----------------------------------------------------------------------------- +# 5. Cache locations (structural: WHERE the JIT/compile caches live in the image). +# These are the mount target for the launcher's persistent host JIT cache. +# ----------------------------------------------------------------------------- +# The image ships NO runtime recipe / tuning / platform ENV. By design, everything +# run-tunable is applied at launch, so this image stays a clean binary/library artifact +# and the same image serves any model/cluster without a rebuild: +# - model-serving recipe (KV_BLOCK_SIZE, KV_CACHE_DTYPE, *_CUDAGRAPH_MODE, *_MORI_BACKEND, +# GPU_MEMORY_UTILIZATION, KV_CACHE_MEMORY_BYTES, VLLM_ROCM_USE_AITER_MLA, ...) +# -> scripts/vllm_dissag/models.yaml (per-model env:, so dense vs MoE differ) +# - ROCm-7.2.3 GPU-RDMA platform env (expandable_segments:False x2, MORI_GPU_ARCHS, +# HSA_ENABLE_IPC_MODE_LEGACY=0, HSA_NO_SCRATCH_RECLAIM) and the MoRI/RDMA fabric +# tuning (MORI_RDMA_TC/SL, MORI_IB_GID_INDEX, MORI_NUM_QP_PER_PE, VLLM_MORIIO_*, ...) +# -> scripts/vllm_dissag/connectors/.env (cluster-editable, no rebuild) +# The slurm launcher forwards both via `docker -e` (platform env must reach PID 1 - +# PyTorch reads alloc-conf at import). Running this image WITHOUT the launcher: set the +# vars you need yourself (see connectors/moriio.env + models.yaml for the values). +ENV AITER_JIT_DIR=/opt/vllm_cache/aiter_jit \ + VLLM_CACHE_ROOT=/opt/vllm_cache/vllm \ + TRITON_CACHE_DIR=/opt/vllm_cache/triton \ + COMGR_CACHE_DIR=/opt/vllm_cache/comgr + +# ----------------------------------------------------------------------------- +# 6. CRITICAL: scrub build-time MoRI JIT state. The `import mori` verification +# steps above compile/lock MoRI EP kernels under /root/.mori/jit on THIS build +# host, leaving stale .hsaco.lock files (ep_internode_v1, ep_internode_v1ll, ...). +# At runtime on the cluster, MoriAll2AllManager finds those locks, waits on a +# build-in-progress whose owner PID is long gone, and DEADLOCKS at ep:0 init. +# A clean image ships /root/.mori empty -> runtime compiles fresh. +# Clearing these makes the from-source image boot clean on 2P2D/4P4D. +# ----------------------------------------------------------------------------- +RUN rm -rf /root/.mori /tmp/mori_jit_* && mkdir -p /root/.mori && \ + echo "JIT_SCRUBBED: /root/.mori + /tmp/mori_jit_* cleared at build end" >> /app/versions.txt + +RUN cat /app/versions.txt 2>/dev/null | tail -20 || true diff --git a/models.json b/models.json index af914f46..689754db 100644 --- a/models.json +++ b/models.json @@ -3420,6 +3420,38 @@ }, "args": "-N 2 -n 2" }, + { + "name": "pyt_vllm_disagg_mori_glm-5.1-fp8", + "url": "", + "dockerfile": "docker/vllm_disagg_inference.glmv5.1", + "scripts": "scripts/vllm_dissag/run_xPyD_models.slurm", + "data": "huggingface", + "n_gpus": "-1", + "owner": "mad.support@amd.com", + "training_precision": "", + "tags": [ + "pyt", + "vllm", + "vllm_disagg", + "mori_ep", + "inference" + ], + "timeout": -1, + "distributed": { + "launcher": "slurm_multi" + }, + "env_vars": { + "DOCKER_IMAGE_NAME": "", + "MODEL_NAME": "GLM-5.1-FP8", + "xP": "1", + "yD": "1", + "RUN_MORI": "1", + "RUN_DEEPEP": "0", + "GLM_SKIP_PATCHERS": "1", + "BENCHMARK_COMBINATIONS": "1024/1024" + }, + "args": "-N 2 -n 2" + }, { "name": "pyt_vllm_disagg_mori_deepseek-v3-5layer", "url": "", diff --git a/scripts/vllm_dissag/apply_glm_aiter_sampling_oob_fix.py b/scripts/vllm_dissag/apply_glm_aiter_sampling_oob_fix.py new file mode 100644 index 00000000..d491b838 --- /dev/null +++ b/scripts/vllm_dissag/apply_glm_aiter_sampling_oob_fix.py @@ -0,0 +1,148 @@ +#!/usr/bin/env python3 +"""Overlay the fixed AITER sampling kernel (ROCm/aiter #3658 + hang cap) into the image. + +DEFECT 2 (the 8k prefill/decode crash): the AITER TopP/TopK sampling kernel +(csrc/cpp_itfs/sampling/sampling.cuh) has two bugs on the released aiter post3 +that this image ships: + + 1. HSA OUT-OF-BOUNDS (ROCm/aiter #3658): SamplingTempStorage::last_valid_id is + never initialized. When a probs row is all-zero / NaN (more likely at long + input, e.g. 8k), the guarded write-back (max_valid != -1) is skipped, and the + fallback `sampled_id = temp_storage.last_valid_id` reads UNINITIALIZED shared + memory -> garbage index -> `probs[row*d + sampled_id]` dereferences OOB and + HSA page-faults ("Memory access fault by GPU node-N"). Deterministic under + CUDA graph (shared-mem residue is stable across replays). This is the silent + worker death at 8k that collapses the disagg DP group -> 503. + Fix: init `last_valid_id = 0` at top of each loop iter + defensive clamp on + the loaded sampled_id before it indexes probs. + + 2. REJECTION-SAMPLING HANG: the bisection `do { ... } while(low < high)` can + spin forever when the [low,high] interval stagnates in float precision on a + degenerate (near-uniform) row -> never-completing HSA signal / hang (the + "sampler hang" that forced the skip-warmup workaround). Fix: cap the loop at + kMaxSamplingRounds=32 (float32 mantissa is exhausted well within 32 rounds, + so healthy distributions always converge via break long before the cap). + +Both fixes land in sampling.cuh. #3658 is MERGED upstream but NOT in the released +aiter post3 (this image). Source: A/B-tested by Shiksha (shikpate); staged fixed +tree at SAMPLING_FIX_DIR. + +METHOD (from Shiksha's validated in-container overlay): copy the whole patched +sampling source dir (.cuh + .py + .jinja) over the container's aiter, then purge +any compiled sampling JIT objects so the kernel recompiles from the fixed source +on next use. + +Idempotent: skips if the fix markers are already present. Model-agnostic at the +kernel level, but invoked from the GLM patch hook. Safe no-op if the staged fix +dir or the target aiter dir is absent. + +Usage: apply_glm_aiter_sampling_oob_fix.py + (vllm_install_dir arg is accepted for hook uniformity but not required; + the aiter dir is resolved via `import aiter`.) +""" +import os +import shutil +import subprocess +import sys + +FIX_DIR = os.environ.get( + "SAMPLING_FIX_DIR", + "/shared_inference/ravgupta/aiter_sampling_fix_3658/sampling_patched", +) +MARKERS = ("last_valid_id = 0", "kMaxSamplingRounds") + + +def _aiter_sampling_dir(): + """Locate the installed aiter sampling source dir (aiter_meta/csrc/...).""" + try: + import aiter # noqa: F401 + except Exception as e: # noqa: BLE001 + print(f"[sampling-fix] aiter not importable ({e}); skipping.") + return None + # The kernel source lives under aiter_meta (sibling of aiter), path is stable. + candidates = [] + try: + import aiter_meta # type: ignore + + candidates.append( + os.path.join(os.path.dirname(aiter_meta.__file__), + "csrc", "cpp_itfs", "sampling") + ) + except Exception: # noqa: BLE001 + pass + # Fallback: search site-packages. + import aiter + sp = os.path.dirname(os.path.dirname(aiter.__file__)) + candidates.append(os.path.join(sp, "aiter_meta", "csrc", "cpp_itfs", "sampling")) + for c in candidates: + if os.path.isdir(c): + return c + print(f"[sampling-fix] could not locate aiter sampling dir (tried {candidates}); skipping.") + return None + + +def main() -> int: + tgt = _aiter_sampling_dir() + if tgt is None: + return 0 # safe no-op + + tgt_cuh = os.path.join(tgt, "sampling.cuh") + if os.path.isfile(tgt_cuh): + cur = open(tgt_cuh, errors="ignore").read() + if all(m in cur for m in MARKERS): + print(f"[sampling-fix] already applied (markers present) in {tgt_cuh}.") + return 0 + + if not os.path.isdir(FIX_DIR): + print(f"[sampling-fix] WARN: staged fix dir {FIX_DIR} not found; leaving image kernel unpatched.", file=sys.stderr) + return 0 + + src_cuh = os.path.join(FIX_DIR, "sampling.cuh") + if not os.path.isfile(src_cuh) or not all(m in open(src_cuh, errors="ignore").read() for m in MARKERS): + print(f"[sampling-fix] WARN: staged {src_cuh} missing/lacks fix markers; skipping.", file=sys.stderr) + return 0 + + # Overlay the whole sampling source dir (.cuh + .py + .jinja), per Shiksha's method. + copied = [] + for fn in os.listdir(FIX_DIR): + s = os.path.join(FIX_DIR, fn) + if os.path.isfile(s): + shutil.copy2(s, os.path.join(tgt, fn)) + copied.append(fn) + print(f"[sampling-fix] overlaid #3658 + hang-cap into {tgt}: {', '.join(sorted(copied))}") + + # Verify. + cur = open(tgt_cuh, errors="ignore").read() + if not all(m in cur for m in MARKERS): + print(f"[sampling-fix] ERROR: markers still absent after overlay in {tgt_cuh}.", file=sys.stderr) + return 1 + + # Purge any compiled sampling JIT objects so the kernel recompiles from source. + purged = 0 + for base in ( + os.path.expanduser("~/.aiter"), "/root/.aiter", "/tmp/aiter", + "/opt/vllm_cache/aiter_jit", os.path.join(os.path.dirname(tgt), "..", "..", "jit"), + ): + if base and os.path.isdir(base): + try: + out = subprocess.run( + ["find", base, "-maxdepth", "6", "-iname", "*sampling_from_probs*"], + capture_output=True, text=True, timeout=60, + ) + for p in out.stdout.split(): + try: + if os.path.isdir(p): + shutil.rmtree(p, ignore_errors=True) + else: + os.remove(p) + purged += 1 + except OSError: + pass + except Exception: # noqa: BLE001 + pass + print(f"[sampling-fix] purged {purged} stale sampling JIT object(s); kernel will recompile from fixed source.") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/vllm_dissag/apply_glm_dsa_indexer_warmup_fix.py b/scripts/vllm_dissag/apply_glm_dsa_indexer_warmup_fix.py new file mode 100755 index 00000000..3eaef8cf --- /dev/null +++ b/scripts/vllm_dissag/apply_glm_dsa_indexer_warmup_fix.py @@ -0,0 +1,251 @@ +#!/usr/bin/env python3 +"""Force-compile the GLM-5.1 DSA sparse-attention indexer Triton kernels at BOOT. + +PROBLEM (root cause of the "first big prompt stalls the whole DP group" hang): + The DSA indexer's Triton kernels are seq-length specialized: + - v1/attention/ops/triton_fp8_mqa_logits.py flips `matrix_instr_nonkdim` + at seq_len<=1024 and launches with grid=[(seq_len,)] (seq_len is a + specialized kernel arg) -> a >1024-row prefill needs a *different* JIT + specialization than a <=1024-row one. + But the boot-time warmup never drives the indexer: + - profile_run() calls _dummy_run(is_profile=True) with force_attention=False + and cudagraph mode NONE -> attn_metadata stays None -> sparse_attn_indexer + takes the `sparse_attn_indexer_fake` path (see the "careful! this will be + None in dummy run" comment in layers/sparse_attn_indexer.py). The real + kernels are never compiled. + - _warmup_and_capture() only sets force_attention=True when the cudagraph + runtime mode is FULL; the DSA indexer builder reports UNIFORM_BATCH, so on + this ROCm/DP build the mixed prefill-decode graphs are PIECEWISE and + force_attention stays False. Even when attention IS forced, capture uses + uniform-decode / small mixed batches -- never a large prefill at + max_num_batched_tokens -- so the >1024 specialization is still absent. + Effect: the first >=8k prompt JIT-compiles the indexer kernel mid-inference on + whichever DP rank happens to receive it. That rank falls out of the DP lockstep + gloo all_reduce (coordinate_batch_across_dp) while it compiles -> the whole DP + group collapses. It is also a general cold-cache robustness hole. + +FIX (surgical, reuses vLLM's OWN metadata construction -- no hand-synthesized +tensors, so zero risk of a bad-input crash at boot): + 1. gpu_model_runner.py: add a `_maybe_warmup_dsa_indexer()` method. It is a + strict NO-OP unless one of the runner's attention backends is (a subclass + of) DeepseekV32IndexerBackend. When present, it runs + `_dummy_run(..., force_attention=True, cudagraph_runtime_mode=NONE)` at TWO + prefill-size regimes -- a small one (<=1024 rows) and a large one + (max_num_batched_tokens, >1024) -- so BOTH Triton specializations compile. + `force_attention=True` makes _dummy_run build a real + DeepseekV32IndexerMetadata via the normal _build_attention_metadata path + (num_prefills>0 because the default dummy batch is multi-token requests and + the indexer decode_threshold is 1), which drives the real + `sparse_attn_indexer` prefill kernels. The whole thing is wrapped in + try/except that only WARNs -- a warmup failure must never crash boot. + 2. gpu_worker.py: call it from compile_or_warm_up_model, right after the + existing warmup loop and before kernel_warmup(). At that point the KV cache + is already allocated (initialize_from_config runs before + compile_or_warm_up_model), which the forced-attention indexer path needs. + +Idempotent + anchor-based (matches the other apply_glm_* patchers): + * Each hunk self-detects if already applied (marker string present) -> no-op. + * Missing anchor -> WARN and skip that hunk (safe across vllm revisions; the + rebase may already warm the indexer natively or have refactored the site). + * Anchor found but the file does not contain the applied marker and the + replace produces no change -> hard error (would silently keep the bug). + * py_compile at the end; hard error if the patched file won't compile. + +Usage: apply_glm_dsa_indexer_warmup_fix.py +""" +import os +import sys + +RUNNER_REL = "v1/worker/gpu_model_runner.py" +WORKER_REL = "v1/worker/gpu_worker.py" + +MARKER = "glm-dsa-indexer-warmup" + +# --- Hunk A: new method inserted immediately before `def capture_model` ------- +# Anchor: the (unique) capture_model definition head in gpu_model_runner.py. +RUNNER_ANCHOR = " def capture_model(self) -> int:\n" + +RUNNER_METHOD = ''' def _maybe_warmup_dsa_indexer(self) -> None: + """Force-compile the DSA sparse-attention indexer Triton kernels at boot. + + NO-OP unless this model actually has a DeepseekV32IndexerBackend (GLM-5.1 + DSA / DeepSeek V3.2). The indexer kernels are seq-length specialized + (triton_fp8_mqa_logits flips matrix_instr_nonkdim at seq_len<=1024 and + launches grid=[(seq_len,)]), and the normal profile/warmup passes never + drive the indexer (attn_metadata is None -> the *_fake path). Without this + the first large prompt JIT-compiles mid-inference and, under DP lockstep, + stalls the whole group. We warm BOTH regimes: a small (<=1024) and a large + (max_num_batched_tokens, >1024) prefill batch, using force_attention=True + so _dummy_run builds a real indexer metadata via the standard path. + """ + # {marker} + try: + from vllm.v1.attention.backends.mla.indexer import ( + DeepseekV32IndexerBackend, + ) + except Exception: # noqa: BLE001 -- backend module absent -> not a DSA build + return + + has_indexer = False + try: + for attn_group in self._attn_group_iterator(): + backend = getattr(attn_group, "backend", None) + if backend is not None and isinstance(backend, type) and issubclass( + backend, DeepseekV32IndexerBackend + ): + has_indexer = True + break + except Exception: # noqa: BLE001 -- iterator shape changed -> stay a no-op + return + if not has_indexer: + return + + # Two prefill-size regimes so both Triton specializations compile. + # Small must be <=1024 rows; large must exceed 1024 (use the real max). + max_tokens = int(self.max_num_tokens) + small = min(512, max_tokens) + sizes = [] + for s in (small, max_tokens): + if s > 0 and s not in sizes: + sizes.append(s) + + logger.info( + "Warming up DSA indexer kernels at prefill sizes %s " + "to avoid mid-inference JIT.", + sizes, + ) + for size in sizes: + try: + self._dummy_run( + size, + cudagraph_runtime_mode=CUDAGraphMode.NONE, + force_attention=True, + skip_eplb=True, + remove_lora=False, + ) + except Exception as e: # noqa: BLE001 -- warmup must NEVER crash boot + logger.warning( + "DSA indexer warmup at size %d failed (%s); the kernel may " + "JIT-compile on first use instead.", + size, + e, + ) + self._sync_device() + +'''.replace("{marker}", MARKER) + +# --- Hunk B: call site in gpu_worker.compile_or_warm_up_model ----------------- +WORKER_ANCHOR = ( + " self.model_runner.maybe_remove_all_loras(" + "self.model_runner.lora_config)\n" + "\n" + " # Warmup and tune the kernels used during model execution before\n" + " # cuda graph capture.\n" + " kernel_warmup(self)\n" +) + +WORKER_REPLACEMENT = ( + " self.model_runner.maybe_remove_all_loras(" + "self.model_runner.lora_config)\n" + "\n" + " # " + MARKER + ": force-compile the DSA sparse-attention indexer\n" + " # Triton kernels now (KV cache is allocated), so a large prompt\n" + " # never JIT-compiles them mid-inference and stalls DP lockstep.\n" + " # No-op unless this model has a DeepseekV32IndexerBackend.\n" + " if hasattr(self.model_runner, \"_maybe_warmup_dsa_indexer\"):\n" + " self.model_runner._maybe_warmup_dsa_indexer()\n" + "\n" + " # Warmup and tune the kernels used during model execution before\n" + " # cuda graph capture.\n" + " kernel_warmup(self)\n" +) + + +def _patch_file(path, tag, anchor, apply_fn, already_marker): + """Return 0 on success/no-op, 1 on hard error.""" + if not os.path.isfile(path): + print(f"[{tag}] {path} not found -- skipping (layout differs).") + return 0 + src = open(path).read() + if already_marker in src: + print(f"[{tag}] already applied ({already_marker} present) in {path} -- no-op.") + return 0 + if anchor not in src: + print( + f"[{tag}] WARN: anchor not found in {path} -- skipping " + "(assuming native warmup / refactor)." + ) + return 0 + new_src = apply_fn(src) + if new_src == src: + print( + f"[{tag}] ERROR: anchor found but patch produced no change in {path}.", + file=sys.stderr, + ) + return 1 + try: + open(path, "w").write(new_src) + except OSError as e: + print(f"[{tag}] ERROR: failed to write patched {path}: {e}", file=sys.stderr) + return 1 + if already_marker not in open(path).read(): + print( + f"[{tag}] ERROR: post-write verification failed in {path}.", + file=sys.stderr, + ) + return 1 + print(f"[{tag}] patched {path} -- 1 hunk.") + return 0 + + +def main() -> int: + if len(sys.argv) != 2: + print(f"usage: {sys.argv[0]} ", file=sys.stderr) + return 2 + vllm_dir = sys.argv[1] + + runner_path = os.path.join(vllm_dir, RUNNER_REL) + worker_path = os.path.join(vllm_dir, WORKER_REL) + + rc = 0 + + # Hunk A: insert the method before capture_model. + rc |= _patch_file( + runner_path, + "glm-dsa-warmup", + RUNNER_ANCHOR, + lambda s: s.replace(RUNNER_ANCHOR, RUNNER_METHOD + RUNNER_ANCHOR, 1), + MARKER, + ) + + # Hunk B: call it from compile_or_warm_up_model. + rc |= _patch_file( + worker_path, + "glm-dsa-warmup", + WORKER_ANCHOR, + lambda s: s.replace(WORKER_ANCHOR, WORKER_REPLACEMENT, 1), + MARKER, + ) + + if rc: + return 1 + + # py-compile sanity for whichever files exist. + try: + import py_compile + + for p in (runner_path, worker_path): + if os.path.isfile(p): + py_compile.compile(p, doraise=True) + print("[glm-dsa-warmup] py_compile OK") + except Exception as e: # noqa: BLE001 + print( + f"[glm-dsa-warmup] ERROR: patched file fails to compile: {e}", + file=sys.stderr, + ) + return 1 + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/vllm_dissag/apply_glm_dsa_kernel_fix.py b/scripts/vllm_dissag/apply_glm_dsa_kernel_fix.py new file mode 100755 index 00000000..33f50dc4 --- /dev/null +++ b/scripts/vllm_dissag/apply_glm_dsa_kernel_fix.py @@ -0,0 +1,85 @@ +#!/usr/bin/env python3 +"""Apply the GLM-5.1 DSA sparse-attention invalid-token kernel fix (vllm #45324). + +The DSA indexer kernel `_convert_req_index_to_global_index_kernel` in + vllm/v1/attention/backends/mla/rocm_aiter_mla_sparse.py +maps invalid token slots to 0 instead of -1. With block-size 1 + DSA sparse MLA +that corrupts KV reads and the model emits `!!!` for every prompt. + +Upstream fix: vllm-project/vllm #45324 -- flip the 0 to -1 in the tl.where call: + is_invalid_tok | (~valid_block), 0, base * BLOCK_SIZE + inblock_off + is_invalid_tok | (~valid_block), -1, base * BLOCK_SIZE + inblock_off + +Design (matches launcher contract -- runs unconditionally for GLM, aborts on real +failure): + * IDEMPOTENT : if already -1, report and exit 0 (no-op). + * SELF-SKIPPING: if the file/anchor is absent (refactored or the rebase already + fixed it differently), report and exit 0 -- do NOT abort, because b10a9f7a may + carry the fix natively. We only fail on the one unambiguous bad state we can + fix and didn't, or on write failure. + * VERIFIES the post-write state. + +Usage: apply_glm_dsa_kernel_fix.py +""" +import os +import re +import sys + +REL = "v1/attention/backends/mla/rocm_aiter_mla_sparse.py" + +# Anchor is the stable right-hand side of the tl.where; the middle operand is the +# 0 (buggy) / -1 (fixed) we toggle. Whitespace-tolerant. +RE_ANY = re.compile( + r"(is_invalid_tok\s*\|\s*\(~valid_block\)\s*,\s*)(-?\d+)(\s*,\s*base\s*\*\s*BLOCK_SIZE\s*\+\s*inblock_off)" +) + + +def main() -> int: + if len(sys.argv) != 2: + print(f"usage: {sys.argv[0]} ", file=sys.stderr) + return 2 + vllm_dir = sys.argv[1] + path = os.path.join(vllm_dir, REL) + + if not os.path.isfile(path): + # File not present on this build -> nothing we can or should do. The + # rebase may use a different sparse backend layout. Do not block launch. + print(f"[glm-dsa] {REL} not found under {vllm_dir} -- skipping (assuming native/refactored).") + return 0 + + src = open(path).read() + m = RE_ANY.search(src) + if not m: + # Anchor gone (refactored / already fixed differently). Don't block. + print(f"[glm-dsa] invalid-token kernel anchor not found in {path} -- skipping (assuming native fix).") + return 0 + + cur = m.group(2) + if cur == "-1": + print(f"[glm-dsa] already fixed (kernel returns -1) in {path} -- no-op.") + return 0 + if cur != "0": + # Unexpected value -- surface it but don't guess. Treat as needs-attention. + print(f"[glm-dsa] ERROR: unexpected invalid-token return value '{cur}' (expected 0 or -1) in {path}.", + file=sys.stderr) + return 1 + + # cur == "0" : the known bug. Flip to -1. + new_src = src[:m.start(2)] + "-1" + src[m.end(2):] + try: + open(path, "w").write(new_src) + except OSError as e: + print(f"[glm-dsa] ERROR: failed to write patched {path}: {e}", file=sys.stderr) + return 1 + + # Verify. + chk = RE_ANY.search(open(path).read()) + if not chk or chk.group(2) != "-1": + print(f"[glm-dsa] ERROR: post-write verification failed in {path}.", file=sys.stderr) + return 1 + print(f"[glm-dsa] patched: invalid-token kernel now returns -1 (vllm #45324) in {path}.") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/vllm_dissag/apply_glm_dsa_moriio_dualkv_fix.py b/scripts/vllm_dissag/apply_glm_dsa_moriio_dualkv_fix.py new file mode 100755 index 00000000..ac6800bf --- /dev/null +++ b/scripts/vllm_dissag/apply_glm_dsa_moriio_dualkv_fix.py @@ -0,0 +1,176 @@ +#!/usr/bin/env python3 +"""Patch the MoRIIO KV connector to handle GLM-5.1 DSA's DUAL KV cache. + +PROBLEM (root cause of the 2P2D "Reaped deferred sends / no finished_sending" stall): + GLM-5.1 (GlmMoeDsaForCausalLM -> deepseek_v2.py) has TWO KV caches per layer: + - main MLA latent KV : MLAAttentionSpec, head_size = kv_lora_rank+rope (~576) + - DSA indexer KV : DeepseekV32IndexerCache, MLAAttentionSpec head_size = index_head_dim (128) + Both are 3D ("use_mla"), but DIFFERENT latent dim -> DIFFERENT per-block byte size. + The MoRIIO connector computes ONE global geometry from `first_kv_cache` and reuses it + for every cache, so the indexer cache is transferred with the main-MLA block size -> + wrong bytes/size -> the RDMA read for that region never reconciles -> completion notify + is never produced -> decode reaps deferred sends after 60s -> request hangs. + +FIX (surgical, per-layer geometry; no behavior change for single-cache MLA/DeepSeek): + 1. register_kv_caches: size each registered region by its OWN tensor (per-cache + region_len), not the global self.block_len. Also fix the local_kv_cache_size + append to use the current cache, not a stale loop var. + 2. _compute_block_transfer_offsets: derive shape from the PER-LAYER tensor + (self.kv_caches[layer_name].shape) instead of the global self.kv_cache_shape, + so transfer_size_byte / strides match that cache. + 3. _read_blocks: compute offsets PER LAYER inside the loop (was computed once from + first_layer and reused for all layers). + +Idempotent + anchor-based: each hunk checks if already applied / anchor present; +missing anchor -> warn-and-skip (so it is safe across connector revisions). A hunk +that finds its OLD anchor but fails to apply is a hard error (would silently keep the bug). + +Usage: apply_glm_dsa_moriio_dualkv_fix.py +""" +import os +import sys + +REL = "distributed/kv_transfer/kv_connector/v1/moriio/moriio_connector.py" + + +def main() -> int: + if len(sys.argv) != 2: + print(f"usage: {sys.argv[0]} ", file=sys.stderr) + return 2 + path = os.path.join(sys.argv[1], REL) + if not os.path.isfile(path): + print(f"[glm-dualkv] {REL} not found -- skipping (connector layout differs).") + return 0 + + src = open(path).read() + orig = src + applied = [] + + # --- Hunk 1: per-cache region_len in register_kv_caches --------------------- + h1_old = """ for cache_or_caches in kv_caches.values(): + cache_list = [cache_or_caches] if use_mla else cache_or_caches + for cache in cache_list: + base_addr = cache.data_ptr() + region_len = self.num_blocks * self.block_len + caches_data.append((base_addr, region_len, cache.device.index, "")) + kv_caches_base_addr.append(base_addr)""" + h1_new = """ for cache_or_caches in kv_caches.values(): + cache_list = [cache_or_caches] if use_mla else cache_or_caches + for cache in cache_list: + base_addr = cache.data_ptr() + # DSA dual-KV fix: size each region by its OWN tensor, not the + # global self.block_len (the DSA indexer cache has a different + # latent dim than the main MLA cache). + region_len = cache.nelement() * cache.element_size() + caches_data.append((base_addr, region_len, cache.device.index, "")) + kv_caches_base_addr.append(base_addr)""" + if "region_len = cache.nelement() * cache.element_size()" in src: + applied.append("h1 (already)") + elif h1_old in src: + src = src.replace(h1_old, h1_new, 1) + applied.append("h1") + else: + print("[glm-dualkv] WARN: h1 anchor (region_len loop) not found -- skipping h1.") + + # --- Hunk 1b: local_kv_cache_size uses current kv_cache, not stale `cache` -- + h1b_old = " self.local_kv_cache_size.append(cache.nelement() * cache.element_size())" + h1b_new = " self.local_kv_cache_size.append(kv_cache.nelement() * kv_cache.element_size())" + if h1b_new in src: + applied.append("h1b (already)") + elif h1b_old in src: + src = src.replace(h1b_old, h1b_new, 1) + applied.append("h1b") + else: + print("[glm-dualkv] WARN: h1b anchor (local_kv_cache_size) not found -- skipping h1b.") + + # --- Hunk 2: per-layer shape in _compute_block_transfer_offsets ------------- + h2_old = """ assert self.kv_cache_shape is not None, "KV caches shape not initialized" + is_mla = len(self.kv_cache_shape) == 3 + stride = self.kv_caches[layer_name].stride() + sz = self.kv_caches[layer_name].element_size() + if is_mla: + blknum, blksize, hs = self.kv_cache_shape + hn = 1 + block_stride = stride[0] + else: + _, blknum, blksize, hn, hs = self.kv_cache_shape""" + h2_new = """ # DSA dual-KV fix: use the PER-LAYER tensor shape, not the global + # self.kv_cache_shape (the DSA indexer cache differs from the main MLA). + _layer_shape = tuple(self.kv_caches[layer_name].shape) + assert len(_layer_shape) > 0, "KV caches shape not initialized" + is_mla = len(_layer_shape) == 3 + stride = self.kv_caches[layer_name].stride() + sz = self.kv_caches[layer_name].element_size() + if is_mla: + blknum, blksize, hs = _layer_shape + hn = 1 + block_stride = stride[0] + else: + _, blknum, blksize, hn, hs = _layer_shape""" + if "_layer_shape = tuple(self.kv_caches[layer_name].shape)" in src: + applied.append("h2 (already)") + elif h2_old in src: + src = src.replace(h2_old, h2_new, 1) + applied.append("h2") + else: + print("[glm-dualkv] WARN: h2 anchor (_compute_block_transfer_offsets head) not found -- skipping h2.") + + # --- Hunk 3: per-layer offsets in _read_blocks ----------------------------- + h3_old = """ first_layer = list(self.layer_name_to_local_kv_cache_metadata.keys())[0] + offs = self._compute_block_transfer_offsets( + first_layer, local_block_ids, remote_block_ids, remote_moriio_meta + ) + + for layer_name in self.layer_name_to_local_kv_cache_metadata: + sess_idx = list(self.layer_name_to_local_kv_cache_metadata.keys()).index( + layer_name + ) + # TODO : apply multi-session batch-read when moriio support it + transfer_status = self.moriio_wrapper.read_remote_data( + offs[2], offs[0], offs[1], sessions[sess_idx] + )""" + h3_new = """ # DSA dual-KV fix: compute offsets PER LAYER (the DSA indexer cache has a + # different per-block size than the main MLA cache, so a single offs reused + # across all layers mis-sizes the indexer transfer -> lost completion notify). + for layer_name in self.layer_name_to_local_kv_cache_metadata: + sess_idx = list(self.layer_name_to_local_kv_cache_metadata.keys()).index( + layer_name + ) + offs = self._compute_block_transfer_offsets( + layer_name, local_block_ids, remote_block_ids, remote_moriio_meta + ) + # TODO : apply multi-session batch-read when moriio support it + transfer_status = self.moriio_wrapper.read_remote_data( + offs[2], offs[0], offs[1], sessions[sess_idx] + )""" + if "compute offsets PER LAYER" in src: + applied.append("h3 (already)") + elif h3_old in src: + src = src.replace(h3_old, h3_new, 1) + applied.append("h3") + else: + print("[glm-dualkv] WARN: h3 anchor (_read_blocks first_layer offsets) not found -- skipping h3.") + + if src != orig: + try: + open(path, "w").write(src) + except OSError as e: + print(f"[glm-dualkv] ERROR: write failed for {path}: {e}", file=sys.stderr) + return 1 + print(f"[glm-dualkv] patched {path} -- hunks: {', '.join(applied)}") + else: + print(f"[glm-dualkv] no changes ({', '.join(applied) or 'nothing applied'}) for {path}") + + # py-compile sanity + try: + import py_compile + py_compile.compile(path, doraise=True) + print("[glm-dualkv] py_compile OK") + except Exception as e: # noqa: BLE001 + print(f"[glm-dualkv] ERROR: patched file fails to compile: {e}", file=sys.stderr) + return 1 + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/vllm_dissag/apply_glm_dsa_moriio_engine_fix.py b/scripts/vllm_dissag/apply_glm_dsa_moriio_engine_fix.py new file mode 100755 index 00000000..9d6d468b --- /dev/null +++ b/scripts/vllm_dissag/apply_glm_dsa_moriio_engine_fix.py @@ -0,0 +1,116 @@ +#!/usr/bin/env python3 +"""Fix MoRIIO WRITE-path per-layer offset caching for GLM-5.1 DSA dual KV cache. + +ROOT CAUSE (proven by instrumentation, job 37594): + GLM-5.1 (GlmMoeDsaForCausalLM) registers num_layers=156 KV caches = 78 main MLA + (per-block latent dim 576) + 78 DSA indexer caches (latent dim 132). TWO geometries. + + moriio_engine.py::MoRIIOEngine._prepare_transfer_plan computes the RDMA transfer + offsets ONCE (for whatever layer arrives first) and caches them on + request_info.transfer_offset, then REUSES that single offset/size tuple for ALL 156 + layers. The 78 indexer layers (dim 132) get written with the main-MLA geometry + (dim 576) -> wrong byte size/offset -> those RDMA writes are malformed; the per-layer + write accounting (writes_done) and/or the remote completion never reconciles -> + the producer's send_notify (gated on writes_done >= num_layers) misbehaves and the + decode side never receives a clean completion -> "Reaped deferred sends / no + finished_sending after 60s" -> request hangs. + +FIX (surgical, no dataclass change): + Cache transfer offsets PER LAYER on the request_info via a dynamically-attached dict + ``_transfer_offset_by_layer`` keyed by layer_name, instead of the single + ``transfer_offset`` slot. Each of the 156 layers then transfers with its OWN geometry + (the underlying _compute_block_transfer_offsets already takes layer_name and, with the + companion dualkv patch h2, reads the per-layer tensor shape). + + Single-geometry models (DeepSeek-V3 / Hunyuan, 1 cache/layer) are unaffected: + every layer has identical geometry, so per-layer caching yields the same offsets. + +Idempotent + anchor-based. A missing anchor warns-and-skips; a found OLD anchor that +fails to apply is a hard error (would silently keep the stall). + +Usage: apply_glm_dsa_moriio_engine_fix.py +""" +import os +import sys + +REL = "distributed/kv_transfer/kv_connector/v1/moriio/moriio_engine.py" + + +def main() -> int: + if len(sys.argv) != 2: + print(f"usage: {sys.argv[0]} ", file=sys.stderr) + return 2 + path = os.path.join(sys.argv[1], REL) + if not os.path.isfile(path): + print(f"[glm-engine] {REL} not found -- skipping (engine layout differs).") + return 0 + + src = open(path).read() + + old = """ # Compute offsets if not cached + if request_info.transfer_offset is None: + offsets = self.worker._compute_block_transfer_offsets( + task.layer_name, + task.local_block_ids, + request_info.block_ids, + remote_moriio_meta, + ) + request_info.transfer_offset = offsets + + # Get session index + layer_names = list(self.worker.layer_name_to_local_kv_cache_metadata.keys()) + sess_idx = layer_names.index(task.layer_name) + + local_off, remote_off, sizes = request_info.transfer_offset""" + + new = """ # DSA dual-KV fix: cache offsets PER LAYER, not once per request. GLM-5.1 has + # two cache geometries (main MLA dim 576 + DSA indexer dim 132); a single + # cached offset reused across all 156 layers mis-sizes the indexer writes and + # the completion never reconciles. Per-layer caching is identical for + # single-geometry models (DeepSeek/Hunyuan). + _off_by_layer = getattr(request_info, "_transfer_offset_by_layer", None) + if _off_by_layer is None: + _off_by_layer = {} + request_info._transfer_offset_by_layer = _off_by_layer + offsets = _off_by_layer.get(task.layer_name) + if offsets is None: + offsets = self.worker._compute_block_transfer_offsets( + task.layer_name, + task.local_block_ids, + request_info.block_ids, + remote_moriio_meta, + ) + _off_by_layer[task.layer_name] = offsets + # keep the legacy single-slot populated (first layer) for any external reader + if request_info.transfer_offset is None: + request_info.transfer_offset = offsets + + # Get session index + layer_names = list(self.worker.layer_name_to_local_kv_cache_metadata.keys()) + sess_idx = layer_names.index(task.layer_name) + + local_off, remote_off, sizes = offsets""" + + if "_transfer_offset_by_layer" in src: + print(f"[glm-engine] already patched (_transfer_offset_by_layer present) -- no-op.") + elif old in src: + src = src.replace(old, new, 1) + open(path, "w").write(src) + print(f"[glm-engine] patched per-layer offset caching in {path}") + else: + print(f"[glm-engine] WARN: anchor (_prepare_transfer_plan offset block) not found -- skipping (engine revision differs).") + # Not fatal: without the anchor we can't safely patch; surface clearly. + return 0 + + try: + import py_compile + py_compile.compile(path, doraise=True) + print("[glm-engine] py_compile OK") + except Exception as e: # noqa: BLE001 + print(f"[glm-engine] ERROR: compile failed: {e}", file=sys.stderr) + return 1 + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/vllm_dissag/apply_glm_dsa_moriio_gate_fix.py b/scripts/vllm_dissag/apply_glm_dsa_moriio_gate_fix.py new file mode 100755 index 00000000..92501ae9 --- /dev/null +++ b/scripts/vllm_dissag/apply_glm_dsa_moriio_gate_fix.py @@ -0,0 +1,133 @@ +#!/usr/bin/env python3 +"""Fix the MoRIIO transfer-completion gate for GLM-5.1 DSA dual KV cache. + +ROOT CAUSE (PROVEN, job 37615 instrumentation): + GLM-5.1 registers num_layers=156 KV caches = 78 main MLA (model.layers.N.self_attn.attn) + + 78 DSA indexer caches (model.layers.N.self_attn.indexer.k_cache). BUT only the 78 + main-MLA layers ever go through the KV-connector save_kv_layer hook -> only they write + -> writes_done caps at 78. The indexer caches are registered (counted in num_layers) + but vLLM NEVER calls save_kv_layer for them (the DSA indexer is a separate attention + component / DeepseekV32IndexerBackend that doesn't use the connector save path; decode + recomputes indexer state from the transferred main latent KV). + + The producer completion gate (moriio_engine.py): + request_info.writes_done += 1 + if request_info.writes_done >= self.worker.num_layers: # 156, never reached + send_notify(...) + caps at writes_done=78 < num_layers=156 -> send_notify NEVER fires -> decode never + gets completion -> "Reaped deferred sends / no finished_sending after 60s" -> stall. + +FIX: + Add self.num_transfer_layers = count of caches that actually transfer (exclude + '.indexer.' caches), with a fallback to num_layers (so single-geometry models - + DeepSeek-V3 / Hunyuan, no indexer - are bit-identical). Gate completion on + num_transfer_layers instead of num_layers. num_layers itself is left unchanged + (it is also used by the Llama-4 per-layer block-window loop, which needs all caches). + +Companion to apply_glm_dsa_moriio_engine_fix.py (per-layer offset caching). This gate +fix is the primary unblocker; the offset fix is correctness insurance for the layers +that DO write (all same geometry here, but harmless). + +Idempotent + anchor-based. Patches BOTH files (connector: define the field; engine: +use it). A found-old-anchor that fails is a hard error. + +Usage: apply_glm_dsa_moriio_gate_fix.py +""" +import os +import sys + +CONN_REL = "distributed/kv_transfer/kv_connector/v1/moriio/moriio_connector.py" +ENG_REL = "distributed/kv_transfer/kv_connector/v1/moriio/moriio_engine.py" + + +def patch_connector(path: str) -> int: + src = open(path).read() + old = " self.num_layers = len(self.kv_caches.keys())" + new = """ self.num_layers = len(self.kv_caches.keys()) + # DSA dual-KV fix: the producer completion gate must count only layers that + # actually transfer via save_kv_layer. GLM-5.1 registers 2 caches/layer (main + # MLA + DSA indexer), but only the main-MLA caches go through save_kv_layer; the + # '.indexer.' caches are registered yet never written. Gating on len(kv_caches) + # would never be reached. Exclude indexer caches; fall back to num_layers for + # single-geometry models (DeepSeek/Hunyuan have no indexer -> identical). + self.num_transfer_layers = ( + len([k for k in self.kv_caches.keys() if ".indexer." not in k]) + or self.num_layers + ) + logger.info( + "[moriio] completion gate: num_transfer_layers=%d (num_layers=%d)", + self.num_transfer_layers, self.num_layers, + )""" + if "self.num_transfer_layers" in src: + print(f"[glm-gate] connector already patched -- no-op.") + return 0 + if old not in src: + print(f"[glm-gate] WARN: connector anchor (num_layers=) not found -- skipping.") + return 0 + src = src.replace(old, new, 1) + open(path, "w").write(src) + print(f"[glm-gate] patched connector: defined num_transfer_layers in {path}") + return 0 + + +def patch_engine(path: str) -> int: + src = open(path).read() + old = " if request_info.writes_done >= self.worker.num_layers:" + new = """ if request_info.writes_done >= getattr( + self.worker, "num_transfer_layers", self.worker.num_layers + ):""" + if 'getattr(\n self.worker, "num_transfer_layers"' in src or "num_transfer_layers" in src: + print(f"[glm-gate] engine already patched -- no-op.") + return 0 + if old not in src: + print(f"[glm-gate] WARN: engine anchor (writes_done gate) not found -- skipping.") + return 0 + src = src.replace(old, new, 1) + open(path, "w").write(src) + print(f"[glm-gate] patched engine: gate on num_transfer_layers in {path}") + return 0 + + +def main() -> int: + if len(sys.argv) != 2: + print(f"usage: {sys.argv[0]} ", file=sys.stderr) + return 2 + base = sys.argv[1] + conn = os.path.join(base, CONN_REL) + eng = os.path.join(base, ENG_REL) + if not os.path.isfile(conn) or not os.path.isfile(eng): + print("[glm-gate] connector/engine not found -- skipping (layout differs).") + return 0 + + # ATOMIC: both halves (connector defines num_transfer_layers, engine gates on + # it) are needed together or not at all. On a restructured image (e.g. mori + # v1.2.1, whose engine replaced the writes_done>=num_layers gate with a sealed + # writes_expected mechanism that already handles hybrid/DSA dual-KV natively), + # the engine anchor is gone. Applying only the connector half would inject a + # dead num_transfer_layers into restructured internals. So if the engine anchor + # is absent, skip BOTH — the native gate already does the right thing. + eng_src = open(eng).read() + eng_gate_present = " if request_info.writes_done >= self.worker.num_layers:" in eng_src + eng_already = "num_transfer_layers" in eng_src + if not eng_gate_present and not eng_already: + print("[glm-gate] engine gate anchor absent (image restructured, e.g. mori " + "v1.2.1 sealed writes_expected) -- skipping BOTH halves (native gate handles DSA).") + return 0 + + rc = patch_connector(conn) or patch_engine(eng) + if rc: + return rc + + try: + import py_compile + py_compile.compile(conn, doraise=True) + py_compile.compile(eng, doraise=True) + print("[glm-gate] py_compile OK (both files)") + except Exception as e: # noqa: BLE001 + print(f"[glm-gate] ERROR: compile failed: {e}", file=sys.stderr) + return 1 + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/vllm_dissag/apply_glm_dsa_moriio_instrument.py b/scripts/vllm_dissag/apply_glm_dsa_moriio_instrument.py new file mode 100755 index 00000000..68e03beb --- /dev/null +++ b/scripts/vllm_dissag/apply_glm_dsa_moriio_instrument.py @@ -0,0 +1,92 @@ +#!/usr/bin/env python3 +"""TEMPORARY instrumentation: does the DSA indexer cache reach the MoRIIO save path? + +Adds logging at two points in moriio_connector.py to settle the dual-KV RCA: + 1. register_kv_caches: log all kv_caches layer names + their shapes + num_layers. + -> shows whether the DeepseekV32IndexerCache is even registered, and its geometry. + 2. _write_blocks_for_req: log each distinct layer_name that actually triggers a write. + -> compare the COUNT/SET of written layers vs num_layers. If indexer layers are in + kv_caches (counted in num_layers) but never written, writes_done can never reach + num_layers -> send_notify never fires -> the stall. + +This is diagnostic only (no behavior change). Remove before any production use. +Idempotent + anchor-safe. + +Usage: apply_glm_dsa_moriio_instrument.py +""" +import os +import sys + +REL = "distributed/kv_transfer/kv_connector/v1/moriio/moriio_connector.py" + + +def main() -> int: + if len(sys.argv) != 2: + print(f"usage: {sys.argv[0]} ", file=sys.stderr) + return 2 + path = os.path.join(sys.argv[1], REL) + if not os.path.isfile(path): + print(f"[glm-instr] {REL} not found -- skipping.") + return 0 + + src = open(path).read() + orig = src + + # --- Point 1: log kv_caches inventory at num_layers assignment ------------- + a1 = " self.num_layers = len(self.kv_caches.keys())" + b1 = """ self.num_layers = len(self.kv_caches.keys()) + # [glm-instr] kv-cache inventory (dual-KV diagnosis) + try: + for _ln, _kv in self.kv_caches.items(): + logger.info("[glm-instr][register] layer=%s shape=%s dtype=%s", + _ln, tuple(_kv.shape), _kv.dtype) + logger.info("[glm-instr][register] num_layers=%d total_kv_caches=%d", + self.num_layers, len(self.kv_caches)) + except Exception as _e: # noqa: BLE001 + logger.info("[glm-instr][register] inventory log failed: %s", _e)""" + if "[glm-instr][register]" in src: + pass + elif a1 in src: + src = src.replace(a1, b1, 1) + else: + print("[glm-instr] WARN: register anchor (num_layers=) not found.") + + # --- Point 2: log each written layer in _write_blocks_for_req ------------- + a2 = " def _write_blocks_for_req(self, req_id: ReqId, meta: ReqMeta, layer_name, kv_layer):" + b2 = (a2 + "\n" + ' # [glm-instr] record which layers actually trigger a KV write\n' + ' try:\n' + ' _seen = getattr(self, "_glm_instr_written_layers", None)\n' + ' if _seen is None:\n' + ' _seen = set(); self._glm_instr_written_layers = _seen\n' + ' if layer_name not in _seen:\n' + ' _seen.add(layer_name)\n' + ' logger.info("[glm-instr][write] NEW layer=%s total_written=%d/%d",\n' + ' layer_name, len(_seen), getattr(self, "num_layers", -1))\n' + ' except Exception as _e: # noqa: BLE001\n' + ' logger.info("[glm-instr][write] log failed: %s", _e)') + if "[glm-instr][write]" in src: + pass + elif a2 in src: + src = src.replace(a2, b2, 1) + else: + print("[glm-instr] WARN: _write_blocks_for_req anchor not found.") + + if src != orig: + open(path, "w").write(src) + print(f"[glm-instr] instrumented {path}") + else: + print(f"[glm-instr] already instrumented / nothing to do for {path}") + + try: + import py_compile + py_compile.compile(path, doraise=True) + print("[glm-instr] py_compile OK") + except Exception as e: # noqa: BLE001 + print(f"[glm-instr] ERROR: compile failed: {e}", file=sys.stderr) + return 1 + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/vllm_dissag/apply_glm_dsa_persistent_kernel_gate_fix.py b/scripts/vllm_dissag/apply_glm_dsa_persistent_kernel_gate_fix.py new file mode 100644 index 00000000..adda36cf --- /dev/null +++ b/scripts/vllm_dissag/apply_glm_dsa_persistent_kernel_gate_fix.py @@ -0,0 +1,129 @@ +#!/usr/bin/env python3 +"""Gate OFF the AITER persistent sparse-MLA kernel for chunked-prefill batches. + +ROOT CAUSE (ROCm/aiter #4076, vLLM #47042 / #47567): + The AITER persistent MLA work-stealing kernel (mla_a8w8_qh16_qseqlen1_gqaratio16_ps, + taken when work_meta_data from get_mla_metadata_v1 is non-None) is NUMERICALLY WRONG + for multi-token (prefill-shaped) batches of qseqlen==1 entries. Pure decode (1 query + token) and fresh single-chunk prefills are correct; the error only appears once a + request becomes a CHUNKED-PREFILL CONTINUATION. The small per-token error COMPOUNDS + through the KV cache across chunked-prefill passes until long-context decode collapses + into repetition/garbage. Failure is gated on CHUNK COUNT, not raw context length + (verified: 22k in 2 chunks = correct, 22k in 3 chunks = garbage). + + On this image (aiter 0.1.16.post3, before the aiter-side kernel fix #3921) GLM-5.1-FP8 + DSA collapses at ~16-18k prompt tokens. The aiter kernel fix is the long-term answer + (AITERKER-132 / aiter #3921); this is the vLLM-side short-term gate (#47567), which + costs ~no perf (decode + single-chunk prefill keep the persistent path). + +FIX (port of vLLM PR #47567, adapted to this image's rocm_aiter_mla_sparse.py::build): + In ROCMAiterMLASparseMetadataBuilder.build(), detect chunked-prefill continuations + (a request with >1 query token this step whose total seq_len exceeds its query_len, + i.e. part of its context was computed in an earlier chunk) and, when ANY request in + the batch is such a continuation: + * skip the get_mla_metadata_v1 persistent-metadata launch, and + * pass work_meta_data=None to the metadata so mla_decode_fwd takes the CORRECT + non-persistent split-KV path. + Decode-only and single-chunk-prefill batches are unchanged (persistent path kept). + + Uses `seg_lengths` (per-request step query lengths, already computed at build() top) + and `common_attn_metadata.seq_lens_cpu[:num_reqs].numpy()` (total seq lens). Both are + present in this image's build(). + +Idempotent + anchor-based + self-skipping. Missing anchor -> warn+skip (safe across +image revisions / if a newer image already carries the aiter kernel fix). A found-old +anchor that fails to apply is a hard error (would silently keep the corruption). + +Usage: apply_glm_dsa_persistent_kernel_gate_fix.py +""" +import os +import sys + +REL = "v1/attention/backends/mla/rocm_aiter_mla_sparse.py" + +# Anchor 1: the persistent-metadata guard. We insert the continuation detection +# just before it and AND it into the condition. +OLD1 = """ if metadata_key != self._prev_metadata_key: + from aiter import get_mla_metadata_v1""" +NEW1 = """ # PERSISTENT-KERNEL GATE (aiter #4076 / vLLM #47567): the persistent + # sparse-MLA work-stealing kernel is numerically wrong for chunked-prefill + # continuation batches; the error compounds and breaks long-context decode. + # Fall back to the correct non-persistent path whenever any request in the + # batch is a chunked-prefill continuation (>1 query token this step AND + # total seq_len > this step's query_len). Decode + single-chunk prefills + # keep the fast persistent path -> no decode-throughput regression. + # Slice to num_reqs and cast to int64 (vLLM #47567 hardening / Rohan138 PR#1) + # so the masks cannot broadcast-mismatch under cudagraph padding. + _step_query_lens = seg_lengths[:num_reqs].astype(np.int64) + _total_seq_lens = common_attn_metadata.seq_lens_cpu[:num_reqs].numpy().astype( + np.int64 + ) + _is_chunked_continuation = (_step_query_lens > 1) & ( + _total_seq_lens > _step_query_lens + ) + _use_persistent = not bool(_is_chunked_continuation.any()) + if _use_persistent and metadata_key != self._prev_metadata_key: + from aiter import get_mla_metadata_v1""" + +# Anchor 2: the metadata construction passes the persistent buffer unconditionally. +# Gate it on _use_persistent. +OLD2 = " work_meta_data=self._mla_work_meta_data," +NEW2 = " work_meta_data=(self._mla_work_meta_data if _use_persistent else None)," + + +def main() -> int: + if len(sys.argv) != 2: + print(f"usage: {sys.argv[0]} ", file=sys.stderr) + return 2 + path = os.path.join(sys.argv[1], REL) + if not os.path.isfile(path): + print(f"[glm-persist] {REL} not found -- skipping (backend layout differs).") + return 0 + + src = open(path).read() + + if "_is_chunked_continuation" in src or "_use_persistent" in src: + print("[glm-persist] already patched (persistent-kernel gate present) -- no-op.") + return 0 + + # Both anchors must be present to apply safely. + if OLD1 not in src: + print("[glm-persist] WARN: persistent-metadata anchor (metadata_key guard) not " + "found -- skipping (image may already carry the aiter kernel fix, or the " + "backend was refactored).") + return 0 + if OLD2 not in src: + print("[glm-persist] ERROR: found the metadata_key guard but NOT the " + "work_meta_data=self._mla_work_meta_data assignment -- refusing partial " + "patch (would leave persistent kernel active). Aborting.", file=sys.stderr) + return 1 + + src = src.replace(OLD1, NEW1, 1) + src = src.replace(OLD2, NEW2, 1) + + try: + open(path, "w").write(src) + except OSError as e: + print(f"[glm-persist] ERROR: write failed for {path}: {e}", file=sys.stderr) + return 1 + + # Verify both edits landed. + chk = open(path).read() + if "_use_persistent = not bool(_is_chunked_continuation.any())" not in chk or \ + "if _use_persistent else None" not in chk: + print("[glm-persist] ERROR: post-write verification failed.", file=sys.stderr) + return 1 + + try: + import py_compile + py_compile.compile(path, doraise=True) + except Exception as e: # noqa: BLE001 + print(f"[glm-persist] ERROR: patched file fails to compile: {e}", file=sys.stderr) + return 1 + + print(f"[glm-persist] patched persistent-kernel gate (aiter #4076 / vLLM #47567) in {path}") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/vllm_dissag/apply_glm_moriio_abort_guard_fix.py b/scripts/vllm_dissag/apply_glm_moriio_abort_guard_fix.py new file mode 100644 index 00000000..c09b4acd --- /dev/null +++ b/scripts/vllm_dissag/apply_glm_moriio_abort_guard_fix.py @@ -0,0 +1,98 @@ +#!/usr/bin/env python3 +"""Guard the MoRIIO connector abort path against a None peer_zmq (mori v1.2.1). + +ROOT CAUSE (observed on the router image, job 199144 decode crash): + When a request is ABORTED before its KV-transfer peer handshake completes, + the connector's release path runs: + + moriio_connector.py::_release_write_prefill_blocks + peer_zmq = get_peer_zmq_from_request_id(request_id, is_producer=False) # -> None + remote_host, _, remote_notify_port = parse_moriio_zmq_address(peer_zmq) # None.split(",") + -> AttributeError: 'NoneType' object has no attribute 'split' + + This only catches ValueError, not the AttributeError from a None peer_zmq, so + the EngineCore dies -> cascades to all decode workers (EngineDeadError) -> decode + is dead. Triggered by any request aborted before the peer handshake (e.g. a + canary/curl that times out during first-token cold JIT). + + The SAME FILE already guards this correctly at the other call site + (request_finished / _should_notify path): `if peer_zmq is not None:` then parse, + else fall back to params. The release path just missed the guard — an + inconsistent-guard bug in the connector. + +FIX (surgical, matches the file's own existing pattern): + In _release_write_prefill_blocks, when the params don't already carry + remote_host/remote_notify_port, guard the peer_zmq lookup: if it is None, log + and return (same graceful bail the existing `except ValueError` already does for + the "missing remote notify address" case). No behavior change when peer_zmq is + valid; single-geometry / non-aborted requests are unaffected. + +Idempotent + anchor-based + self-skipping (no-ops if the anchor is absent/already +guarded, so it is safe across connector revisions and other images). A found-old +anchor that fails to apply is a hard error (would leave the crash). + +Usage: apply_glm_moriio_abort_guard_fix.py +""" +import os +import sys + +REL = "distributed/kv_transfer/kv_connector/v1/moriio/moriio_connector.py" + +# The buggy two lines: fetch peer_zmq (may be None) then parse it unguarded. +OLD = """ peer_zmq = get_peer_zmq_from_request_id(request_id, is_producer=False) + remote_host, _, remote_notify_port = parse_moriio_zmq_address(peer_zmq)""" + +NEW = """ peer_zmq = get_peer_zmq_from_request_id(request_id, is_producer=False) + # Abort-path guard: a request aborted before the KV peer handshake + # has peer_zmq=None; parse_moriio_zmq_address(None) would raise + # AttributeError and kill the EngineCore. Bail gracefully like the + # ValueError case below (matches the guarded call site elsewhere). + if peer_zmq is None: + logger.warning( + "Cannot release WRITE prefill blocks for request %s: " + "no peer zmq address (aborted before peer handshake)", + request_id, + ) + return + remote_host, _, remote_notify_port = parse_moriio_zmq_address(peer_zmq)""" + + +def main() -> int: + if len(sys.argv) != 2: + print(f"usage: {sys.argv[0]} ", file=sys.stderr) + return 2 + path = os.path.join(sys.argv[1], REL) + if not os.path.isfile(path): + print(f"[glm-abort] {REL} not found -- skipping (connector layout differs).") + return 0 + + src = open(path).read() + if "no peer zmq address (aborted before peer handshake)" in src: + print("[glm-abort] already patched -- no-op.") + return 0 + if OLD not in src: + # Anchor absent: either the release path was refactored or this image + # already guards it. Do not block launch. + print("[glm-abort] release-path anchor not found -- skipping (assuming " + "native guard / refactored).") + return 0 + + src = src.replace(OLD, NEW, 1) + try: + open(path, "w").write(src) + except OSError as e: + print(f"[glm-abort] ERROR: write failed for {path}: {e}", file=sys.stderr) + return 1 + + try: + import py_compile + py_compile.compile(path, doraise=True) + except Exception as e: # noqa: BLE001 + print(f"[glm-abort] ERROR: patched file fails to compile: {e}", file=sys.stderr) + return 1 + print(f"[glm-abort] patched _release_write_prefill_blocks None-guard in {path}") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/vllm_dissag/connectors/moriio.sh b/scripts/vllm_dissag/connectors/moriio.sh index 44c64d68..c26a0944 100644 --- a/scripts/vllm_dissag/connectors/moriio.sh +++ b/scripts/vllm_dissag/connectors/moriio.sh @@ -124,13 +124,96 @@ _moriio_build_kv_transfer_config() { } connector_runtime_patch() { - # No-op: the MoRIIO multi-node disagg fixes (vLLM PR#39276 notify-path, #41751 LL - # split, DP-rank hash-failsafe) are committed in-source in the vLLM the image is - # built from (see the Dockerfile VLLM_REF). There is no runtime .py patcher — that - # would be a drifting duplicate of fixes that already live upstream in the fork. - # If you ever run an image WITHOUT these fixes baked, use an image that has them - # (rebuild from the pinned VLLM_REF) rather than patching a stock image at runtime. - return 0 + # MoRIIO multi-node disagg fixes (vLLM PR#39276 notify-path, #41751 LL split, + # DP-rank hash-failsafe) are committed in-source in the vLLM the image is built + # from (Dockerfile VLLM_REF). There is no generic runtime .py patcher for those — + # that would be a drifting duplicate of fixes already upstream in the fork. + # + # EXCEPTION — GLM-5.1-FP8 (GlmMoeDsaForCausalLM, MLA + DSA sparse attention): + # DSA is a NEW attention family the MoRIIO connector was never built for. It adds + # a 2nd KV cache per layer (indexer) with a different geometry, which the + # single-geometry connector mis-handles -> disagg KV transfer stalls; plus a DSA + # invalid-token kernel bug (#45324) that produces `!!!`. These are model-specific + # code gaps, applied here as idempotent, anchor-based, self-skipping .py patchers + # (they no-op cleanly if the fix is native/refactored on the chosen image). Gated + # on MODEL_NAME so DeepSeek/other models are a pure no-op (byte-identical to before). + # The MoRI version is pinned by the Dockerfile MORI_REF (post-1.2.1 main with the + # large-transfer notify/mapping fixes #424/#436/#432 baked in); if a newer MoRI is + # needed, update MORI_REF and rebuild the image — no runtime library swap here. + [ "${MODEL_NAME:-}" = "GLM-5.1-FP8" ] || return 0 + _glm_dsa_runtime_patch +} + +# GLM-5.1 DSA patchers (see connector_runtime_patch). Ported from MAD-private #338. +# Resolves the vLLM install dir, then applies the 4 required patchers in order, +# aborting on a hard failure (a real failure means GLM emits garbage or stalls, so +# failing at launch is correct). Patchers self-skip (rc 0) when their anchor is +# absent, so an image that already carries or refactored a fix no-ops cleanly. +_glm_dsa_runtime_patch() { + # GLM_SKIP_PATCHERS=1: the serving image already carries the GLM-5.1 DSA fixes + # in-source (e.g. the #47766 stack image built from raviguptaamd/vllm@ + # glm5.1-dsa-wideEP_on_shikpate_06_29_customer). Skip ALL runtime patchers — they + # are redundant, and the persistent-gate/sampling-overlay patchers would actively + # REGRESS a baked image (turn persistent MLA off / overwrite stock aiter kernels). + if [ "${GLM_SKIP_PATCHERS:-0}" = "1" ]; then + echo "[glm] GLM_SKIP_PATCHERS=1: image carries DSA fixes in-source; skipping runtime patchers." + return 0 + fi + local _patch_dir="${SCRIPT_DIR:-$(cd "$(dirname "${BASH_SOURCE[0]:-$0}")/.." && pwd)}" + local _vllm_dir + _vllm_dir="$(python3 -c 'import vllm, os; print(os.path.dirname(vllm.__file__))' 2>/dev/null || true)" + if [ -z "${_vllm_dir}" ] || [ ! -d "${_vllm_dir}" ]; then + echo "Error: [glm] cannot locate vLLM install dir for DSA patchers. Aborting." >&2 + exit 1 + fi + echo "[glm] MODEL_NAME=GLM-5.1-FP8: applying DSA runtime patchers against ${_vllm_dir}" + + # Ordered list of REQUIRED patchers (all abort on hard failure). + # GLM_PERSIST_GATE=0 skips the persistent-MLA accuracy gate (debug only: to test + # whether the non-persistent kernel it routes to is what crashes disagg at >=8k). + local _gate_patcher="apply_glm_dsa_persistent_kernel_gate_fix.py" + [ "${GLM_PERSIST_GATE:-1}" = "0" ] && _gate_patcher="" + local _p + for _p in \ + apply_glm_dsa_kernel_fix.py \ + apply_glm_dsa_moriio_dualkv_fix.py \ + apply_glm_dsa_moriio_engine_fix.py \ + apply_glm_dsa_moriio_gate_fix.py \ + apply_glm_moriio_abort_guard_fix.py \ + ${_gate_patcher} \ + apply_glm_aiter_sampling_oob_fix.py; do + local _py="${_patch_dir}/${_p}" + if [ ! -f "${_py}" ]; then + echo "Error: [glm] required patcher ${_py} not found. Aborting." >&2 + exit 1 + fi + echo "[glm] applying ${_p}" + python3 "${_py}" "${_vllm_dir}" 2>&1 || { + echo "Error: [glm] ${_p} failed — GLM-5.1 would emit garbage or stall. Aborting." >&2 + exit 1 + } + done + + # Optional DSA indexer boot-warmup (GLM_INDEXER_WARMUP=1). Force-compiles the DSA + # indexer kernels at boot so they never JIT mid-inference. Opt-in because it drives a + # large (>=8k) prefill forward at boot: on stacks where that forward faults it makes + # the fault DETERMINISTIC at boot (useful for debugging) rather than on first request. + if [ "${GLM_INDEXER_WARMUP:-0}" = "1" ]; then + local _warm="${_patch_dir}/apply_glm_dsa_indexer_warmup_fix.py" + if [ -f "${_warm}" ]; then + echo "[glm] applying DSA indexer boot-warmup (GLM_INDEXER_WARMUP=1)" + python3 "${_warm}" "${_vllm_dir}" 2>&1 || echo "Warning: [glm] indexer-warmup patch failed (non-fatal)." + fi + fi + + # Optional diagnostic instrumentation (GLM_INSTRUMENT=1). Non-fatal. + if [ "${GLM_INSTRUMENT:-0}" = "1" ]; then + local _instr="${_patch_dir}/apply_glm_dsa_moriio_instrument.py" + if [ -f "${_instr}" ]; then + echo "[glm] applying instrumentation (GLM_INSTRUMENT=1): apply_glm_dsa_moriio_instrument.py" + python3 "${_instr}" "${_vllm_dir}" 2>&1 || echo "Warning: [glm] instrumentation failed (non-fatal)." + fi + fi } # connector_launch_worker [dp_start_rank] @@ -219,7 +302,7 @@ connector_launch_worker() { --all2all-backend "${_all2all}" \ --trust-remote-code \ --distributed-timeout-seconds "${DISTRIBUTED_TIMEOUT_SECONDS:-7200}" \ - "${exec_args[@]}" "${extra_args[@]}" "${kv_args[@]}" + "${exec_args[@]}" "${extra_args[@]}" "${kv_args[@]}" "${model_args[@]}" WORKER_PID=0; return 0 fi @@ -242,6 +325,7 @@ connector_launch_worker() { "${exec_args[@]}" \ "${extra_args[@]}" \ "${kv_args[@]}" \ + "${model_args[@]}" \ 2>&1 | tee /run_logs/${SLURM_JOB_ID}/${log_prefix}_NODE${NODE_RANK}.log >/dev/null & WORKER_PID=$! return 0 diff --git a/scripts/vllm_dissag/keepalive_bench.sh b/scripts/vllm_dissag/keepalive_bench.sh new file mode 100755 index 00000000..68d80dcb --- /dev/null +++ b/scripts/vllm_dissag/keepalive_bench.sh @@ -0,0 +1,18 @@ +#!/bin/bash +# Keepalive with LIGHT heartbeat traffic: holds the disagg server up AND sends a +# tiny request every ~20s so prefill discovery/ping stays registered (idle sleep +# lets the prefill ZMQ ping die ~2min in). Runs KEEPALIVE_MINS (default 90). +: "${KEEPALIVE_MINS:=90}" +PORT="${BENCHMARK_PORT:-30000}" +MODEL="/mnt/m2m_nobackup/models_blog/GLM-5.1-FP8" +echo "[keepalive] light-traffic hold ${KEEPALIVE_MINS}min on :${PORT}" +_end=$(( $(date +%s) + KEEPALIVE_MINS*60 )) +i=0 +while [ "$(date +%s)" -lt "$_end" ]; do + curl -s -m 30 "http://127.0.0.1:${PORT}/v1/completions" -H "Content-Type: application/json" \ + -d "{\"model\":\"${MODEL}\",\"prompt\":\"hi\",\"max_tokens\":1,\"temperature\":0}" >/dev/null 2>&1 + i=$((i+1)) + [ $((i % 3)) -eq 0 ] && echo "[keepalive] heartbeat $i, $(( (_end-$(date +%s))/60 ))min left" + sleep 20 +done +echo "[keepalive] done" diff --git a/scripts/vllm_dissag/models.yaml b/scripts/vllm_dissag/models.yaml index 23d66059..9eec3e7d 100644 --- a/scripts/vllm_dissag/models.yaml +++ b/scripts/vllm_dissag/models.yaml @@ -160,10 +160,17 @@ _deepseek_recipe_env: &deepseek_recipe_env # above. No tp: blocks — TP is unsupported for these models. DeepSeek-V3: env: *deepseek_recipe_env + # NEWER-BASE ADAPTATION (isolated to this model): Shiksha's newer vLLM base added + # TritonMLAMetadataBuilder._reserve_attn_logits_workspace(), which pre-reserves the + # decode split-KV logits workspace at WORST CASE + # (max_num_seqs x q_heads x max_kv_splits(max_model_len) x lse_dim x fp32). With + # DSV3's defaults (max_num_seqs=256, max_model_len~163k) this reserves ~128 GiB and + # OOMs at KV init. Cap max-num-seqs + max-model-len to bound the workspace. This is a + # per-model dp: flag (model_args), fully isolated -- does NOT touch GLM/other recipes. prefill: - dp: "" + dp: "--max-num-seqs 64 --max-model-len 32768" decode: - dp: "" + dp: "--max-num-seqs 64 --max-model-len 32768" DeepSeek-V3-5layer: env: *deepseek_recipe_env @@ -178,3 +185,94 @@ DeepSeek-R1: dp: "" decode: dp: "" + +# ============================ MoE + DSA (wideEP only) ============================ + +# GLM-5.1-FP8 (zai-org/GLM-5.1-FP8, arch GlmMoeDsaForCausalLM): MLA + DeepSeek +# Sparse Attention (DSA). 78 layers (3 dense + 75 MoE), 256 routed experts top-8 +# + 1 shared, FP8 block 128. wideEP-only (see WIDE_EP_ONLY_MODELS in the slurm). +# +# Validated DOCKER_IMAGE_NAME (submit-time, not set here — #171 requires it explicit): +# rocmshared/pytorch-private:vllm-wideep_06_29_2026_Shiksha_dp16_2p2d_mori_v1.2.1_aiter_v0.1.16.post3_nightlybase_mori121 +# On this image the DSA patchers self-adapt: only the invalid-token kernel fix +# (#45324) applies; the MoRIIO dual-KV geometry + completion-gate fixes are NATIVE +# (moriio_layout.py per-layer geometry + engine sealed writes_expected), so those +# patchers cleanly no-op. On the older b10a9f7a image all 4 patchers apply. +# +# GLM differs from DeepSeek in 3 recipe-defining ways (both are MLA MoE, but GLM +# is DSA-sparse): +# - KV_BLOCK_SIZE=1 (DSA sparse indexer REQUIRES block-size 1; DS uses 16) +# - VLLM_ROCM_USE_AITER_MLA=1 (GLM MLA path ON via AITER sparse; DS sets 0) +# - prefill EAGER (NONE), decode PIECEWISE (prefill cudagraph capture deadlocks +# on this stack; decode PIECEWISE captures cleanly on DSA and is a ~3.7x ITL win +# (validated: ~69ms vs ~264ms eager). Global VLLM_CUDAGRAPH_MODE=NONE as the +# safe floor; per-role PREFILL=NONE / DECODE=PIECEWISE override it.) +# Note: block=1 + AITER_MLA=1 are ALREADY the moriio.sh connector defaults — GLM +# keeps them; DeepSeek is the one that overrides them off. Set here explicitly so +# the recipe is self-documenting and robust to connector default changes. +# +# DSA adds a 2nd KV cache per layer (indexer); the MoRIIO connector needs the GLM +# DSA patchers (kernel #45324 + dual-KV geometry + per-layer offset + completion +# gate) applied by connector_runtime_patch in connectors/moriio.sh (gated on this +# MODEL_NAME). Those are code patches, not flags — nothing to add here for them. +# +# dp_flags carry the GLM tool/reasoning parsers (AMD GLM recipe). They are applied +# to BOTH roles by compose() in vllm_disagg.sh and reach `vllm serve` via the +# connector's model_args. block-size / kv-cache-dtype / all2all / cudagraph come +# from the env: recipe above (the connector emits them), so the dp: blocks are +# empty like the DeepSeek family. +# +# LONG-CONTEXT CAVEAT (vLLM #40018): the ROCM_AITER_MLA_SPARSE prefill indexer +# corrupts output for prompts beyond ~16-18k tokens on this image (mori v1.2.1) — +# coherent + correct needle retrieval up to 14k, garbage (repetition collapse, +# unique-ratio ~0.1) at ~18.7k. This is an UPSTREAM kernel bug, not the MAD port. +# TESTED: pinning --max-model-len 32768 does NOT move the threshold (workspace is +# sized max_model_len*40 but the corruption onset is a fixed ~18k token count in +# the gather/logits kernel, not a buffer-scaling artifact). So no config knob +# helps; it needs the complete upstream prefill fix in a newer image. Left at the +# native max_model_len (do not cap — capping gives no accuracy benefit and only +# limits usable context). Serve prompts <~14k for correct output on this image. +GLM-5.1-FP8: + env: + VLLM_USE_V1: "1" + VLLM_ROCM_USE_AITER: "1" + VLLM_ROCM_USE_AITER_RMSNORM: "1" + VLLM_ROCM_USE_AITER_MLA: "1" + KV_BLOCK_SIZE: "1" + KV_CACHE_DTYPE: "fp8" + GPU_MEMORY_UTILIZATION: "0.80" + VLLM_CUDAGRAPH_MODE: "NONE" + PREFILL_CUDAGRAPH_MODE: "NONE" + DECODE_CUDAGRAPH_MODE: "PIECEWISE" + CUDAGRAPH_CAPTURE_SIZES: "1 2 4 8 16 32 64 128 256" + VLLM_ALL2ALL_BACKEND: "mori_high_throughput" + PREFILL_MORI_BACKEND: "mori_high_throughput" + DECODE_MORI_BACKEND: "mori_low_latency" + MORI_SHMEM_HEAP_SIZE: "17179869184" + # DSA sparse-indexer logits-buffer cap (crash fix). The indexer prefill computes an + # M*N fp32 logits buffer; split_indexer_prefill_chunks only sub-chunks the query dim + # when M*N*4 > this budget. Default 512MB lets an 8192-token prefill build a single + # 268MB buffer + launch the fp8_mqa_logits kernel at grid=(8192,), which HARD-FAULTS + # the worker on gfx942 (silent GPU fault -> DP group collapse -> 503 at >=8k prompts). + # Capping at 64MB forces M-dim sub-chunking (~2k tokens/chunk) so the buffer and the + # kernel launch stay bounded. Root cause: vllm/v1/attention/ops/triton_fp8_mqa_logits.py + # fp8_mqa_logits_gfx942; chunking logic: mla/indexer.py split_indexer_prefill_chunks. + VLLM_SPARSE_INDEXER_MAX_LOGITS_MB: "64" + # NCCL heartbeat watchdog: at long context (>~8k) a DP rank's sparse-MLA/MoE + # all2all collective can exceed the default HeartbeatMonitor timeout -> + # ProcessGroupNCCL::HeartbeatMonitor::runLoop() declares the rank dead and + # tears down the whole process group -> prefill EngineCore crashes -> 503. + # (Confirmed root cause of the 8k+ prefill crash; #338 EP-landmine.) Disable + # the monitor-triggered teardown and extend timeouts so long-ctx collectives + # complete instead of being watchdog-killed. + TORCH_NCCL_ENABLE_MONITORING: "0" + TORCH_NCCL_HEARTBEAT_TIMEOUT_SEC: "1800" + TORCH_NCCL_DUMP_ON_TIMEOUT: "0" + TORCH_NCCL_BLOCKING_WAIT: "0" + TORCH_NCCL_ASYNC_ERROR_HANDLING: "1" + NCCL_IB_TIMEOUT: "22" + dp_flags: "--tool-call-parser glm47 --reasoning-parser glm45 --enable-auto-tool-choice --chat-template-content-format string" + prefill: + dp: "" + decode: + dp: "" diff --git a/scripts/vllm_dissag/run_xPyD_models.slurm b/scripts/vllm_dissag/run_xPyD_models.slurm index c71fc7e8..98281f57 100755 --- a/scripts/vllm_dissag/run_xPyD_models.slurm +++ b/scripts/vllm_dissag/run_xPyD_models.slurm @@ -68,6 +68,22 @@ for f in "${REQUIRED_FILES[@]}"; do done echo "Running from: $(pwd)" +# ------------------------------------------------------------------------------ +# models.yaml env precedence: capture which recipe knobs the USER explicitly set +# at submit time. The driver (vllm_disagg.sh) uses this to let models.yaml `env:` +# OVERRIDE image-baked ENV defaults (e.g. a DeepSeek-tuned image bakes +# KV_BLOCK_SIZE=16 / VLLM_ROCM_USE_AITER_MLA=0, which would otherwise shadow a +# model's own recipe — GLM-5.1 DSA needs block=1 + AITER MLA on), while a genuine +# submit-time `-e VAR=...` still wins. Precedence: image-baked < models.yaml < submit -e. +# Captured HERE (before the slurm sets any defaults) so it reflects user intent only. +_RECIPE_ENV_KEYS="VLLM_USE_V1 VLLM_ROCM_USE_AITER VLLM_ROCM_USE_AITER_RMSNORM VLLM_ROCM_USE_AITER_MLA KV_BLOCK_SIZE KV_CACHE_DTYPE KV_CACHE_MEMORY_BYTES GPU_MEMORY_UTILIZATION VLLM_CUDAGRAPH_MODE PREFILL_CUDAGRAPH_MODE DECODE_CUDAGRAPH_MODE CUDAGRAPH_CAPTURE_SIZES VLLM_ALL2ALL_BACKEND PREFILL_MORI_BACKEND DECODE_MORI_BACKEND MORI_SHMEM_HEAP_SIZE" +MODELS_YAML_PROTECT="" +for _k in $_RECIPE_ENV_KEYS; do + [ -n "${!_k+x}" ] && MODELS_YAML_PROTECT="${MODELS_YAML_PROTECT} ${_k}" +done +export MODELS_YAML_PROTECT="${MODELS_YAML_PROTECT# }" +echo "models.yaml protect-list (submit-time overrides): '${MODELS_YAML_PROTECT}'" + # ------------------------ # Print current time in UTC and PST formats # ------------------------ @@ -87,6 +103,7 @@ VALID_MODELS=( \ "DeepSeek-R1" \ "Qwen3-32B" \ "Qwen3-30B-A3B" \ + "GLM-5.1-FP8" \ ) # Models allowed for CONNECTOR=moriio WIDE_EP=1 (MoRI-EP; legacy RUN_MORI=1) @@ -94,6 +111,7 @@ MORI_EP_VALID_MODELS=( \ "DeepSeek-V3" \ "DeepSeek-V3-5layer" \ "DeepSeek-R1" \ + "GLM-5.1-FP8" \ ) # Models allowed for CONNECTOR=rixl WIDE_EP=1 EP_BACKEND=deepep (legacy RUN_DEEPEP=1) @@ -179,7 +197,10 @@ WIDE_EP="${WIDE_EP:-0}" # the MoRI-EP / DeepEP recipe (block=16, MLA off, per-role cudagraph). Running them # in TP mode is unsupported — the TP argv would double the model's own # --compilation-config and drop the mandatory +quant_fp8 op. Reject early. -WIDE_EP_ONLY_MODELS=( "DeepSeek-V3" "DeepSeek-V3-5layer" "DeepSeek-R1" ) +# GLM-5.1-FP8 (GlmMoeDsaForCausalLM, MLA+DSA) is validated only under MoRI-EP +# wideEP disagg (block=1, AITER sparse MLA on, per-role all2all). The moriio+TP +# ("Stage B") path is untested for DSA, so reject WIDE_EP=0 for it too. +WIDE_EP_ONLY_MODELS=( "DeepSeek-V3" "DeepSeek-V3-5layer" "DeepSeek-R1" "GLM-5.1-FP8" ) model_is_wide_ep_only() { local m="$1" for x in "${WIDE_EP_ONLY_MODELS[@]}"; do [[ "$m" == "$x" ]] && return 0; done @@ -427,11 +448,14 @@ BENCHMARK_COMBINATIONS="${BENCHMARK_COMBINATIONS:-}" # Benchmark script selector: BENCHMARK_SCRIPT tag -> file run by the launcher. # sweep (default) -> benchmark_xPyD.sh (general concurrency sweep) # long_context -> benchmark_long_context.sh (per-shape warmup, c=1-first) +# keepalive -> keepalive_bench.sh (hold server up KEEPALIVE_MINS +# for external accuracy probes) BENCHMARK_SCRIPT="${BENCHMARK_SCRIPT:-sweep}" case "$BENCHMARK_SCRIPT" in sweep) BENCHMARK_SCRIPT_FILE="benchmark_xPyD.sh" ;; long_context) BENCHMARK_SCRIPT_FILE="benchmark_long_context.sh" ;; - *) echo "Error: invalid BENCHMARK_SCRIPT='$BENCHMARK_SCRIPT' (valid: sweep, long_context)" >&2; exit 1 ;; + keepalive) BENCHMARK_SCRIPT_FILE="keepalive_bench.sh" ;; + *) echo "Error: invalid BENCHMARK_SCRIPT='$BENCHMARK_SCRIPT' (valid: sweep, long_context, keepalive)" >&2; exit 1 ;; esac if [[ ! -f "$BENCHMARK_SCRIPT_FILE" ]]; then echo "Error: selected benchmark script '$BENCHMARK_SCRIPT_FILE' not found in $(pwd)." >&2 @@ -583,6 +607,9 @@ docker run --rm \ ${VLLM_ALL2ALL_BACKEND:+-e VLLM_ALL2ALL_BACKEND=$VLLM_ALL2ALL_BACKEND} \ ${PREFILL_MORI_BACKEND:+-e PREFILL_MORI_BACKEND=$PREFILL_MORI_BACKEND} \ ${DECODE_MORI_BACKEND:+-e DECODE_MORI_BACKEND=$DECODE_MORI_BACKEND} \ + -e MODELS_YAML_PROTECT="${MODELS_YAML_PROTECT:-}" \ + ${GLM_PERSIST_GATE:+-e GLM_PERSIST_GATE=$GLM_PERSIST_GATE} \ + ${GLM_SKIP_PATCHERS:+-e GLM_SKIP_PATCHERS=$GLM_SKIP_PATCHERS} \ ${KV_BLOCK_SIZE:+-e KV_BLOCK_SIZE=$KV_BLOCK_SIZE} \ ${KV_CACHE_MEMORY_BYTES:+-e KV_CACHE_MEMORY_BYTES=$KV_CACHE_MEMORY_BYTES} \ ${VLLM_ROCM_USE_AITER_MLA:+-e VLLM_ROCM_USE_AITER_MLA=$VLLM_ROCM_USE_AITER_MLA} \ @@ -590,6 +617,7 @@ docker run --rm \ ${KV_CACHE_DTYPE:+-e KV_CACHE_DTYPE=$KV_CACHE_DTYPE} \ ${MORIIO_TOY_PROXY:+-e MORIIO_TOY_PROXY=$MORIIO_TOY_PROXY} \ ${BENCHMARK_SCRIPT_FILE:+-e BENCHMARK_SCRIPT_FILE=$BENCHMARK_SCRIPT_FILE} \ + ${KEEPALIVE_MINS:+-e KEEPALIVE_MINS=$KEEPALIVE_MINS} \ ${PREFILL_CUDAGRAPH_MODE:+-e PREFILL_CUDAGRAPH_MODE=$PREFILL_CUDAGRAPH_MODE} \ ${DECODE_CUDAGRAPH_MODE:+-e DECODE_CUDAGRAPH_MODE=$DECODE_CUDAGRAPH_MODE} \ ${CUDAGRAPH_CAPTURE_SIZES:+-e CUDAGRAPH_CAPTURE_SIZES="$CUDAGRAPH_CAPTURE_SIZES"} \ diff --git a/scripts/vllm_dissag/vllm_disagg.sh b/scripts/vllm_dissag/vllm_disagg.sh index 06fbf84f..acbabbd4 100755 --- a/scripts/vllm_dissag/vllm_disagg.sh +++ b/scripts/vllm_dissag/vllm_disagg.sh @@ -160,17 +160,33 @@ MODEL_CONFIG_DECODE="" if [[ -n "$MODEL_NAME" && -f "$MODELS_YAML" ]]; then export MODELS_YAML MODEL_NAME PARALLEL_MODE # 1) Export per-model env: block FIRST (so connector ${VAR:-default} yields to it). - # Only set a var that is NOT already in the environment, so a submit-time - # `docker -e VAR=...` (already exported) WINS over the yaml value. Precedence: - # connector default < models.yaml env: < submit-time -e. + # Precedence: image-baked ENV < models.yaml env: < submit-time -e. + # models.yaml MUST override image-baked ENV: a DeepSeek-tuned disagg image + # bakes KV_BLOCK_SIZE=16 / VLLM_ROCM_USE_AITER_MLA=0 / VLLM_CUDAGRAPH_MODE= + # PIECEWISE etc. as container ENV, which would otherwise shadow a model's own + # recipe (GLM-5.1 DSA needs block=1 + AITER sparse MLA on). But a genuine + # submit-time `-e VAR=...` must still win. The slurm can tell the two apart + # (it runs on the host) and passes MODELS_YAML_PROTECT = the space-separated + # list of keys the USER set at submit; the driver protects only those. When + # MODELS_YAML_PROTECT is unset (script run directly, no slurm), fall back to + # the old "skip if in env" behavior so nothing regresses. _yaml_env="$(python3 - <<'PY' import os, yaml, shlex m = yaml.safe_load(open(os.environ["MODELS_YAML"])) or {} cfg = m.get(os.environ["MODEL_NAME"]) or {} +protect_raw = os.environ.get("MODELS_YAML_PROTECT") +have_protect = protect_raw is not None +protect = set((protect_raw or "").split()) for k, v in (cfg.get("env") or {}).items(): - # skip if already present in the environment (submit-time -e override wins) - if k in os.environ: - continue + if have_protect: + # 3-tier: yaml overrides baked ENV; only a user submit-time -e (in the + # protect-list) wins over yaml. + if k in protect: + continue + else: + # No protect-list (direct run): legacy behavior — any existing env wins. + if k in os.environ: + continue print(f'export {k}={shlex.quote(str(v))}') PY )" From 740898414054022a40339acfeb5720b4e47fbe36 Mon Sep 17 00:00:00 2001 From: raviguptaamd Date: Wed, 8 Jul 2026 15:51:15 +0000 Subject: [PATCH 02/41] vllm_dissag: make NIAH harness thinking-model-aware (GLM-5.1) benchmark_niah.py mis-scored thinking models: it never disabled thinking and read only content + reasoning_content. GLM-5.1 emits chain-of-thought into the `reasoning` field and leaves `content` empty until the final answer, so with a small max_tokens the answer never lands in content -> a false 0/10 even when generation is correct. - Add chat_template_kwargs.enable_thinking=false so the answer goes to content. - Also score the `reasoning` field as a fallback. Verified against GLM-5.1-FP8: correct 9-10/10 retrieval across 2k-35k on all tested topologies (EP8/EP16/EP32) after the fix. Co-Authored-By: Claude --- scripts/vllm_dissag/benchmark_niah.py | 11 ++++++++++- 1 file changed, 10 insertions(+), 1 deletion(-) diff --git a/scripts/vllm_dissag/benchmark_niah.py b/scripts/vllm_dissag/benchmark_niah.py index 0cdd027e..b1b0264d 100755 --- a/scripts/vllm_dissag/benchmark_niah.py +++ b/scripts/vllm_dissag/benchmark_niah.py @@ -50,6 +50,11 @@ def run(n_words): ], "temperature": 0.0, "max_tokens": MAXTOK, + # Thinking models (e.g. GLM-5.1) emit chain-of-thought into a separate + # reasoning field and leave `content` empty until the final answer; with a + # small max_tokens the answer never appears in `content` and the score is a + # false 0/10. Disable thinking so the answer lands in `content` directly. + "chat_template_kwargs": {"enable_thinking": False}, } data = json.dumps(body).encode() req = urllib.request.Request(URL, data=data, headers={"Content-Type": "application/json"}) @@ -59,7 +64,11 @@ def run(n_words): except Exception as e: print("words=%6d ERROR %s" % (n_words, e), flush=True) return None - text = ((msg.get("content") or "") + " " + (msg.get("reasoning_content") or "")).lower() + # Score content plus any reasoning field (some servers surface CoT as + # `reasoning` or `reasoning_content`) so a thinking model is never mis-scored. + text = ((msg.get("content") or "") + " " + + (msg.get("reasoning_content") or "") + " " + + (msg.get("reasoning") or "")).lower() found = sorted(a for a in ANIMALS if a in text) print("words=%6d found=%2d/10 %s" % (n_words, len(found), found), flush=True) return len(found) From 9cfe987385eb4bc5da66c20fa34040c688cef0b1 Mon Sep 17 00:00:00 2001 From: raviguptaamd Date: Wed, 8 Jul 2026 16:16:52 +0000 Subject: [PATCH 03/41] vllm_dissag: NIAH multi-seed support (NIAH_SEEDS) for variance-aware accuracy Needle layout is seeded, so a single run is deterministic (bit-exact on the same stack) but can't tell a real accuracy dip from single-needle variance. Add NIAH_SEEDS (default 0,1,2) to run each context length across multiple needle layouts; the summary now reports mean/min/max across seeds. Backward compatible: NIAH_SEEDS=0 reproduces the prior single-seed behavior. Co-Authored-By: Claude --- scripts/vllm_dissag/benchmark_niah.py | 29 ++++++++++++++++++--------- 1 file changed, 20 insertions(+), 9 deletions(-) diff --git a/scripts/vllm_dissag/benchmark_niah.py b/scripts/vllm_dissag/benchmark_niah.py index b1b0264d..82a55663 100755 --- a/scripts/vllm_dissag/benchmark_niah.py +++ b/scripts/vllm_dissag/benchmark_niah.py @@ -8,6 +8,8 @@ # NIAH_MODEL model name/tag the server serves (required — the served path) # NIAH_WORDS comma list of context sizes in words (default 2000,8000,20000,35000) # NIAH_MAXTOK max_tokens for the answer (default 2048) +# NIAH_SEEDS comma list of needle-layout seeds (default 0,1,2); summary reports +# mean/min/max across seeds to separate real accuracy from variance # NIAH_TIMEOUT per-request timeout seconds (default 1800) import os, sys, json, random, urllib.request @@ -16,6 +18,10 @@ WORDS = [int(x) for x in os.environ.get("NIAH_WORDS", "2000,8000,20000,35000").split(",") if x.strip()] MAXTOK = int(os.environ.get("NIAH_MAXTOK", "2048")) TIMEOUT = float(os.environ.get("NIAH_TIMEOUT", "1800")) +# Needle layout is seeded, so a single run is deterministic (bit-exact repro on the +# same stack). Run multiple seeds to distinguish real accuracy from single-needle +# variance; the summary reports mean/min/max across seeds. Default 0,1,2. +SEEDS = [int(x) for x in os.environ.get("NIAH_SEEDS", "0,1,2").split(",") if x.strip()] FILLER = ( "table chair window bottle pencil garden river mountain coffee planet " @@ -41,12 +47,12 @@ def make_haystack(n_words, seed=0): return " ".join(words) -def run(n_words): +def run(n_words, seed=0): body = { "model": MODEL, "messages": [ {"role": "system", "content": SYSTEM}, - {"role": "user", "content": "Find the animals in this list:\n\n" + make_haystack(n_words)}, + {"role": "user", "content": "Find the animals in this list:\n\n" + make_haystack(n_words, seed)}, ], "temperature": 0.0, "max_tokens": MAXTOK, @@ -70,7 +76,7 @@ def run(n_words): + (msg.get("reasoning_content") or "") + " " + (msg.get("reasoning") or "")).lower() found = sorted(a for a in ANIMALS if a in text) - print("words=%6d found=%2d/10 %s" % (n_words, len(found), found), flush=True) + print("words=%6d seed=%d found=%2d/10 %s" % (n_words, seed, len(found), found), flush=True) return len(found) @@ -79,14 +85,19 @@ def main(): print("NIAH_MODEL must be set (the served model path/name)", file=sys.stderr) sys.exit(2) print("=== NIAH retrieval test ===", flush=True) - print("url=%s model=%s sizes=%s" % (URL, MODEL, WORDS), flush=True) - results = {} + print("url=%s model=%s sizes=%s seeds=%s" % (URL, MODEL, WORDS, SEEDS), flush=True) + results = {} # n_words -> list of scores across seeds (None on error) for n in WORDS: - results[n] = run(n) - print("=== NIAH summary ===", flush=True) + results[n] = [run(n, s) for s in SEEDS] + print("=== NIAH summary (mean/min/max across %d seed(s)) ===" % len(SEEDS), flush=True) for n in WORDS: - v = results[n] - print(" words=%6d found=%s/10" % (n, "ERR" if v is None else v), flush=True) + vals = [v for v in results[n] if v is not None] + if not vals: + print(" words=%6d ERR" % n, flush=True) + continue + mean = sum(vals) / len(vals) + print(" words=%6d mean=%.1f/10 min=%d max=%d (n=%d)" + % (n, mean, min(vals), max(vals), len(vals)), flush=True) if __name__ == "__main__": From 74c69b7f40b3e5cb0f9769cfe56029d3a2bd6d4f Mon Sep 17 00:00:00 2001 From: raviguptaamd Date: Fri, 10 Jul 2026 07:09:21 +0000 Subject: [PATCH 04/41] vllm_dissag: NIAH gate robust to cold-start JIT (warmup + readiness probe) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit On a freshly-booted node the first request of each context shape pays the full JIT/kernel-autotune compile (minutes). The NIAH harness scored the FIRST request, so cold compile landed on a scored/gated request -> false 0/10 or timeout, failing the accuracy gate and skipping the perf sweep. Root-caused by reproducing on a cold boot (0 results) vs a warm server (all pass) on the same image. Fixes: - benchmark_niah.py: add a warmup pass (NIAH_WARMUP=1 default) — one throwaway request per context length before scoring, with a generous timeout, failures tolerated. Scored requests are then always warm. - benchmark_niah.py: distinguish TIMEOUT/ERROR from a wrong answer. Timeouts return a sentinel (excluded from mean, never counted as 0/10); summary flags NO-RESULT with guidance instead of silently reporting 0. - benchmark_niah.sh: replace the blind `sleep 10` with a /v1/models readiness poll (up to 5 min), and forward NIAH_WARMUP. Verified: patched harness on the warm server passes 10/10; cold-boot repro no longer produces false 0/10 because compile happens in the warmup pass. Co-Authored-By: Claude --- scripts/vllm_dissag/benchmark_niah.py | 59 ++++++++++++++++++++++----- scripts/vllm_dissag/benchmark_niah.sh | 18 +++++++- 2 files changed, 64 insertions(+), 13 deletions(-) diff --git a/scripts/vllm_dissag/benchmark_niah.py b/scripts/vllm_dissag/benchmark_niah.py index 82a55663..43d81850 100755 --- a/scripts/vllm_dissag/benchmark_niah.py +++ b/scripts/vllm_dissag/benchmark_niah.py @@ -11,6 +11,12 @@ # NIAH_SEEDS comma list of needle-layout seeds (default 0,1,2); summary reports # mean/min/max across seeds to separate real accuracy from variance # NIAH_TIMEOUT per-request timeout seconds (default 1800) +# NIAH_WARMUP 1 (default) = send one throwaway request per context length BEFORE +# scoring, so the first-hit JIT/kernel-autotune compile happens outside +# the scored/gated window. On a freshly-booted node the first request of +# a shape can take minutes to compile; without warmup that lands on the +# first scored request -> false 0/10 or timeout. Warmup failures are +# tolerated (logged, not fatal). Set 0 to disable. import os, sys, json, random, urllib.request URL = os.environ.get("NIAH_URL", "http://127.0.0.1:30000/v1/chat/completions") @@ -22,6 +28,10 @@ # same stack). Run multiple seeds to distinguish real accuracy from single-needle # variance; the summary reports mean/min/max across seeds. Default 0,1,2. SEEDS = [int(x) for x in os.environ.get("NIAH_SEEDS", "0,1,2").split(",") if x.strip()] +WARMUP = os.environ.get("NIAH_WARMUP", "1") == "1" +# Warmup uses a generous timeout (cold compile of a long-context shape can take minutes) +# and never fails the run — its only job is to trigger compilation before scoring. +WARMUP_TIMEOUT = max(TIMEOUT, 1800.0) FILLER = ( "table chair window bottle pencil garden river mountain coffee planet " @@ -47,7 +57,8 @@ def make_haystack(n_words, seed=0): return " ".join(words) -def run(n_words, seed=0): +def _request(n_words, seed, max_tokens, timeout): + """POST one NIAH request; return (message_dict, error_str). Exactly one is non-None.""" body = { "model": MODEL, "messages": [ @@ -55,7 +66,7 @@ def run(n_words, seed=0): {"role": "user", "content": "Find the animals in this list:\n\n" + make_haystack(n_words, seed)}, ], "temperature": 0.0, - "max_tokens": MAXTOK, + "max_tokens": max_tokens, # Thinking models (e.g. GLM-5.1) emit chain-of-thought into a separate # reasoning field and leave `content` empty until the final answer; with a # small max_tokens the answer never appears in `content` and the score is a @@ -65,10 +76,26 @@ def run(n_words, seed=0): data = json.dumps(body).encode() req = urllib.request.Request(URL, data=data, headers={"Content-Type": "application/json"}) try: - with urllib.request.urlopen(req, timeout=TIMEOUT) as r: - msg = json.loads(r.read())["choices"][0]["message"] + with urllib.request.urlopen(req, timeout=timeout) as r: + return json.loads(r.read())["choices"][0]["message"], None except Exception as e: - print("words=%6d ERROR %s" % (n_words, e), flush=True) + return None, str(e) + + +def warmup(n_words): + """One throwaway request per length so first-hit compile happens off the scored path. + Never fatal: a warmup timeout just means the shape is still compiling; the scored + request will pay whatever remains (bounded by NIAH_TIMEOUT).""" + _, err = _request(n_words, seed=0, max_tokens=8, timeout=WARMUP_TIMEOUT) + status = "ok" if err is None else ("timeout/err: %s" % err) + print("words=%6d [warmup] %s" % (n_words, status), flush=True) + + +def run(n_words, seed=0): + # Sentinel: None = timeout/transport error (NOT a wrong answer); int = score 0..10. + msg, err = _request(n_words, seed, MAXTOK, TIMEOUT) + if err is not None: + print("words=%6d seed=%d TIMEOUT/ERROR %s" % (n_words, seed, err), flush=True) return None # Score content plus any reasoning field (some servers surface CoT as # `reasoning` or `reasoning_content`) so a thinking model is never mis-scored. @@ -85,19 +112,29 @@ def main(): print("NIAH_MODEL must be set (the served model path/name)", file=sys.stderr) sys.exit(2) print("=== NIAH retrieval test ===", flush=True) - print("url=%s model=%s sizes=%s seeds=%s" % (URL, MODEL, WORDS, SEEDS), flush=True) - results = {} # n_words -> list of scores across seeds (None on error) + print("url=%s model=%s sizes=%s seeds=%s warmup=%s" % (URL, MODEL, WORDS, SEEDS, WARMUP), flush=True) + # Warmup pass: compile every shape once before scoring, so cold JIT never lands on a + # scored/gated request (the common cause of false 0/10 or timeout on a fresh boot). + if WARMUP: + print("=== NIAH warmup (one throwaway request per length) ===", flush=True) + for n in WORDS: + warmup(n) + results = {} # n_words -> list of scores across seeds (None = timeout/error, not a wrong answer) for n in WORDS: results[n] = [run(n, s) for s in SEEDS] print("=== NIAH summary (mean/min/max across %d seed(s)) ===" % len(SEEDS), flush=True) for n in WORDS: - vals = [v for v in results[n] if v is not None] + scored = results[n] + vals = [v for v in scored if v is not None] + n_to = sum(1 for v in scored if v is None) # timeouts/errors, excluded from mean if not vals: - print(" words=%6d ERR" % n, flush=True) + print(" words=%6d NO-RESULT (%d/%d timed out or errored — likely cold compile; " + "raise NIAH_TIMEOUT or keep NIAH_WARMUP=1)" % (n, n_to, len(scored)), flush=True) continue mean = sum(vals) / len(vals) - print(" words=%6d mean=%.1f/10 min=%d max=%d (n=%d)" - % (n, mean, min(vals), max(vals), len(vals)), flush=True) + extra = (" [%d timeout/err excluded]" % n_to) if n_to else "" + print(" words=%6d mean=%.1f/10 min=%d max=%d (n=%d)%s" + % (n, mean, min(vals), max(vals), len(vals), extra), flush=True) if __name__ == "__main__": diff --git a/scripts/vllm_dissag/benchmark_niah.sh b/scripts/vllm_dissag/benchmark_niah.sh index ba49a359..d366b5e4 100755 --- a/scripts/vllm_dissag/benchmark_niah.sh +++ b/scripts/vllm_dissag/benchmark_niah.sh @@ -14,15 +14,29 @@ LOG="/run_logs/${SLURM_JOB_ID}/niah_${SLURM_JOB_ID}_${timestamp}_xP${xP}_yD${yD} echo "==== NIAH long-context retrieval test ====" echo "port=${BENCHMARK_PORT} model=${MODEL_PATH} sizes=${NIAH_WORDS:-2000,8000,20000,35000}" -# Give the router a moment to be fully ready for chat completions. -sleep 10 +# Wait until the router actually serves before starting (replaces a blind sleep). On a +# fresh boot the router may register a few seconds after the workers report ready; poll +# /v1/models until it answers, up to ~5 min. Non-fatal: fall through if the probe can't +# confirm (the harness's own warmup + timeout still protect the run). +_ready=0 +for _i in $(seq 1 60); do + if curl -s -o /dev/null -w '%{http_code}' --max-time 5 \ + "http://127.0.0.1:${BENCHMARK_PORT}/v1/models" 2>/dev/null | grep -q '^200$'; then + _ready=1; echo "[niah] router ready after ~$((_i*5))s"; break + fi + sleep 5 +done +[ "$_ready" = 1 ] || echo "[niah] WARN: router readiness not confirmed in 300s; proceeding (warmup will absorb)" # The server registers the model under its path (served_model_name = MODEL_PATH). +# NIAH_WARMUP=1 (harness default): first-hit JIT compiles off the scored path so a cold +# boot does not produce false 0/10 or timeouts on the first scored request. NIAH_URL="http://127.0.0.1:${BENCHMARK_PORT}/v1/chat/completions" \ NIAH_MODEL="${MODEL_PATH}" \ NIAH_WORDS="${NIAH_WORDS:-2000,8000,20000,35000}" \ NIAH_MAXTOK="${NIAH_MAXTOK:-2048}" \ NIAH_TIMEOUT="${NIAH_TIMEOUT:-1800}" \ +NIAH_WARMUP="${NIAH_WARMUP:-1}" \ python3 "${DIR}/benchmark_niah.py" 2>&1 | tee -a "${LOG}" echo "NIAH results -> ${LOG}" From 8d0ecabb999dffa8135146522210ebf8dbc418ba Mon Sep 17 00:00:00 2001 From: raviguptaamd Date: Sun, 16 Aug 2026 04:17:14 +0000 Subject: [PATCH 05/41] [GLM-5.1] v0.27 level-set: 3.4x decode speedup, 200K context, RDMA + launcher fixes Validated on MI300X, 8 nodes, image rocmshared/pytorch-private:glm5.1-vllm027-b8 (base ci_base-dedbf6be8b + vLLM raviguptaamd/vllm@glm5.1-dsa-wideEP_on_vllm-v0.27 + aiter e03fa6040 + MoRI 42e895472b08 + router #181). PERF FIX (models.yaml decode.dp) -- the headline change: --max-num-batched-tokens 2048 on the DECODE role only. max_num_batched_tokens is a chunked-prefill SCHEDULER knob, but it also sizes the MoRI EP dispatch buffer (fused_moe/layer.py -> all2all_utils.py -> all2all.py max_num_inp_token_per_rank). At the 8192 default a decode instance ran an 8192-token-wide all2all every step, per layer, x78 layers, while decoding a handful of tokens: a fixed ~302ms/step floor, ~320x this model's HBM-bandwidth bound. Prefill keeps 8192 (it genuinely dispatches wide batches). 1024/64 con=8, warm: TPOT TTFT out tok/s 1P/1D 302 -> 88.0 ms 2431 -> 906 ms 24.9 -> 78.8 2P/2D 302 -> 94.1 ms 1633 ms 66.7 Published reference: 1P/1D ~89ms, 2P/2D ~91ms -> matched within 3%. Accuracy unaffected: NIAH 2k-200k clean on both topologies, no length collapse, memfault=0, latencies equal-or-better at every length. 200K validated (5.7x beyond the previously published 35K ceiling). Dockerfile: base -> ci_base-dedbf6be8b (matches the fork's upstream base), VLLM_REF -> the v0.27 branch, and WITH_MORI_BUILD/WITH_AITER_BUILD now default to 1 so a plain `docker build` reproduces the validated stack. Previously they defaulted to 0, which silently used the base's bundled aiter 0.1.19 -- that GPU-faults on the GLM DSA decode kernel. The pinned aiter e03fa6040 / MoRI 42e895472b08 must not be bumped without re-running long-context NIAH. connectors/moriio.sh: per-role env split (PREFILL_*/DECODE_* -> VLLM_MORI_*), mirroring the existing PREFILL/DECODE_MORI_BACKEND pattern -- models.yaml env: applies to BOTH roles, but prefill and decode need opposite values here. Also injects use_inductor_graph_partition (pairs with the vLLM splitting_ops fix). connectors/moriio.env: RDMA fabric -- MORI_IB_GID_INDEX=3 (RoCEv2 IPv4), MORI_RDMA_DEVICES/NCCL_IB_HCA restricted to the 8 GPU-local NICs (leaving the mgmt NICs in makes QPs form over a non-routable fabric -> ibverbs.cpp:189 timeouts), NCCL/GLOO control sockets on eth0. run_xPyD_models.slurm: libionic bind-mount requires a regular file after symlink resolve (a dangling symlink gave "OCI runtime create ... not a directory", container exit 125); prefer FABRIC_SUBNET over `hostname -I` first IP (nodes list a 10.224 overlay first, which made the socket_barrier advertise an unreachable NIC -> "Waiting for nodes" hang); GLM_KERNEL_PATCH/GLM_BACKEND_PATCH bind-mount hooks to test .py fixes without a rebuild; forward the new per-role env keys. vllm_disagg.sh: same FABRIC_SUBNET IP-selection fix for host_ip. benchmark_xPyD.sh: per-shape warmup at the REAL isl/osl before each shape's cells. The global warmup is isl=osl=32/con=1, which never exercises a 1024/8192/28672 prefill path or the decode cudagraph batch sizes, so the first measured cell absorbed residual JIT (observed 302ms vs ~88ms steady-state). Warmup output goes to a separate _SHAPEWARMUP.log so it cannot pollute the CSV. models.yaml (GLM-5.1-FP8): decode.dp perf fix above; recipe = prefill eager + mori_high_throughput, decode PIECEWISE cudagraph + mori_low_latency; VLLM_USE_LAYERNAME=0; VLLM_SPARSE_INDEXER_MAX_LOGITS_MB=64; NCCL heartbeat/timeout knobs for long-context collectives. Full operational playbook (including the dead ends) in skills_vllm_disagg.md. Co-Authored-By: Claude --- ...gg_inference.glmv5.1.ubuntu.amd.Dockerfile | 130 ++++++++++-------- scripts/vllm_dissag/benchmark_xPyD.sh | 26 ++++ scripts/vllm_dissag/connectors/moriio.env | 13 +- scripts/vllm_dissag/connectors/moriio.sh | 32 ++++- scripts/vllm_dissag/models.yaml | 34 ++++- scripts/vllm_dissag/run_xPyD_models.slurm | 26 +++- scripts/vllm_dissag/vllm_disagg.sh | 7 +- 7 files changed, 198 insertions(+), 70 deletions(-) diff --git a/docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile b/docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile index f61a8819..8da85d77 100644 --- a/docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile +++ b/docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile @@ -53,18 +53,19 @@ # the moriep all-to-all combine at EP32 scale -> deferred to future work. Use 1P/1D # and 2P/2D only. (BASE_IMAGE is a gated nightly; override --build-arg BASE_IMAGE=...) # ============================================================================= -# Reconstructs the validated v1.2.1 (mori121) runtime stack by applying the recipe's -# component pins ON TOP of the open ROCm vLLM ci_base, cloning each source from -# public Git (no local build-contexts). Mirrors dist-inf-cookbook -# Dockerfile.vllm.mori121_shareable: +# Builds the GLM-5.1 runtime stack by applying component pins ON TOP of a +# purpose-built ROCm/vLLM/MoRI base, cloning each overridden source from public Git +# (no local build-contexts): # -# - BASE: rocm/vllm-dev:ci_base-0fcd9b99... (open ROCm 7.2 / cp312 CI base). -# - MoRI -> built from ROCm/MoRI @ v1.2.1 (BUILD_UMBP=OFF). +# - BASE: rocmshared/pytorch-private:vllm-rocm_07_22_2026_shikpate_mori1.2.3 +# (ROCm + torch + a bundled vLLM/MoRI 1.2.3 stack). The stages below deliberately +# OVERRIDE the base's vLLM/MoRI/AITER with the pins we validate for GLM DSA. +# - MoRI -> built from ROCm/MoRI @ 42e895472b08 (validated for GLM DSA, BUILD_UMBP=OFF). +# (main LATEST 120d2de broke the connector KV-notify handshake -- see note at MORI_REF.) # - AITER -> STOCK ROCm/aiter @ e03fa6040 compiled from source + flydsl 0.1.7-0.1.9; # stale JIT wiped. (#47766 keeps persistent MLA ON -> aiter native gqa64 fold.) -# - vLLM -> COMPILED from shikamd123/vllm @ -# vllm_2p2d_wide-ep_write_shikpate_test_06_29_customer (Wide-EP multi-pod PD, the -# connector/router reference for the 2P2D DP=EP=16 topology). Full compile: it is +# - vLLM -> COMPILED from raviguptaamd/vllm @ glm5.1-dsa-wideEP_on_shik_0721 +# (Shiksha 7/21 WideEP base + GLM DSA edits + sparse-MLA guard fix). Full compile: # a different commit than the base's, so a .py-only overlay would be ABI-mismatched. # - RDMA fix (expandable_segments:False x2 + HSA_ENABLE_IPC_MODE_LEGACY=0) is NOT baked # here — it lives in scripts/vllm_dissag/connectors/.env and the launcher @@ -78,12 +79,11 @@ # Build context = repo root: # docker build -f docker/vllm_disagg_inference.ubuntu.amd.Dockerfile -t / . # -# BASE_IMAGE is the open rocm/vllm-dev ci_base pinned by the validated recipe -# (dist-inf-cookbook Dockerfile.vllm.mori121_shareable). Override --build-arg +# BASE_IMAGE is the purpose-built ROCm/vLLM/MoRI base above. Override --build-arg # BASE_IMAGE=... to build on a different ROCm base. vLLM compile is long (~30-60 min). # ============================================================================= -ARG BASE_IMAGE=rocm/vllm-dev:ci_base-0fcd9b99cc9d63202da4c858d8ebc6582c9e2491 +ARG BASE_IMAGE=rocm/vllm-dev:ci_base-dedbf6be8b1afa17a6220473b9c8c98242ac1c03 FROM ${BASE_IMAGE} ENTRYPOINT [] @@ -109,8 +109,12 @@ ARG NIC_COMPILATION_ARCH="cx7" # backends produced a MoRI that deadlocked at the cross-node EP all-to-all init. # ----------------------------------------------------------------------------- ARG MORI_REPO=https://github.com/ROCm/mori.git -# 42e895472b08: MoRI main tip past v1.2.1, validated by MAD-private #338 for GLM-5.1 -# DSA WideEP disagg (v1.2.1 large-transfer notify path was insufficient at high EP). +# 42e895472b08: validated MoRI tip for GLM DSA WideEP disagg. The v0.27 base bundles +# amd_mori 1.0.0, but the bundled build regressed GLM DSA (b1: GPU fault on the aiter +# DSA decode kernel), so we build MoRI from source at this pinned commit by DEFAULT +# (WITH_MORI_BUILD=1). Set --build-arg WITH_MORI_BUILD=0 only to fall back to the +# base's bundled mori for debugging. +ARG WITH_MORI_BUILD=1 ARG MORI_REF=42e895472b08 ENV MORI_GPU_ARCHS=gfx942 # Newer MoRI added the UMBP subsystem which requires gRPC (grpcpp/grpcpp.h) not @@ -125,50 +129,49 @@ RUN sed -i 's|http://|https://|g' /etc/apt/sources.list 2>/dev/null || true && \ apt-get update && apt-get install -y --no-install-recommends \ git build-essential cmake ninja-build ccache libssl-dev pkg-config curl ca-certificates && \ pip install meson==0.64.0 "pybind11[global]" tqdm prettytable && \ - pip uninstall -y amd_mori amd-mori amd-mori-nightly mori 2>/dev/null || true && \ - rm -rf /tmp/mori-src && \ - git clone --recursive "${MORI_REPO}" /tmp/mori-src && \ - cd /tmp/mori-src && git checkout "${MORI_REF}" && git submodule update --init --recursive && \ - BUILD_UMBP=OFF pip install . && \ - python3 -c "import mori, mori.io, mori.ops; print('MoRI OK at', mori.__path__[0])" && \ - mkdir -p /app && echo "MORI_REF=${MORI_REF}@$(git -C /tmp/mori-src rev-parse HEAD)" >> /app/versions.txt && \ - rm -rf /tmp/mori-src + mkdir -p /app && \ + if [ "${WITH_MORI_BUILD}" != "1" ]; then \ + python3 -c "import mori, mori.io, mori.ops; print('MoRI (bundled) OK at', mori.__path__[0])" && \ + echo "MORI_REF=BUNDLED (base amd_mori, WITH_MORI_BUILD=0)" >> /app/versions.txt ; \ + else \ + pip uninstall -y amd_mori amd-mori amd-mori-nightly mori 2>/dev/null || true && \ + rm -rf /tmp/mori-src && \ + git clone --recursive "${MORI_REPO}" /tmp/mori-src && \ + cd /tmp/mori-src && git checkout "${MORI_REF}" && git submodule update --init --recursive && \ + BUILD_UMBP=OFF pip install . && \ + python3 -c "import mori, mori.io, mori.ops; print('MoRI OK at', mori.__path__[0])" && \ + echo "MORI_REF=${MORI_REF}@$(git -C /tmp/mori-src rev-parse HEAD)" >> /app/versions.txt && \ + rm -rf /tmp/mori-src ; \ + fi # ----------------------------------------------------------------------------- -# 2. AITER: build STOCK upstream ROCm/aiter @ e03fa6040 from source (NO fork, -# NO gqa64-fold patch). Under vLLM #47766 the sparse-MLA persistent path stays -# ON, so GLM's gqa=64 decode hits aiter's PRE-EXISTING persistent gqa64->16 fold -# (aiter/mla.py: `nhead in range(32,128+1,16) and persistent_mode`); the fork's -# extra non-persistent fold is never exercised, so stock is sufficient. -# Validated by MAD-private #338: 1P/1D EP8 + 2P/2D EP16 NIAH PASS on this exact -# aiter tip under #47766. Pin the exact commit (the one tested), not the release -# wheel. Then invalidate the stale prewarmed JIT cache compiled against the old .so. +# 2. AITER: the v0.27 base bundles amd-aiter 0.1.19 (+ flydsl 0.2.4), but bundled 0.1.19 +# GPU-faults on the GLM DSA decode kernel mla_a8w8_qh64_gqaratio64_v3 (confirmed b1 on +# this v0.27 base, same regression as the old stack). So we build aiter from source at +# the validated commit e03fa6040 by DEFAULT (WITH_AITER_BUILD=1). aiter > e03fa6040 +# reintroduces the fault; do not bump without re-running long-ctx NIAH. Set +# --build-arg WITH_AITER_BUILD=0 only to fall back to the bundled aiter for debugging. # ----------------------------------------------------------------------------- ARG AITER_REPO=https://github.com/ROCm/aiter.git +ARG WITH_AITER_BUILD=1 ARG AITER_REF=e03fa6040 -RUN echo "Compiling STOCK AITER (no fork) from ${AITER_REPO}@${AITER_REF}" && \ - rm -rf /tmp/aiter-src && \ - git clone --recursive "${AITER_REPO}" /tmp/aiter-src && \ - cd /tmp/aiter-src && git checkout "${AITER_REF}" && \ - git submodule update --init --recursive && \ - (pip uninstall -y amd_aiter amd-aiter aiter 2>/dev/null || true) && \ - pip install --no-build-isolation --no-deps -v . && \ - pip install --no-deps -U "flydsl>=0.1.7,<0.1.9" && \ - echo "AITER_REF=${AITER_REF}@$(git rev-parse HEAD) (stock ROCm/aiter, no fork)" >> /app/versions.txt && \ - rm -rf /tmp/aiter-src && \ - python3 - <<'PYEOF' -# Verify aiter/mla.py installed + has the persistent gqa64 fold, WITHOUT importing -# aiter/torch (torch->amdsmi->libamd_smi.so is not loadable at build: no GPU in sandbox). -import glob, pathlib -cands = glob.glob("/usr/local/lib/python*/dist-packages/aiter/mla.py") + \ - glob.glob("/usr/lib/python*/dist-packages/aiter/mla.py") -assert cands, "aiter/mla.py not found in site-packages after install" -src = pathlib.Path(cands[0]).read_text() -assert "persistent_mode" in src, f"AITER persistent fold path MISSING in {cands[0]}" -print("STOCK AITER OK (persistent gqa64 fold path present):", cands[0]) -PYEOF -RUN rm -rf /opt/vllm_cache/aiter_jit /root/.aiter && echo "cleared stale AITER JIT cache" && \ - echo "AITER_REF=${AITER_REF} (stock)" >> /app/versions.txt +RUN if [ "${WITH_AITER_BUILD}" != "1" ]; then \ + echo "AITER: using BUNDLED base aiter (WITH_AITER_BUILD=0)" && \ + python3 -c "import importlib.metadata as m; print('aiter (bundled)', m.version('amd-aiter'))" && \ + echo "AITER_REF=BUNDLED (base amd-aiter, WITH_AITER_BUILD=0)" >> /app/versions.txt ; \ + else \ + echo "Compiling STOCK AITER (no fork) from ${AITER_REPO}@${AITER_REF}" && \ + rm -rf /tmp/aiter-src && \ + git clone --recursive "${AITER_REPO}" /tmp/aiter-src && \ + cd /tmp/aiter-src && git checkout "${AITER_REF}" && \ + git submodule update --init --recursive && \ + (pip uninstall -y amd_aiter amd-aiter aiter 2>/dev/null || true) && \ + pip install --no-build-isolation --no-deps -v . && \ + pip install --no-deps -U "flydsl>=0.1.7,<0.1.9" && \ + echo "AITER_REF=${AITER_REF}@$(git rev-parse HEAD) (stock ROCm/aiter, no fork)" >> /app/versions.txt && \ + rm -rf /tmp/aiter-src && \ + rm -rf /opt/vllm_cache/aiter_jit /root/.aiter && echo "cleared stale AITER JIT cache" ; \ + fi # ----------------------------------------------------------------------------- # 3. vLLM: compile from source at the 06_29 validated Wide-EP WRITE-mode branch @@ -181,7 +184,14 @@ RUN rm -rf /opt/vllm_cache/aiter_jit /root/.aiter && echo "cleared stale AITER J # VLLM_REPO/REF are a PUBLIC GitHub repo + branch (the Wide-EP WRITE-mode vLLM the # dist-inf-cookbook mori121 image builds from). Override to your own vLLM fork/branch. ARG VLLM_REPO=https://github.com/raviguptaamd/vllm.git -ARG VLLM_REF=glm5.1-dsa-wideEP_on_shik_latest +# glm5.1-dsa-wideEP_on_vllm-v0.27 (HEAD cda3648602) = upstream v0.27 tip dedbf6be8b + 7 +# ROCm/DSA commits. Core 3: per-req-ctx metadata key (#47766), DSA indexer KV transfer +# (reworked onto upstream's native MoRIIO connector), invalid-token sentinel. Plus 4 +# v0.27 fixes: concat_and_cache_mla positional (stable-ABI), splitting_ops out of the +# compiled graph (MLA "unknown parameter type"), sparse-indexer bounds-guard, and the +# decisive sentinel -1->0 (cda3648602 — aiter mla_decode_fwd derefs -1 -> GPU fault at +# disagg long-ctx). NIAH-validated 1P/1D + 2P/1D + 1P/2D, 2k-35k, decode PIECEWISE. +ARG VLLM_REF=glm5.1-dsa-wideEP_on_vllm-v0.27 ENV VLLM_TARGET_DEVICE=rocm \ PYTORCH_ROCM_ARCH=${PYTORCH_ROCM_ARCH} \ MAX_JOBS=${MAX_JOBS} @@ -203,12 +213,14 @@ def get(names): except PackageNotFoundError: pass return None av = get(("amd-aiter", "amd_aiter", "aiter")) -# Stock source build of ROCm/aiter@e03fa6040 reports 0.1.17.dev195+ge03fa6040. -# Verify the aiter install survived the vLLM install (present + carries the e03fa6040 -# commit tag) rather than pinning a release version string. -assert av and "e03fa6040" in av, f"AITER missing/downgraded (want e03fa6040 build): {av!r}" +# Verify the aiter install survived the vLLM install (present, not silently downgraded +# to a base-bundled wheel). We pin aiter by commit (e03fa6040), whose reported version +# string varies by build, so assert presence rather than a hardcoded commit substring. Do NOT +# `import aiter` here: it pulls torch->amdsmi->libamd_smi.so, not loadable in the no-GPU +# build sandbox (same reason the Stage-2 verify reads mla.py from disk instead). +assert av, "AITER missing after vLLM install (expected bundled 0.1.19 or source-built ref)" import mori, mori.io, mori.ops -print("Post-vLLM check OK: AITER", av, "+ MoRI importable") +print("Post-vLLM check OK: AITER", av, "present + MoRI importable") PYEOF # ----------------------------------------------------------------------------- diff --git a/scripts/vllm_dissag/benchmark_xPyD.sh b/scripts/vllm_dissag/benchmark_xPyD.sh index b8851d24..699c9a4e 100755 --- a/scripts/vllm_dissag/benchmark_xPyD.sh +++ b/scripts/vllm_dissag/benchmark_xPyD.sh @@ -40,6 +40,32 @@ for i in $(seq 1 $BENCHMARK_ITR); do echo "Running the benchserving script for iter: $i" | tee -a ${LOG}_CONCURRENCY.log >/dev/null for combo in "${COMBINATIONS[@]}"; do IFS="/" read -r isl osl <<< "$combo" + # Per-shape warmup at the REAL isl/osl, low concurrency. The global warmup above + # is isl=osl=32/con=1, which never exercises this shape's prefill path, its Triton/ + # aiter kernel variants, or the decode cudagraph batch sizes -- so without this the + # FIRST measured cell of each shape absorbs all the residual JIT and reports a + # wildly inflated TPOT (observed 302ms vs ~89ms steady-state). Measured cells must + # start from a warm graph. Skip with SHAPE_WARMUP=0. + if [[ "${SHAPE_WARMUP:-1}" == "1" ]]; then + _w_con="${SHAPE_WARMUP_CON:-4}" + _w_prompts="${SHAPE_WARMUP_PROMPTS:-8}" + echo "[WARMUP] shape isl $isl osl $osl con ${_w_con} prompts ${_w_prompts}" \ + | tee -a ${LOG}_CONCURRENCY.log >/dev/null + timeout "${SHAPE_WARMUP_TIMEOUT:-2400}" vllm bench serve \ + --model $MODEL_PATH \ + --backend vllm \ + --host 127.0.0.1 \ + --port $BENCHMARK_PORT \ + --dataset-name "random" \ + --random-input-len $isl \ + --random-output-len $osl \ + --random-prefix-len 0 \ + --num-prompts ${_w_prompts} \ + --request-rate "inf" \ + --ignore-eos \ + --max-concurrency ${_w_con} \ + 2>&1 | tee -a ${LOG}_SHAPEWARMUP.log >/dev/null + fi for con in $CON; do p_con=$(($con * 2)) if [ "$p_con" -lt 16 ]; then diff --git a/scripts/vllm_dissag/connectors/moriio.env b/scripts/vllm_dissag/connectors/moriio.env index 29ed23ba..780f0214 100644 --- a/scripts/vllm_dissag/connectors/moriio.env +++ b/scripts/vllm_dissag/connectors/moriio.env @@ -23,7 +23,18 @@ MORI_RDMA_TC=41 MORI_RDMA_SL=0 MORI_IO_SL=1 MORI_IB_ENABLE_RELAXED_ORDERING=1 -MORI_IB_GID_INDEX=1 +MORI_IB_GID_INDEX=3 +# RDMA NIC allowlist (MI300 + CX7 / RoCE). Without this MoRI auto-enumerates ALL ibv +# devices incl. the mgmt NICs (mlx5_1=eth0, mlx5_6=eth1 on the 10.158 mgmt net), and +# tries to establish QPs over a non-routable/mgmt fabric -> ibverbs.cpp:189 "Connection +# timed out" at the prefill->decode KV transfer. Restrict to the 8 GPU-RoCE NICs +# (rdma0-7 on the 10.224 fabric) per the dist-inf-cookbook cluster-rdma-env-recommender. +# NCCL/GLOO use eth0 for their (non-RDMA) control sockets. Override per-fabric if needed. +MORI_RDMA_DEVICES=mlx5_0,mlx5_2,mlx5_3,mlx5_4,mlx5_5,mlx5_7,mlx5_8,mlx5_9 +NCCL_IB_HCA=mlx5_0,mlx5_2,mlx5_3,mlx5_4,mlx5_5,mlx5_7,mlx5_8,mlx5_9 +NCCL_IB_GID_INDEX=3 +NCCL_SOCKET_IFNAME=eth0 +GLOO_SOCKET_IFNAME=eth0 MORI_NUM_QP_PER_PE=8 VLLM_MORIIO_QP_PER_TRANSFER=2 VLLM_MORIIO_NUM_WORKERS=4 diff --git a/scripts/vllm_dissag/connectors/moriio.sh b/scripts/vllm_dissag/connectors/moriio.sh index c26a0944..33d5322d 100644 --- a/scripts/vllm_dissag/connectors/moriio.sh +++ b/scripts/vllm_dissag/connectors/moriio.sh @@ -245,12 +245,21 @@ connector_launch_worker() { else _cudagraph_mode="${PREFILL_CUDAGRAPH_MODE:-$_cudagraph_mode}" fi + # v0.27 MLA fix: the *_kv_cache_update op dispatches the STABLE-ABI concat_and_cache_mla + # whose boxed kernel does NOT compose inside the Dynamo-FX-partitioned compiled graph + # -> "RuntimeError: unknown parameter type" on the first real MLA decode (passes boot + + # warmup via the fake path, then crashes). splitting_ops-list membership alone doesn't + # cut the graph there. use_inductor_graph_partition=true moves partitioning to inductor + # codegen (after all passes), splitting at cudagraph_unsafe ops incl. the KV-update so + # it runs as an eager boundary. Toggle via USE_INDUCTOR_GRAPH_PARTITION (default 1). + local _igp_json="" + [[ "${USE_INDUCTOR_GRAPH_PARTITION:-1}" == "1" ]] && _igp_json=',"use_inductor_graph_partition":true' if [[ -n "$_cudagraph_mode" && "$_cudagraph_mode" != "NONE" ]]; then local _capture_sizes="${CUDAGRAPH_CAPTURE_SIZES:-1 2 4 8 16 32 64 128 256}" - exec_args+=(--compilation-config '{"cudagraph_mode":"'"${_cudagraph_mode}"'","custom_ops":["+quant_fp8"]}') + exec_args+=(--compilation-config '{"cudagraph_mode":"'"${_cudagraph_mode}"'","custom_ops":["+quant_fp8"]'"${_igp_json}"'}') exec_args+=(--cudagraph-capture-sizes ${_capture_sizes}) else - exec_args+=(--compilation-config '{"cudagraph_mode":"NONE","custom_ops":["+quant_fp8"]}') + exec_args+=(--compilation-config '{"cudagraph_mode":"NONE","custom_ops":["+quant_fp8"]'"${_igp_json}"'}') fi # Per-model flags from models.yaml (driver-exported; empty if none). @@ -265,6 +274,25 @@ connector_launch_worker() { local _all2all="${PREFILL_MORI_BACKEND}" [[ "$log_prefix" == "decode" ]] && _all2all="${DECODE_MORI_BACKEND}" + # Per-role MoRI EP buffer width. VLLM_MORI_MAX_TOKENS_PER_RANK sizes the + # dispatch/combine buffer; unset (0) it inherits max_num_batched_tokens -- a + # chunked-prefill SCHEDULER knob (8192) -- so a decode instance moves an + # 8192-token-wide buffer every step, per layer: ~302ms vs ~88ms TPOT (3.4x). + # Prefill and decode want OPPOSITE values (prefill genuinely dispatches wide + # batches and must keep the large buffer), but models.yaml env: applies to BOTH + # roles -- so split it here, mirroring PREFILL/DECODE_MORI_BACKEND above. + if [[ "$log_prefix" == "decode" ]]; then + [[ -n "${DECODE_MORI_MAX_TOKENS_PER_RANK:-}" ]] && \ + export VLLM_MORI_MAX_TOKENS_PER_RANK="${DECODE_MORI_MAX_TOKENS_PER_RANK}" + # Recv capacity must still cover vLLM's profiling dummy run (which pushes + # max_num_batched_tokens tokens) even though steady-state dispatch is narrow. + [[ -n "${DECODE_MORI_MAX_TOTAL_RECV_TOKENS:-}" ]] && \ + export VLLM_MORI_MAX_TOTAL_RECV_TOKENS="${DECODE_MORI_MAX_TOTAL_RECV_TOKENS}" + else + export VLLM_MORI_MAX_TOKENS_PER_RANK="${PREFILL_MORI_MAX_TOKENS_PER_RANK:-0}" + export VLLM_MORI_MAX_TOTAL_RECV_TOKENS="${PREFILL_MORI_MAX_TOTAL_RECV_TOKENS:-0}" + fi + local extra_args=() kv_args=() if [[ "$role" == "master" ]]; then extra_args+=(--api-server-count=${_GPUS_PER_NODE}) diff --git a/scripts/vllm_dissag/models.yaml b/scripts/vllm_dissag/models.yaml index 9eec3e7d..e0017704 100644 --- a/scripts/vllm_dissag/models.yaml +++ b/scripts/vllm_dissag/models.yaml @@ -235,6 +235,14 @@ DeepSeek-R1: GLM-5.1-FP8: env: VLLM_USE_V1: "1" + # v0.27: layer_name is wrapped in a torch OpaqueBase (LayerName) and passed through + # the unified_mla_kv_cache_update / unified_mla_attention custom ops. On this ROCm + # torch 2.12 build the opaque-type boxing FAILS -> "RuntimeError: unknown parameter + # type" at torch/_ops.py on the first real MLA decode forward (the fake/compile path + # returns early, so it passes boot+warmup then crashes on first request -> DP gloo + # cascade -> 0 prefill/0 decode). VLLM_USE_LAYERNAME=0 makes LayerNameType=str (the + # pre-2.11 path), so a plain string passes through the op. No image rebuild needed. + VLLM_USE_LAYERNAME: "0" VLLM_ROCM_USE_AITER: "1" VLLM_ROCM_USE_AITER_RMSNORM: "1" VLLM_ROCM_USE_AITER_MLA: "1" @@ -248,6 +256,21 @@ GLM-5.1-FP8: VLLM_ALL2ALL_BACKEND: "mori_high_throughput" PREFILL_MORI_BACKEND: "mori_high_throughput" DECODE_MORI_BACKEND: "mori_low_latency" + # MoRI EP dispatch/combine buffer width. Without this it inherits + # max_num_batched_tokens (8192) -- a chunked-prefill SCHEDULER setting -- so every + # decode step moves an 8192-token-wide buffer per layer x78 layers regardless of the + # real batch. That is a fixed ~300ms/step floor (~320x this model's HBM-bandwidth + # bound). Sizing it for the actual decode batch gives TPOT 302ms -> 88ms (3.4x), + # matching the published EP8 figure. Needs vLLM >= e8c186f71b. + # Prefill is unaffected (it genuinely dispatches wide batches). + # Per-role (moriio.sh -> VLLM_MORI_MAX_TOKENS_PER_RANK). MUST be >= max_num_seqs * + # num_experts_per_token (top-8 here): the value is also the RECV capacity + # (mori MaxNumTokensToRecv = worldSize * this). 512 tripped a device assert + # "Total recv token overflow". 2048 is safe and still 4x smaller than 8192. Decode wants + # the buffer sized to its real batch; prefill must KEEP the wide buffer (0 = inherit + # max_num_batched_tokens) because it genuinely dispatches 8192-token chunks. + # Recv capacity stays wide enough for the profiling dummy run (8192 tokens) and + # bursty top-8 routing: worldSize * ceil(65536/worldSize) >= 8192 for EP8..EP32. MORI_SHMEM_HEAP_SIZE: "17179869184" # DSA sparse-indexer logits-buffer cap (crash fix). The indexer prefill computes an # M*N fp32 logits buffer; split_indexer_prefill_chunks only sub-chunks the query dim @@ -275,4 +298,13 @@ GLM-5.1-FP8: prefill: dp: "" decode: - dp: "" + # PERF: the MoRI EP dispatch buffer width is max_num_batched_tokens (via + # FusedMoEConfig.max_num_tokens -> all2all.py max_num_inp_token_per_rank), so the + # decode role otherwise runs an 8192-token-wide all2all every step: ~302ms TPOT. + # mori bounds recv capacity BY the send width (MaxNumTokensToRecvPerRank returns + # min(ceil(maxTotalRecvTokens/ws), maxNumInpTokenPerRank)), so the buffer must still + # cover vLLM's profiling dummy run -- it cannot be shrunk via env alone. Lowering + # this knob on the DECODE role lowers both consistently. + # Value must stay >= typical prompt length: 512 gave 87.9ms TPOT but 13.1s TTFT + # (a 1024-token prompt could not be admitted in one step). 2048 keeps TTFT healthy. + dp: "--max-num-batched-tokens 2048" diff --git a/scripts/vllm_dissag/run_xPyD_models.slurm b/scripts/vllm_dissag/run_xPyD_models.slurm index 98281f57..52d046a1 100755 --- a/scripts/vllm_dissag/run_xPyD_models.slurm +++ b/scripts/vllm_dissag/run_xPyD_models.slurm @@ -76,7 +76,7 @@ echo "Running from: $(pwd)" # model's own recipe — GLM-5.1 DSA needs block=1 + AITER MLA on), while a genuine # submit-time `-e VAR=...` still wins. Precedence: image-baked < models.yaml < submit -e. # Captured HERE (before the slurm sets any defaults) so it reflects user intent only. -_RECIPE_ENV_KEYS="VLLM_USE_V1 VLLM_ROCM_USE_AITER VLLM_ROCM_USE_AITER_RMSNORM VLLM_ROCM_USE_AITER_MLA KV_BLOCK_SIZE KV_CACHE_DTYPE KV_CACHE_MEMORY_BYTES GPU_MEMORY_UTILIZATION VLLM_CUDAGRAPH_MODE PREFILL_CUDAGRAPH_MODE DECODE_CUDAGRAPH_MODE CUDAGRAPH_CAPTURE_SIZES VLLM_ALL2ALL_BACKEND PREFILL_MORI_BACKEND DECODE_MORI_BACKEND MORI_SHMEM_HEAP_SIZE" +_RECIPE_ENV_KEYS="VLLM_USE_V1 DECODE_MORI_MAX_TOTAL_RECV_TOKENS PREFILL_MORI_MAX_TOTAL_RECV_TOKENS DECODE_MORI_MAX_TOKENS_PER_RANK PREFILL_MORI_MAX_TOKENS_PER_RANK VLLM_MORI_MAX_TOKENS_PER_RANK VLLM_MORI_WARP_NUM_PER_BLOCK VLLM_MORI_BLOCK_NUM VLLM_MORI_RDMA_BLOCK_NUM VLLM_USE_LAYERNAME VLLM_ROCM_USE_AITER VLLM_ROCM_USE_AITER_RMSNORM VLLM_ROCM_USE_AITER_MLA KV_BLOCK_SIZE KV_CACHE_DTYPE KV_CACHE_MEMORY_BYTES GPU_MEMORY_UTILIZATION VLLM_CUDAGRAPH_MODE PREFILL_CUDAGRAPH_MODE DECODE_CUDAGRAPH_MODE CUDAGRAPH_CAPTURE_SIZES VLLM_ALL2ALL_BACKEND PREFILL_MORI_BACKEND DECODE_MORI_BACKEND MORI_SHMEM_HEAP_SIZE" MODELS_YAML_PROTECT="" for _k in $_RECIPE_ENV_KEYS; do [ -n "${!_k+x}" ] && MODELS_YAML_PROTECT="${MODELS_YAML_PROTECT} ${_k}" @@ -426,15 +426,22 @@ echo "" # Node information USER_NAME=$(whoami) MASTER_NODE=$(echo "$SELECTED_NODES" | head -n 1) -MASTER_ADDR=$(srun --nodes=1 --ntasks=1 --time=00:20:00 --nodelist="$MASTER_NODE" bash -c 'hostname -I') -MASTER_ADDR=$(echo "$MASTER_ADDR" | awk 'NR==1 {print $1}') +# Pick the routable fabric IP, not just hostname -I's first entry. These nodes expose +# multiple NICs (e.g. a 10.224.x overlay listed BEFORE the routable 10.158.x fabric); +# taking $1 blindly can advertise an unreachable addr -> prefill/decode barrier hangs +# "Waiting for nodes" forever. Prefer FABRIC_SUBNET (default 10.158.), fall back to $1. +FABRIC_SUBNET="${FABRIC_SUBNET:-10.158.}" +# From a "hostname -I" line, return the first IP on FABRIC_SUBNET, else the first IP. +_pick_fabric_ip() { + awk -v pfx="$FABRIC_SUBNET" '{f=$1; for(i=1;i<=NF;i++) if(index($i,pfx)==1){f=$i; break} print f}' +} +MASTER_ADDR=$(srun --nodes=1 --ntasks=1 --time=00:20:00 --nodelist="$MASTER_NODE" bash -c 'hostname -I' | _pick_fabric_ip) MASTER_PORT=39566 # Choose an open port IPS=() for NODE in $SELECTED_NODES; do - IP=$(srun --nodes=1 --ntasks=1 --time=00:20:00 --nodelist="$NODE" bash -c 'hostname -I') - IP=$(echo "$IP" | awk 'NR==1 {print $1}') + IP=$(srun --nodes=1 --ntasks=1 --time=00:20:00 --nodelist="$NODE" bash -c 'hostname -I' | _pick_fabric_ip) IPS+=("$IP") done @@ -549,7 +556,12 @@ done for _pattern in libmlx5.so* libionic*.so* libbnxt_re*.so* libefa.so* libhns.so*; do for _vlib in $_LIBDIR/${_pattern}; do - [ -e "$_vlib" ] && _RDMA_MOUNTS="$_RDMA_MOUNTS -v $_vlib:$_vlib:ro" + # Require a regular file AFTER symlink resolution: these mounts are built on + # ONE node but applied on ALL nodes, and vendor NIC libs (e.g. libionic.so.1) + # can be a DANGLING symlink on some nodes -> bind-mount fails "not a directory" + # -> container create exit 125. `-f` (follows symlink, requires regular file) + # skips those; the fabric in use (mlx5) is still mounted where present. + [ -f "$_vlib" ] && _RDMA_MOUNTS="$_RDMA_MOUNTS -v $_vlib:$_vlib:ro" done done @@ -578,6 +590,8 @@ docker run --rm \ -v $NIXL_REPO_DIR:$NIXL_COOKBOOK_PATH \ -v /tmp/vllm_cache:/tmp/vllm_cache \ ${_JIT_CACHE_MOUNT} \ + ${GLM_KERNEL_PATCH:+-v ${GLM_KERNEL_PATCH}:/usr/local/lib/python3.12/dist-packages/vllm/v1/attention/ops/rocm_aiter_mla_sparse.py:ro} \ + ${GLM_BACKEND_PATCH:+-v ${GLM_BACKEND_PATCH}:/usr/local/lib/python3.12/dist-packages/vllm/v1/attention/backends/mla/rocm_aiter_mla_sparse.py:ro} \ $_RDMA_MOUNTS \ --entrypoint /bin/bash \ -e SLURM_JOB_ID=$SLURM_JOB_ID \ diff --git a/scripts/vllm_dissag/vllm_disagg.sh b/scripts/vllm_dissag/vllm_disagg.sh index acbabbd4..da2afb00 100755 --- a/scripts/vllm_dissag/vllm_disagg.sh +++ b/scripts/vllm_dissag/vllm_disagg.sh @@ -86,7 +86,12 @@ IFS=',' read -ra IP_ARRAY <<< "${IPADDRS}" echo "Listing NIXL_COOKBOOK_PATH: ${NIXL_COOKBOOK_PATH:-}" [[ -n "${NIXL_COOKBOOK_PATH:-}" ]] && ls "${NIXL_COOKBOOK_PATH}" -host_ip=$(hostname -I | awk '{print $1}') +# Prefer the routable fabric IP (FABRIC_SUBNET, default 10.158.) over hostname -I's +# first entry: nodes with a 10.224.x overlay listed first would bind the socket_barrier +# / advertise host_ip on an unreachable NIC -> prefill<->decode barrier hangs "Waiting +# for nodes". Matches the IPADDRS selection in run_xPyD_models.slurm. Falls back to $1. +FABRIC_SUBNET="${FABRIC_SUBNET:-10.158.}" +host_ip=$(hostname -I | awk -v pfx="$FABRIC_SUBNET" '{f=$1; for(i=1;i<=NF;i++) if(index($i,pfx)==1){f=$i; break} print f}') host_name=$(hostname) # ============================================================================= From 0724197ae5a00a44cb306ae00f66ad7e4f89af51 Mon Sep 17 00:00:00 2001 From: raviguptaamd Date: Sun, 16 Aug 2026 04:24:53 +0000 Subject: [PATCH 06/41] [GLM-5.1] Default WITH_NIXL=0 (lean MoRI-EP-only image) GLM-5.1 is served over MoRI-EP (all2all) + MoRI-IO (KV transfer); the rixl connector's UCX/RIXL/rocSHMEM/DeepEP stack is never selected by this recipe, so building it only lengthens the build and ships unused transports. 0 is also how the validated image was actually built: glm5.1-vllm027-b8 used --build-arg WITH_NIXL=0, so the previous default of 1 meant a plain `docker build` did NOT reproduce the tested artifact. With this change, and with WITH_MORI_BUILD/WITH_AITER_BUILD already defaulting to 1, a no-flag build now matches the validated stack exactly. Set --build-arg WITH_NIXL=1 if you need the rixl connector from this same Dockerfile. Co-Authored-By: Claude --- .../vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile | 9 ++++++--- 1 file changed, 6 insertions(+), 3 deletions(-) diff --git a/docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile b/docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile index 8da85d77..edd4d797 100644 --- a/docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile +++ b/docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile @@ -92,9 +92,12 @@ WORKDIR /app ARG GFX_COMPILATION_ARCH="gfx942" ARG PYTORCH_ROCM_ARCH="gfx942" ARG MAX_JOBS=32 -# NIXL/RIXL transport for the rixl connector. Default 1 => all connectors built -# (UCX/RIXL/rocSHMEM/DeepEP). Set --build-arg WITH_NIXL=0 for a lean MoRI-EP-only image. -ARG WITH_NIXL=1 +# NIXL/RIXL transport for the rixl connector. GLM-5.1 is served over MoRI-EP + MoRI-IO, +# so the UCX/RIXL/rocSHMEM/DeepEP stack is dead weight here: it lengthens the build and +# ships transports this recipe never selects. Default 0 => lean MoRI-EP-only image, which +# is also exactly how the validated image (glm5.1-vllm027-b8) was built. Set +# --build-arg WITH_NIXL=1 only if you need the rixl connector from this same Dockerfile. +ARG WITH_NIXL=0 ARG NIC_COMPILATION_ARCH="cx7" # ----------------------------------------------------------------------------- From 09d84ce535682d3ba50c396a581e66c7cf9182c2 Mon Sep 17 00:00:00 2001 From: raviguptaamd Date: Sun, 16 Aug 2026 04:28:57 +0000 Subject: [PATCH 07/41] [GLM-5.1] Add long-context NIAH harness + vllm-disagg operational playbook niah_200k.py: needle-in-a-haystack sweep that validated GLM-5.1-FP8 to 200,049 tokens on both 1P/1D (EP8) and 2P/2D (EP16). Reports found/10, latency, and the server-reported prompt_tokens per length, and writes JSON. Model id is overridable via NIAH_MODEL so it is not GLM-specific. The existing benchmark_niah.* stop well short of this range; this covers the 64k-200k band. skills_vllm_disagg.md: operational playbook for vLLM PD-disaggregated WideEP on MI300X (MoRI-EP + MoRI-IO), written from this enablement. Documents, with measurements: - benchmarking method: ALWAYS discard the first post-boot run (cold Triton JIT made TTFT read 13.4s vs 906ms warm; with prefill eager the JIT cost lands in TTFT, not TPOT), and sanity-check against the HBM-bandwidth bound before blaming a kernel - the max_num_batched_tokens trap: a chunked-prefill SCHEDULER knob also sizes the MoRI EP dispatch buffer, so decode ran an 8192-token-wide all2all every step (302ms -> 88ms TPOT once sized for the real batch) - the DSA sentinel landmine: the invalid sparse-index sentinel must be 0, not -1, because aiter's mla_decode_fwd dereferences it (only bites at disagg long context) - three documented DEAD ENDS so they are not retried, including why mori's max_total_recv_tokens cannot decouple recv from send capacity (the clamp is a min()) - cache/boot behaviour (three caches with different rules, the aiter baton lock, measured boot times), readiness signals for multi-node topologies, per-role env plumbing, and RDMA fabric verification (a node can be SLURM-'alloc' with a dead fabric - verify with ping/ib_write_bw before blaming code) Co-Authored-By: Claude --- scripts/vllm_dissag/niah_200k.py | 86 ++++++++ scripts/vllm_dissag/skills_vllm_disagg.md | 245 ++++++++++++++++++++++ 2 files changed, 331 insertions(+) create mode 100755 scripts/vllm_dissag/niah_200k.py create mode 100644 scripts/vllm_dissag/skills_vllm_disagg.md diff --git a/scripts/vllm_dissag/niah_200k.py b/scripts/vllm_dissag/niah_200k.py new file mode 100755 index 00000000..65273bc1 --- /dev/null +++ b/scripts/vllm_dissag/niah_200k.py @@ -0,0 +1,86 @@ +#!/usr/bin/env python3 +"""Needle-in-a-haystack sweep including long context (validated to 200K tokens). + +Hides 10 animal names at even intervals in a filler-word haystack and asks the model to +list them back. Reports found/10, end-to-end latency, and the server-reported +prompt_tokens per length, and can dump the whole run to JSON. + +Usage: + niah_200k.py [lengths_csv] [out_json] + base_url e.g. http://127.0.0.1:20005 (prefill/serve port, or the router) + lengths_csv comma-separated WORD counts (default: 2k..200k) + out_json optional path to write results + + NIAH_MODEL= override the served model id (default below) + +Lengths are given in WORDS to stay comparable with earlier published runs. On this filler +the GLM tokenizer lands ~1 token/word, so words ~= tokens (the script prints the actual +prompt_tokens so you can check). + +NOTE: the FIRST request after a server boot pays cold Triton JIT and can take >80s with a +prefill instance running eager. Warm the server (or use a generous timeout) before +treating any latency number here as steady-state. +""" +import json, os, sys, time, random, urllib.request + +BASE = sys.argv[1] if len(sys.argv) > 1 else "http://127.0.0.1:20005" +LENGTHS = [int(x) for x in (sys.argv[2].split(",") if len(sys.argv) > 2 else + "2000,8000,16000,20000,28000,35000,64000,100000,150000,200000".split(","))] +OUT = sys.argv[3] if len(sys.argv) > 3 else None +URL = BASE.rstrip("/") + "/v1/chat/completions" +MODEL = os.environ.get("NIAH_MODEL", "/mnt/m2m_nobackup/models_blog/GLM-5.1-FP8") + +FILLER = ("table chair window bottle pencil garden river mountain coffee planet " + "engine guitar pillow ticket basket candle market silver button orange").split() +ANIMALS = ["elephant", "giraffe", "kangaroo", "penguin", "dolphin", + "tiger", "rhinoceros", "octopus", "crocodile", "panda"] +SYS = ("You read a word list and pick out the animals. Reply with a single " + "comma-separated list of lowercase animal names. Output nothing else.") + + +def hay(n, seed=0): + rng = random.Random(seed) + w = [rng.choice(FILLER) for _ in range(n)] + step = max(n // (len(ANIMALS) + 1), 1) + for i, a in enumerate(ANIMALS): + w[min((i + 1) * step, len(w) - 1)] = a + return " ".join(w) + + +def run(n, seed=0, timeout=1800): + body = { + "model": MODEL, + "messages": [ + {"role": "system", "content": SYS}, + {"role": "user", "content": "Find the animals in this list:\n\n" + hay(n, seed)}, + ], + "temperature": 0, + "max_tokens": 128, + "chat_template_kwargs": {"enable_thinking": False}, + } + req = urllib.request.Request(URL, data=json.dumps(body).encode(), + headers={"Content-Type": "application/json"}) + t = time.time() + try: + r = json.loads(urllib.request.urlopen(req, timeout=timeout).read()) + m = r["choices"][0]["message"] + txt = ((m.get("content") or "") + " " + (m.get("reasoning") or "")).lower() + found = sorted(a for a in ANIMALS if a in txt) + u = r.get("usage") or {} + rec = {"words": n, "seed": seed, "found": len(found), "latency_s": round(time.time() - t, 1), + "prompt_tokens": u.get("prompt_tokens"), "animals": found} + print("words=%7d tok=%-7s found=%2d/10 (%6.1fs) %s" % ( + n, rec["prompt_tokens"], rec["found"], rec["latency_s"], found), flush=True) + return rec + except Exception as e: + rec = {"words": n, "seed": seed, "found": -1, "latency_s": round(time.time() - t, 1), + "error": str(e)[:200]} + print("words=%7d ERROR (%.1fs) %s" % (n, rec["latency_s"], rec["error"]), flush=True) + return rec + + +results = [run(n) for n in LENGTHS] +if OUT: + with open(OUT, "w") as f: + json.dump(results, f, indent=2) + print("wrote", OUT, flush=True) diff --git a/scripts/vllm_dissag/skills_vllm_disagg.md b/scripts/vllm_dissag/skills_vllm_disagg.md new file mode 100644 index 00000000..ee778064 --- /dev/null +++ b/scripts/vllm_dissag/skills_vllm_disagg.md @@ -0,0 +1,245 @@ +# skills_vllm_disagg.md + +Hard-won operational knowledge for **vLLM PD-disaggregated WideEP serving on AMD MI300X** +(MoRI-EP all-to-all + MoRI-IO RDMA KV transfer), learned while bringing GLM-5.1-FP8 +(MLA + DeepSeek Sparse Attention) onto vLLM v0.27. + +Everything below is *measured*, not theorised. Where a belief turned out to be wrong, +the wrong belief is kept alongside the correction — those are the expensive lessons. + +--- + +## 1. Benchmarking methodology (read this first — it invalidated three of my conclusions) + +### 1.1 ALWAYS discard the first bench run after a boot +The first real request after startup pays **cold Triton JIT** for the sparse/indexer +kernels (`_indexer_k_quant_and_cache_kernel`, `generate_sparse_seqlen_kernel`, +`_convert_req_index_to_global_index_kernel`). Measured on 1P/1D EP8, identical bench: + +| run | TTFT | TPOT | +|---|---|---| +| 1st after boot (cold) | 13,451 ms | 88.7 ms | +| 2nd (warm) | **906 ms** | 88.0 ms | + +TTFT moved **14.9x**; TPOT barely moved. A single cold run made me invent (and act on) +a false "scheduler admission" theory twice. Discard it, or warm up explicitly. + +### 1.2 Why cold JIT lands on TTFT and not TPOT +With the standard recipe **prefill = eager (`CUDAGraphMode.NONE`)**, decode = PIECEWISE: +- prefill has no graph capture at boot -> its kernels compile lazily on the **first real + request** -> the whole compile cost is inside TTFT. +- decode captured graphs at boot ("Graph capturing finished") -> already warm -> TPOT is + correct even on the cold run. +This asymmetry is diagnostic: *cold-JIT symptoms show up in TTFT only*. + +### 1.3 The harness warmup is not a warmup +`benchmark_xPyD.sh` warms at `isl=32 osl=32 con=1` — which never exercises a 1024/8192/28672 +prefill path, nor the decode cudagraph batch sizes. The first *measured* cell therefore +absorbs residual JIT. Fix applied: a per-shape warmup at the real ISL/OSL before each +shape's cells (writes to a separate `_SHAPEWARMUP.log` so it can't pollute the CSV). + +### 1.4 Client timeouts read as server failures +A 50-80s curl timeout against a cold server returns an **empty body**, which looks exactly +like a crash. Use >=300s on the first request. This produced a false "total failure" +verdict during EP32 debugging. + +### 1.5 Sanity-check against physics before blaming a kernel +GLM-5.1-FP8 activates ~37.7B params/token. At 5.3 TB/s HBM, fp8: +- full model on one rank: **7.1 ms/step** +- 1/8 of experts per rank (EP8): **0.9 ms/step** + +Measured 290-300 ms => ~320x the bound. That immediately rules out "compute" or "bandwidth" +and says *stall / oversized transfer*. Do this arithmetic early; it saves hours. + +--- + +## 2. The big perf trap: `max_num_batched_tokens` sizes the MoRI EP buffer + +### 2.1 The chain +``` +vllm/model_executor/layers/fused_moe/layer.py:349 + max_num_tokens = max_num_batched_tokens # 8192 default (SCHEDULER knob) +vllm/model_executor/layers/fused_moe/all2all_utils.py:181 + max_num_tokens_per_dp_rank = moe.max_num_tokens +vllm/distributed/device_communicators/all2all.py (MoriAll2AllManager) + max_num_inp_token_per_rank = +``` +`max_num_batched_tokens` is a **chunked-prefill scheduler** setting. Using it to size the +EP dispatch/combine buffer means a **decode** instance runs an 8192-token-wide all-to-all +**every step, per layer, x78 layers**, while decoding a handful of tokens. + +### 2.2 Signature of this bug +- fixed per-step cost: TPOT identical at concurrency 1, 4, 8, 16 +- independent of KV length: 292 ms at isl=128, 297 ms at isl=1024 +- batching still scales perfectly (con=1 -> 8 gave 8.3x throughput, TPOT flat) +- orders of magnitude above the bandwidth bound +=> "constant oversized transfer", not compute. + +### 2.3 The fix that works today (no code change) +Lower `--max-num-batched-tokens` **on the decode role only** (`models.yaml` `decode.dp:`): + +| decode mnbt | TPOT | TTFT (warm) | out tok/s | +|---|---|---|---| +| 8192 (default) | 302.5 ms | 2431 ms | 24.9 | +| **2048** | **88.0 ms** | **906 ms** | **78.8** | + +3.4x faster decode, and it also *improved* TTFT and throughput. Prefill keeps 8192 (it +genuinely dispatches wide batches). + +### 2.4 What does NOT work (dead ends — do not repeat) +- **`VLLM_MORI_MAX_TOKENS_PER_RANK` alone** (512 or 2048): device assert at boot + ``` + mori .../dispatch_combine/intranode.hpp:134 + `destTokId < config.MaxNumTokensToRecv() && + "Total recv token overflow: increase maxTotalRecvTokens"' + ``` + because vLLM's **profiling/warmup dummy run deliberately pushes + `max_num_batched_tokens` (8192) tokens** through the model. The EP buffer must survive + that even though steady-state decode never needs it. +- **`max_total_recv_tokens` to decouple recv from send: IMPOSSIBLE in current mori.** + ``` + MaxNumTokensToRecvPerRank(): + if maxTotalRecvTokens > 0: + perRank = ceil(maxTotalRecvTokens / worldSize) + return perRank < maxNumInpTokenPerRank ? perRank : maxNumInpTokenPerRank # min() + return maxNumInpTokenPerRank + ``` + It returns **min(perRank, send_width)** — it can only *lower* recv capacity, never raise + it above the send width. `send=1024, recv=65536` behaves identically to leaving it unset. + Recv capacity is structurally bounded by send width. + +### 2.5 The proper upstream fix (not yet done) +Bound the **profiling/dummy run on a decode instance by `max_num_seqs`** instead of +`max_num_batched_tokens`. Then the EP buffer can be narrow and profiling never exceeds it. +This lives entirely in vLLM. (Alternative: a mori change allowing recv > send.) + +### 2.6 MoRI has no tuning knobs (parity gap vs DeepEP) +DeepEP exposes `VLLM_DEEPEP_BUFFER_SIZE_MB`; MoRI hardcodes `warp_num_per_block`, +`block_num`, `rdma_block_num` and inherits its token width from an unrelated scheduler +default. Added `VLLM_MORI_*` knobs for parity (all defaulting to current values). + +--- + +## 3. Accuracy: the DSA sparse-index sentinel landmine + +**Never set the invalid/OOB sentinel to `-1`** in +`_convert_req_index_to_global_index_kernel` +(`vllm/v1/attention/backends/mla/rocm_aiter_mla_sparse.py`). It must be `0`. + +Why: aiter's `mla_decode_fwd` **dereferences** `paged_kv_indices`, so `-1` becomes +`kv_cache + (-1)*stride` -> page-aligned GPU memory access fault -> worker dies -> gloo +DP all-reduce collapse. `0` is masked out by `paged_kv_indptr`/`last_page_len`. + +**Why it hides:** it only fires at **disaggregated long context**. Decode inherits the +prefill's long `seq_lens` -> `generate_sparse_seqlen` widens `paged_kv_indptr` -> `-1` +padding entries land *inside* a live indptr range and get dereferenced. Short prompts and +non-disagg decode never emit an in-range `-1`, so it passes every quick test. +`HIP_LAUNCH_BLOCKING=1` does **not** help (data-dependent OOB, not an async race). + +--- + +## 4. Debugging playbook — isolate the layer before optimising + +The ladder that found the perf bug, in order of cheapness: + +| Suspect | Test | What it showed | +|---|---|---| +| router / KV transfer | bench **prefill-direct** (`:20005`) vs via router (`:30000`) | 341 ms vs 303 ms TPOT -> router + MoRI-IO exonerated (router *does* add TTFT) | +| JIT warmup | re-run on a 30-min-warm server | identical -> not warmup | +| cudagraph config | grep `cudagraph_mode`, `Graph capturing finished`, GiB captured | config correct | +| KV length | sweep isl 128 / 1024 / 8192 at con=1 | flat -> fixed floor, not KV | +| batching | con=1 vs 8 | TPOT flat, throughput 8.3x -> batching fine | +| all2all backend | force `PREFILL_MORI_BACKEND=mori_low_latency` | still broken -> not an HT-kernel bug | + +**Also:** compare against a *known-good* build. Diffing our stack against the passing +v0.25 blog stack showed MoRI, aiter and the recipe were **byte-identical** — only the base +image and vLLM differed. That single comparison exonerated two whole components. + +--- + +## 5. Operational gotchas + +### 5.1 Three caches, three different rules +| cache | keyed by | bake into image? | +|---|---|---| +| `aiter_jit` | (aiter commit, gfx arch) | **yes** — topology-independent, and it is the cold-boot long pole | +| `vllm` (torch.compile/inductor), `triton` | model cfg + batch sizes + cudagraph mode + **topology** (EP8/16/32) | **no** — a stale graph is a real trap; wipe when changing shapes | +| `comgr` | ROCm code objects | harmless | + +Host persistence: `VLLM_CACHE_PERSIST=1` -> +`/mnt/m2m_nobackup/$USER/vllm_jit_cache/a256` -> `/opt/vllm_cache`. +**Keyed by image ID**, so every new image = full cold rebuild. + +If baking `aiter_jit` into an image: **scrub `lock_module_*`, `.ninja_log`, `/root/.mori`, +`/tmp/mori_jit_*` first**, or a fresh container waits on a baton nobody holds -> boot hang. +Verify: a fresh container must start with **zero** ninja/hipcc/clang processes. + +### 5.2 Measured boot times (8 nodes, model on local NVMe) +| scenario | time | +|---|---| +| cold aiter, 2 nodes | ~25 min | +| cold aiter, 4 nodes (2 cold decode) | ~41 min | +| cold aiter, 8 nodes (4 cold decode) | ~106 min | +| **warm cache, 8 nodes** | **~11 min** | + +Changing `max_num_batched_tokens` invalidates the torch.compile cache -> full recompile. + +### 5.3 The aiter baton lock looks like a hang but isn't +`[aiter] waiting for baton release at /opt/vllm_cache/aiter_jit/build/lock_` — +one worker compiles, the rest block. Diagnose by counting build procs *inside the +container*: 100-160 = actively building; **0 on all nodes** = genuinely wedged. + +### 5.4 Readiness: don't count "Application startup complete" +Only nodes running an API server print it. For 2P/2D and 4P/4D the non-master DP ranks +never will. Judge readiness by: router `All servers healthy` + `Graph capturing finished` ++ `GPU KV cache size` per node. + +### 5.5 Verify teardown on EVERY node +`docker rm -f` can leave a container alive on one node; a stale container then collides +with the new run (observed: 2 containers on one node -> prefill shut down mid-boot). +Always re-check `docker ps -q | wc -l == 0` everywhere before relaunching. + +### 5.6 Per-role env: `models.yaml env:` applies to BOTH roles +Prefill and decode often need **opposite** values (e.g. EP buffer width). The pattern is +`PREFILL_*` / `DECODE_*` keys in `env:`, split inside `connectors/moriio.sh` (mirrors the +existing `PREFILL_MORI_BACKEND` / `DECODE_MORI_BACKEND`). Verify it landed by reading +`/proc//environ` **inside the container** — the value is exported into the +server process, not the container shell, so `docker exec env` shows nothing. + +### 5.7 RDMA fabric +- GID index **3** = RoCEv2 IPv4 (check `show_gids` / `sysfs .../gid_attrs/types`). +- Restrict `MORI_RDMA_DEVICES` / `NCCL_IB_HCA` to the 8 GPU-local NICs; leave the mgmt NICs + out or QPs try to form over a non-routable fabric -> `ibverbs.cpp:189 Connection timed out`. +- NCCL/GLOO control sockets on `eth0` (mgmt); KV data on the RDMA NICs. +- **Verify the fabric before blaming code**: `ping -I rdma0` matrix, then `ib_write_bw` + (healthy pair measured 386 Gb/s). A node can be `alloc` in SLURM with a **dead** fabric — + SLURM does not detect this. One dead node cost a whole 2P/2D campaign. +- `ib_write_bw` needs both endpoints co-alive: run it as a single 2-task srun step, not two + separate sruns (a backgrounded server dies when its srun returns). + +### 5.8 Images +Push to a registry (`docker push`) rather than relying on `docker save | ssh | docker load` +serially — parallel `docker pull` across 7 nodes is far faster and removes the +"image missing on one node" failure that silently stalls the launcher barrier. + +--- + +## 6. Open items (state as of 2026-08-16) + +- **4P/4D EP32 silent output corruption.** Cluster boots healthy (memfault=0), but output is + garbage. MoRI + aiter + recipe are byte-identical to the passing v0.25 stack; only base + + vLLM differ => **v0.27 regression**. Both `mori_high_throughput` and `mori_low_latency` + corrupt at EP32 while both are clean at EP8/EP16 => an EP-width (>16 ranks) issue, not a + kernel-specific one. Next probe: 3P/3D (EP24) to test "any EP>16" vs "exactly 32". +- **Proper upstream fix for §2.5** (bound decode profiling by `max_num_seqs`). +- **Prewarmed image** (§5.1) to kill the cold-boot and first-request-JIT costs. + +--- + +## 7. One-line summary of the two bugs found + +1. **Accuracy:** a `-1` sentinel that aiter dereferences -> GPU fault, but only at disagg + long context. Use `0`. +2. **Perf:** the MoE all-to-all buffer is sized from a chunked-prefill *scheduler* knob, so + decode moves an 8192-wide buffer every step. Lower `--max-num-batched-tokens` on the + decode role: **302 ms -> 88 ms TPOT**. From 550603f8db531c2b970a92722e37add83ec98a2f Mon Sep 17 00:00:00 2001 From: raviguptaamd Date: Sun, 16 Aug 2026 04:33:07 +0000 Subject: [PATCH 08/41] [GLM-5.1] Remove dead-end MoRI EP env plumbing and its misleading comments Review cleanup. The VLLM_MORI_MAX_TOKENS_PER_RANK / VLLM_MORI_MAX_TOTAL_RECV_TOKENS per-role plumbing was written while chasing the decode-TPOT bug and is NOT what fixed it (the fix is `--max-num-batched-tokens 2048` on decode.dp). Worse, the comments asserted that max_total_recv_tokens keeps recv capacity large enough for vLLM's profiling dummy run -- which is false and was disproved by measurement: mori's MaxNumTokensToRecvPerRank() = min(ceil(maxTotalRecvTokens / worldSize), maxNumInpTokenPerRank) is a min(), so maxTotalRecvTokens can only LOWER recv capacity, never raise it above the send width. Anyone following those comments and setting the knobs would hit "Total recv token overflow" at boot (observed at 512, 2048, and with recv=65536). Removed: the per-role export block in moriio.sh, the stale models.yaml comment block, and the six dead keys from _RECIPE_ENV_KEYS. Replaced with a short NOTE in moriio.sh pointing at the real fix and at skills_vllm_disagg.md for the measurements and dead ends. No functional change to the validated configuration: the knobs defaulted to 0/unset, so the b8 runs never exercised them. Co-Authored-By: Claude --- .../vllm_dissag/.nfs0000000016f44d2b00008188 | 685 ++++++++++++++++++ scripts/vllm_dissag/connectors/moriio.sh | 27 +- scripts/vllm_dissag/models.yaml | 12 +- scripts/vllm_dissag/run_xPyD_models.slurm | 2 +- 4 files changed, 697 insertions(+), 29 deletions(-) create mode 100755 scripts/vllm_dissag/.nfs0000000016f44d2b00008188 diff --git a/scripts/vllm_dissag/.nfs0000000016f44d2b00008188 b/scripts/vllm_dissag/.nfs0000000016f44d2b00008188 new file mode 100755 index 00000000..52d046a1 --- /dev/null +++ b/scripts/vllm_dissag/.nfs0000000016f44d2b00008188 @@ -0,0 +1,685 @@ +#!/bin/bash +#SBATCH --job-name=vllm-pd # Specify a custom string for your slurm batch job +#SBATCH -N 2 # Default 2 nodes (1P/1D); override with sbatch -N for larger topologies +#SBATCH --ntasks-per-node=1 +#SBATCH --spread-job +#SBATCH --gres=gpu:8 # Request 8 GPUs and 8 NICs (use --gres if specific GPU resources are needed) +#SBATCH --time=24:00:00 # Set a time limit for the job (HH:MM:SS) +#SBATCH --output="/shared_inference/%u/model_blog_logs/slurm-%j.out" +#SBATCH --error="/shared_inference/%u/model_blog_logs/slurm-%j.err" + + +# ------------------------ +# Auto-detect and cd to the script directory so that $(pwd) always +# points to the folder containing the server scripts, regardless of +# where the user called sbatch from. +# +# Priority: +# 1. BASH_SOURCE — works for direct invocation (bash script.sh) +# 2. SLURM_SUBMIT_DIR + script path — when sbatch submits from repo root, +# SLURM_SUBMIT_DIR is the CWD, not the script dir. Append the relative +# path from the SBATCH command to get the actual script directory. +# 3. SLURM_SUBMIT_DIR alone — last resort (assumes sbatch was run from +# the script directory). +# ------------------------ +_resolve_script_dir() { + # Try BASH_SOURCE first (works for direct invocation) + if [[ -n "${BASH_SOURCE[0]:-}" ]]; then + local _d + _d="$(cd "$(dirname "${BASH_SOURCE[0]}")" 2>/dev/null && pwd)" + if [[ -n "$_d" && -f "$_d/vllm_disagg.sh" ]]; then + echo "$_d" + return 0 + fi + fi + + # Try SLURM_SUBMIT_DIR + relative script path + # When madengine does: sbatch scripts/vllm_dissag/run_xPyD_models.slurm + # SLURM_SUBMIT_DIR = repo root, so we need to append the dirname + if [[ -n "${SLURM_SUBMIT_DIR:-}" ]]; then + local _candidate="$SLURM_SUBMIT_DIR/scripts/vllm_dissag" + if [[ -f "$_candidate/vllm_disagg.sh" ]]; then + echo "$_candidate" + return 0 + fi + # Maybe they ran sbatch from the script dir itself + if [[ -f "$SLURM_SUBMIT_DIR/vllm_disagg.sh" ]]; then + echo "$SLURM_SUBMIT_DIR" + return 0 + fi + fi + + # Fallback + echo "." + return 1 +} + +SCRIPT_DIR="$(_resolve_script_dir)" +cd "$SCRIPT_DIR" || { echo "Error: cannot cd to $SCRIPT_DIR" >&2; exit 1; } + +REQUIRED_FILES=("vllm_disagg.sh" "parallelism.sh" "connectors/rixl.sh" "connectors/moriio.sh" "models.yaml" "benchmark_xPyD.sh" "parse_to_csv.py" "socket_barrier.py" "socket_wait.py" "connectors/moriio.env" "connectors/rixl.env") +for f in "${REQUIRED_FILES[@]}"; do + if [[ ! -f "$f" ]]; then + echo "Error: Required file '$f' not found in $(pwd)." >&2 + echo "Please run sbatch from the scripts/vllm_dissag/ directory, e.g.:" >&2 + echo " cd MAD/scripts/vllm_dissag && sbatch run_xPyD_models.slurm" >&2 + exit 1 + fi +done +echo "Running from: $(pwd)" + +# ------------------------------------------------------------------------------ +# models.yaml env precedence: capture which recipe knobs the USER explicitly set +# at submit time. The driver (vllm_disagg.sh) uses this to let models.yaml `env:` +# OVERRIDE image-baked ENV defaults (e.g. a DeepSeek-tuned image bakes +# KV_BLOCK_SIZE=16 / VLLM_ROCM_USE_AITER_MLA=0, which would otherwise shadow a +# model's own recipe — GLM-5.1 DSA needs block=1 + AITER MLA on), while a genuine +# submit-time `-e VAR=...` still wins. Precedence: image-baked < models.yaml < submit -e. +# Captured HERE (before the slurm sets any defaults) so it reflects user intent only. +_RECIPE_ENV_KEYS="VLLM_USE_V1 DECODE_MORI_MAX_TOTAL_RECV_TOKENS PREFILL_MORI_MAX_TOTAL_RECV_TOKENS DECODE_MORI_MAX_TOKENS_PER_RANK PREFILL_MORI_MAX_TOKENS_PER_RANK VLLM_MORI_MAX_TOKENS_PER_RANK VLLM_MORI_WARP_NUM_PER_BLOCK VLLM_MORI_BLOCK_NUM VLLM_MORI_RDMA_BLOCK_NUM VLLM_USE_LAYERNAME VLLM_ROCM_USE_AITER VLLM_ROCM_USE_AITER_RMSNORM VLLM_ROCM_USE_AITER_MLA KV_BLOCK_SIZE KV_CACHE_DTYPE KV_CACHE_MEMORY_BYTES GPU_MEMORY_UTILIZATION VLLM_CUDAGRAPH_MODE PREFILL_CUDAGRAPH_MODE DECODE_CUDAGRAPH_MODE CUDAGRAPH_CAPTURE_SIZES VLLM_ALL2ALL_BACKEND PREFILL_MORI_BACKEND DECODE_MORI_BACKEND MORI_SHMEM_HEAP_SIZE" +MODELS_YAML_PROTECT="" +for _k in $_RECIPE_ENV_KEYS; do + [ -n "${!_k+x}" ] && MODELS_YAML_PROTECT="${MODELS_YAML_PROTECT} ${_k}" +done +export MODELS_YAML_PROTECT="${MODELS_YAML_PROTECT# }" +echo "models.yaml protect-list (submit-time overrides): '${MODELS_YAML_PROTECT}'" + +# ------------------------ +# Print current time in UTC and PST formats +# ------------------------ +echo "=== Job Start Time ===" +echo "UTC Time: $(TZ=UTC date '+%Y-%m-%d %H:%M:%S %Z')" +echo "PST Time: $(TZ=America/Los_Angeles date '+%Y-%m-%d %H:%M:%S %Z')" +echo "=======================" +echo "" + +# Define valid model names (must have a models.yaml entry) +VALID_MODELS=( \ + "Llama-3.1-405B-Instruct-FP8-KV" \ + "amd-Llama-3.3-70B-Instruct-FP8-KV" \ + "DeepSeek-V3" \ + "DeepSeek-V3-5layer" \ + "gpt-oss-120b" \ + "DeepSeek-R1" \ + "Qwen3-32B" \ + "Qwen3-30B-A3B" \ + "GLM-5.1-FP8" \ +) + +# Models allowed for CONNECTOR=moriio WIDE_EP=1 (MoRI-EP; legacy RUN_MORI=1) +MORI_EP_VALID_MODELS=( \ + "DeepSeek-V3" \ + "DeepSeek-V3-5layer" \ + "DeepSeek-R1" \ + "GLM-5.1-FP8" \ +) + +# Models allowed for CONNECTOR=rixl WIDE_EP=1 EP_BACKEND=deepep (legacy RUN_DEEPEP=1) +DEEPEP_VALID_MODELS=( \ + "DeepSeek-V3" \ + "DeepSeek-V3-5layer" \ + "DeepSeek-R1" \ +) + +MODEL_NAME="${MODEL_NAME:-None}" + +validate_model_name() { + local is_valid_model=false + + for model in "${VALID_MODELS[@]}"; do + if [[ "$MODEL_NAME" == "$model" ]]; then + is_valid_model=true + break + fi + done + + if ! $is_valid_model; then + printf "Error: Invalid MODEL_NAME: '%s'\nValid models are:\n" "$MODEL_NAME" + for model in "${VALID_MODELS[@]}"; do + printf " - %s\n" "$model" + done + exit 1 + fi + + echo "MODEL_NAME '$MODEL_NAME' is valid." + return 0 +} + +validate_model_name "${MODEL_NAME}" + +model_allows_mori_ep() { + local m="$1" + for x in "${MORI_EP_VALID_MODELS[@]}"; do + [[ "$m" == "$x" ]] && return 0 + done + return 1 +} + +model_allows_deepep() { + local m="$1" + for x in "${DEEPEP_VALID_MODELS[@]}"; do + [[ "$m" == "$x" ]] && return 0 + done + return 1 +} + +# --------------------------------------------------------------------------- +# Axis selection -> single launcher (vllm_disagg.sh). +# Two axes: CONNECTOR={rixl|moriio} x WIDE_EP={0=TP|1=wideEP}; EP_BACKEND only +# when WIDE_EP=1 (rixl->deepep, moriio->mori). Legacy RUN_MORI / RUN_DEEPEP are +# still honored via a back-compat shim. The launcher itself does the final +# CONNECTOR/WIDE_EP/EP_BACKEND validation; here we resolve them + gate the model +# against the existing allowlists. +# --------------------------------------------------------------------------- +RUN_FILE="vllm_disagg.sh" +_run_mori="${RUN_MORI:-0}" +_run_deepep="${RUN_DEEPEP:-0}" + +if [[ "$_run_mori" == "1" && "$_run_deepep" == "1" ]]; then + echo "Error: Both RUN_MORI and RUN_DEEPEP are set to 1. Set only one." >&2 + exit 1 +fi + +# Back-compat: legacy flags map onto the axes when CONNECTOR is not set explicitly. +if [[ -z "${CONNECTOR:-}" ]]; then + if [[ "$_run_mori" == "1" ]]; then + CONNECTOR=moriio; WIDE_EP="${WIDE_EP:-1}"; EP_BACKEND="${EP_BACKEND:-mori}" + elif [[ "$_run_deepep" == "1" ]]; then + CONNECTOR=rixl; WIDE_EP="${WIDE_EP:-1}"; EP_BACKEND="${EP_BACKEND:-deepep}" + else + # Default keeps the historical "no flags" behavior: rixl + TP. + CONNECTOR=rixl; WIDE_EP="${WIDE_EP:-0}" + fi +fi +WIDE_EP="${WIDE_EP:-0}" + +# Models that ONLY run wideEP (DP/EP), never TP: the DeepSeek family is served with +# the MoRI-EP / DeepEP recipe (block=16, MLA off, per-role cudagraph). Running them +# in TP mode is unsupported — the TP argv would double the model's own +# --compilation-config and drop the mandatory +quant_fp8 op. Reject early. +# GLM-5.1-FP8 (GlmMoeDsaForCausalLM, MLA+DSA) is validated only under MoRI-EP +# wideEP disagg (block=1, AITER sparse MLA on, per-role all2all). The moriio+TP +# ("Stage B") path is untested for DSA, so reject WIDE_EP=0 for it too. +WIDE_EP_ONLY_MODELS=( "DeepSeek-V3" "DeepSeek-V3-5layer" "DeepSeek-R1" "GLM-5.1-FP8" ) +model_is_wide_ep_only() { + local m="$1" + for x in "${WIDE_EP_ONLY_MODELS[@]}"; do [[ "$m" == "$x" ]] && return 0; done + return 1 +} + +# Model allowlist gate (kept, sglang-style). wideEP modes use the per-backend +# allowlists; TP uses the global VALID_MODELS (already validated above) minus the +# wideEP-only models. +if [[ "$WIDE_EP" == "0" ]]; then + if model_is_wide_ep_only "$MODEL_NAME"; then + echo "Error: MODEL_NAME '$MODEL_NAME' is wideEP-only (set WIDE_EP=1). TP mode is not supported for it." >&2 + printf "wideEP-only models:\n"; for m in "${WIDE_EP_ONLY_MODELS[@]}"; do printf " - %s\n" "$m"; done + exit 1 + fi +elif [[ "$WIDE_EP" == "1" && "${CONNECTOR}" == "moriio" ]]; then + if ! model_allows_mori_ep "$MODEL_NAME"; then + echo "Error: CONNECTOR=moriio WIDE_EP=1 but MODEL_NAME '$MODEL_NAME' is not in MORI_EP_VALID_MODELS" >&2 + printf "MoRI EP allowed models:\n"; for m in "${MORI_EP_VALID_MODELS[@]}"; do printf " - %s\n" "$m"; done + exit 1 + fi +elif [[ "$WIDE_EP" == "1" && "${CONNECTOR}" == "rixl" ]]; then + if ! model_allows_deepep "$MODEL_NAME"; then + echo "Error: CONNECTOR=rixl WIDE_EP=1 (deepep) but MODEL_NAME '$MODEL_NAME' is not in DEEPEP_VALID_MODELS" >&2 + printf "DeepEP allowed models:\n"; for m in "${DEEPEP_VALID_MODELS[@]}"; do printf " - %s\n" "$m"; done + exit 1 + fi +fi + +export CONNECTOR WIDE_EP EP_BACKEND +echo "Launcher: $RUN_FILE (CONNECTOR=${CONNECTOR} WIDE_EP=${WIDE_EP} EP_BACKEND=${EP_BACKEND:-}) for model '$MODEL_NAME'" + +# --------------------------------------------------------------------------- +# Connector platform env: per-connector .env holds the ROCm-7.2.3 +# runtime env that MUST reach the container at PID 1 (e.g. expandable_segments:False +# for GPU-RDMA registration). Source the resolved connector's file and collect its +# KEY=VALUE lines into CONNECTOR_ENV_ARGS as `-e KEY=${KEY:-VALUE}` pairs, so a +# submit-time export of the same name still overrides. Forwarded in the docker run. +# --------------------------------------------------------------------------- +CONNECTOR_ENV_FILE="${SCRIPT_DIR}/connectors/${CONNECTOR}.env" +CONNECTOR_ENV_ARGS="" +if [[ -f "$CONNECTOR_ENV_FILE" ]]; then + echo "Loading connector platform env: $CONNECTOR_ENV_FILE" + while IFS= read -r _line; do + [[ "$_line" =~ ^[[:space:]]*# || -z "${_line// }" ]] && continue + _k="${_line%%=*}"; _v="${_line#*=}" + CONNECTOR_ENV_ARGS+=" -e ${_k}=${!_k:-$_v}" # submit-time export of $_k wins + done < "$CONNECTOR_ENV_FILE" +else + echo "WARN: connector env file not found: $CONNECTOR_ENV_FILE" >&2 +fi + +if [[ -z "${DOCKER_IMAGE_NAME:-}" ]]; then + echo "Error: DOCKER_IMAGE_NAME is not set. Please export DOCKER_IMAGE_NAME before running." >&2 + echo " There is no public prebuilt image. Build your own from the provided Dockerfile:" >&2 + echo " docker build -f docker/vllm_disagg_inference.ubuntu.amd.Dockerfile \\" >&2 + echo " -t /vllm-disagg:local . # all connectors (WITH_NIXL=1 default)" >&2 + echo " (add --build-arg WITH_NIXL=0 for a lean MoRI-EP-only image)" >&2 + echo " then: export DOCKER_IMAGE_NAME=/vllm-disagg:local" >&2 + exit 1 +fi +export DOCKER_IMAGE_NAME + +# Set current directory to be REPO directory with all relevant scripts +NIXL_REPO_DIR=$(pwd) +LOG_PATH="${LOG_PATH:-/shared_inference/${USER}/model_blog_logs}" + +xP="${xP:-1}" #-> Number of Prefill Servers +yD="${yD:-1}" #-> Number of Decode Servers + +MODEL_DIR="${MODEL_DIR:-"/shared_inference/models_blog/"}" + + +# ------------------------ +# Model path validation and selection across all nodes +# ------------------------ +echo "Looking for model: $MODEL_NAME" +echo "Checking model availability across all allocated nodes..." + +# Get all allocated nodes +ALL_NODES=$(scontrol show hostnames "$SLURM_JOB_NODELIST") +TOTAL_NODES=$(echo "$ALL_NODES" | wc -l) + +echo "Total allocated nodes: $TOTAL_NODES" +echo "Nodes: $(echo "$ALL_NODES" | tr '\n' ' ')" + +# Function to check model path on all nodes +check_model_path() { + local path=$1 + local check_name=$2 + + echo "Checking $check_name: $path" + + # Run check on all nodes in parallel + srun --nodes=$SLURM_NNODES --ntasks=$SLURM_NNODES /bin/bash -c " + if [ -d '$path' ]; then + echo \"\$(hostname): ✓ Found $path\" + exit 0 + else + echo \"\$(hostname): ✗ Missing $path\" + exit 1 + fi + " + + # Check if all nodes succeeded (exit code 0) + local exit_code=$? + if [ $exit_code -eq 0 ]; then + echo "✓ $check_name available on ALL nodes" + return 0 + else + echo "✗ $check_name NOT available on all nodes" + return 1 + fi +} + +# Check /mnt/m2m_nobackup/models_blog first +MODEL_PATH_1="/mnt/m2m_nobackup/models_blog/$MODEL_NAME" +if check_model_path "$MODEL_PATH_1" "/mnt/m2m_nobackup/models_blog"; then + MODEL_PATH="$MODEL_PATH_1" + echo "" + echo "✓ Selected MODEL_PATH: $MODEL_PATH (available on all nodes)" +# Check /shared-inference/models_blog +elif check_model_path "/shared_inference/models_blog/$MODEL_NAME" "/shared_inference/models_blog"; then + MODEL_PATH="/shared_inference/models_blog/$MODEL_NAME" + echo "" + echo "✓ Selected MODEL_PATH: $MODEL_PATH (available on all nodes)" +elif check_model_path "$MODEL_DIR/$MODEL_NAME" "$MODEL_DIR"; then + MODEL_PATH="$MODEL_DIR/$MODEL_NAME" + echo "" + echo "✓ Selected MODEL_PATH: $MODEL_PATH (available on all nodes)" +else + echo "" + echo "✗ FATAL ERROR: Model '$MODEL_NAME' not found on ALL allocated nodes in either:" + echo " - /mnt/m2m_nobackup/models_blog/$MODEL_NAME" + echo " - /shared_inference/models_blog/$MODEL_NAME" + echo "" + echo "Model must be accessible from all nodes for distributed execution." + echo "Please ensure the model is available on all allocated nodes." + exit 1 +fi + +echo "Final MODEL_PATH: $MODEL_PATH" +echo "" + + +# Calculate NUM_NODES based on xP and yD +NUM_NODES=$((xP + yD)) +echo "Calculated NUM_NODES: $NUM_NODES (xP=$xP + yD=$yD, proxy co-located on prefill master)" + +# DeepEP configuration (only exported when RUN_DEEPEP=1) +if [[ "$_run_deepep" == "1" ]]; then + export PREFILL_DEEPEP_BACKEND="${PREFILL_DEEPEP_BACKEND:-deepep_high_throughput}" + export DECODE_DEEPEP_BACKEND="${DECODE_DEEPEP_BACKEND:-deepep_low_latency}" + export ENABLE_DBO="${ENABLE_DBO:-false}" + export DBO_COMM_SMS="${DBO_COMM_SMS:-}" + export ENABLE_PROFILING="${ENABLE_PROFILING:-false}" + echo "DeepEP config: PREFILL_BACKEND=$PREFILL_DEEPEP_BACKEND DECODE_BACKEND=$DECODE_DEEPEP_BACKEND DBO=$ENABLE_DBO" +fi + +# ------------------------ +# Extract first NUM_NODES from SLURM allocation and update SLURM variables +# ------------------------ +echo "Original SLURM allocation:" +echo "SLURM_JOB_NODELIST: $SLURM_JOB_NODELIST" +echo "SLURM_NNODES: $SLURM_NNODES" +echo "SLURM_NTASKS: $SLURM_NTASKS" + +# Get the full nodelist and extract first NUM_NODES +FULL_NODELIST=$(scontrol show hostnames "$SLURM_JOB_NODELIST") +SELECTED_NODES=$(echo "$FULL_NODELIST" | head -n $NUM_NODES) +NEW_SLURM_NODELIST=$(echo "$SELECTED_NODES" | paste -sd,) + +# Update SLURM environment variables +export SLURM_NNODES=$NUM_NODES +export SLURM_NTASKS=$NUM_NODES +export SLURM_JOB_NUM_NODES=$NUM_NODES +export SLURM_NPROCS=$NUM_NODES +export SLURM_JOB_NODELIST="$NEW_SLURM_NODELIST" +export SLURM_NODELIST="$NEW_SLURM_NODELIST" + +# Keep other SLURM variables as they were or set defaults +export SLURM_TASKS_PER_NODE="1(x$NUM_NODES)" + +export SLURM_CLUSTER_NAME="${SLURM_CLUSTER_NAME}" +export SLURM_JOB_CPUS_PER_NODE="${SLURM_JOB_CPUS_PER_NODE}" +export SLURM_JOB_PARTITION="${SLURM_JOB_PARTITION}" +export SLURM_JOBID="${SLURM_JOBID:-$SLURM_JOB_ID}" +export SLURM_JOB_QOS="${SLURM_JOB_QOS:-normal}" +export SLURM_JOB_ACCOUNT="${SLURM_JOB_ACCOUNT}" +export SLURM_NTASKS_PER_NODE=1 +export SLURM_SUBMIT_HOST="${SLURM_SUBMIT_HOST}" +export SLURM_JOB_ID="${SLURM_JOB_ID}" +export SLURM_CONF="${SLURM_CONF:-/etc/slurm/slurm.conf}" +export SLURM_JOB_NAME="${SLURM_JOB_NAME:-1p1d_bench-serving}" + +echo "" +echo "Updated SLURM Environment Variables:" +echo "SLURM_JOB_ID: $SLURM_JOB_ID" +echo "SLURM_JOB_NODELIST: $SLURM_JOB_NODELIST" +echo "SLURM_NNODES: $SLURM_NNODES" +echo "SLURM_NTASKS: $SLURM_NTASKS" +echo "SLURM_TASKS_PER_NODE: $SLURM_TASKS_PER_NODE" +echo "SLURM_JOB_CPUS_PER_NODE: $SLURM_JOB_CPUS_PER_NODE" +echo "SLURM_JOB_PARTITION: $SLURM_JOB_PARTITION" +echo "SLURM_JOB_NUM_NODES: $SLURM_JOB_NUM_NODES" +echo "SLURM_JOBID: $SLURM_JOBID" +echo "SLURM_JOB_QOS: $SLURM_JOB_QOS" +echo "SLURM_NODELIST: $SLURM_NODELIST" +echo "SLURM_JOB_ACCOUNT: $SLURM_JOB_ACCOUNT" +echo "SLURM_NPROCS: $SLURM_NPROCS" +echo "SLURM_SUBMIT_HOST: $SLURM_SUBMIT_HOST" +echo "SLURM_CONF: $SLURM_CONF" +echo "SLURM_JOB_NAME: $SLURM_JOB_NAME" +echo "SLURM_NTASKS_PER_NODE: $SLURM_NTASKS_PER_NODE" +#echo "SLURM_SUBMIT_DIR: $SLURM_SUBMIT_DIR" +echo "SLURM_CLUSTER_NAME: $SLURM_CLUSTER_NAME" +echo "ulimit: $(ulimit -a)" +echo "" +echo "Selected nodes for execution:" +echo "$SELECTED_NODES" +echo "" + +# Node information +USER_NAME=$(whoami) +MASTER_NODE=$(echo "$SELECTED_NODES" | head -n 1) +# Pick the routable fabric IP, not just hostname -I's first entry. These nodes expose +# multiple NICs (e.g. a 10.224.x overlay listed BEFORE the routable 10.158.x fabric); +# taking $1 blindly can advertise an unreachable addr -> prefill/decode barrier hangs +# "Waiting for nodes" forever. Prefer FABRIC_SUBNET (default 10.158.), fall back to $1. +FABRIC_SUBNET="${FABRIC_SUBNET:-10.158.}" +# From a "hostname -I" line, return the first IP on FABRIC_SUBNET, else the first IP. +_pick_fabric_ip() { + awk -v pfx="$FABRIC_SUBNET" '{f=$1; for(i=1;i<=NF;i++) if(index($i,pfx)==1){f=$i; break} print f}' +} +MASTER_ADDR=$(srun --nodes=1 --ntasks=1 --time=00:20:00 --nodelist="$MASTER_NODE" bash -c 'hostname -I' | _pick_fabric_ip) +MASTER_PORT=39566 # Choose an open port + +IPS=() + +for NODE in $SELECTED_NODES; do + IP=$(srun --nodes=1 --ntasks=1 --time=00:20:00 --nodelist="$NODE" bash -c 'hostname -I' | _pick_fabric_ip) + IPS+=("$IP") +done + +echo "Selected node IPs: ${IPS[*]}" | sed 's/ /,/g' + +NIXL_COOKBOOK_PATH="/opt/nixl-vllm-cookbook" +BENCHMARK_ITR="${BENCHMARK_ITR:-1}" +BENCHMARK_CON="${BENCHMARK_CON:-}" +BENCHMARK_COMBINATIONS="${BENCHMARK_COMBINATIONS:-}" + +# Benchmark script selector: BENCHMARK_SCRIPT tag -> file run by the launcher. +# sweep (default) -> benchmark_xPyD.sh (general concurrency sweep) +# long_context -> benchmark_long_context.sh (per-shape warmup, c=1-first) +# keepalive -> keepalive_bench.sh (hold server up KEEPALIVE_MINS +# for external accuracy probes) +BENCHMARK_SCRIPT="${BENCHMARK_SCRIPT:-sweep}" +case "$BENCHMARK_SCRIPT" in + sweep) BENCHMARK_SCRIPT_FILE="benchmark_xPyD.sh" ;; + long_context) BENCHMARK_SCRIPT_FILE="benchmark_long_context.sh" ;; + keepalive) BENCHMARK_SCRIPT_FILE="keepalive_bench.sh" ;; + *) echo "Error: invalid BENCHMARK_SCRIPT='$BENCHMARK_SCRIPT' (valid: sweep, long_context, keepalive)" >&2; exit 1 ;; +esac +if [[ ! -f "$BENCHMARK_SCRIPT_FILE" ]]; then + echo "Error: selected benchmark script '$BENCHMARK_SCRIPT_FILE' not found in $(pwd)." >&2 + exit 1 +fi +echo "BENCHMARK_SCRIPT='$BENCHMARK_SCRIPT' -> $BENCHMARK_SCRIPT_FILE" + +NNODES=$NUM_NODES + +echo "MASTER_NODE is ${MASTER_NODE}" +echo "MASTER_ADDR is ${MASTER_ADDR}" +echo "MASTER_PORT is ${MASTER_PORT}" +echo "NNODES is ${NNODES}" +echo "REPO Directory is ${NIXL_REPO_DIR}" + +if [ ! -d "$LOG_PATH" ]; then + mkdir -p "$LOG_PATH" + echo "Created directory: $LOG_PATH" +else + echo "Directory already exists: $LOG_PATH" +fi + +export CONNECTOR_ENV_ARGS="$CONNECTOR_ENV_ARGS" +export LOG_PATH=$LOG_PATH +export NIXL_REPO_DIR=$NIXL_REPO_DIR +export NIXL_COOKBOOK_PATH=$NIXL_COOKBOOK_PATH +export NNODES=$NNODES +export MASTER_ADDR=$MASTER_ADDR +export MASTER_PORT=$MASTER_PORT +export MODEL_PATH=$MODEL_PATH +export xP=$xP +export yD=$yD +export MODEL_NAME=$MODEL_NAME +export USER_NAME=$USER_NAME +export IPADDRS="$(echo "${IPS[*]}" | sed 's/ /,/g')" +export BENCHMARK_ITR=$BENCHMARK_ITR +export BENCHMARK_CON="${BENCHMARK_CON}" +export BENCHMARK_COMBINATIONS="${BENCHMARK_COMBINATIONS}" +export BENCHMARK_SCRIPT_FILE="${BENCHMARK_SCRIPT_FILE}" + +export DOCKER_CONT_NAME="container_${MODEL_NAME}_${SLURM_JOB_ID}" +export RUN_FILE_FULL="$NIXL_COOKBOOK_PATH/${RUN_FILE}" + +# Use only the selected nodes for srun execution +SELECTED_NODELIST_SRUN=$(echo "$SELECTED_NODES" | paste -sd,) + +srun --nodelist="$SELECTED_NODELIST_SRUN" bash -c ' +echo "Rank $SLURM_PROCID on $(hostname)"; +docker ps -q | xargs --no-run-if-empty docker stop; +docker rm -f $DOCKER_CONT_NAME 2>/dev/null || true; +fuser -k 5000/tcp 2>/dev/null || true; +fuser -k 2222/tcp 2>/dev/null || true; +fuser -k 15000/tcp 2>/dev/null || true; +sleep 2; +docker pull $DOCKER_IMAGE_NAME 2>/dev/null || true; + +# --- Create host-local compilation cache dirs (ext4, survives container restarts) --- +mkdir -p /tmp/vllm_cache/{aiter_jit,triton,vllm,comgr} 2>/dev/null || true; + +# --- Persistent JIT cache mount --- +# The image points AITER_JIT_DIR/TRITON_CACHE_DIR/VLLM_CACHE_ROOT/COMGR_CACHE_DIR at +# /opt/vllm_cache. Mount a host dir there so AITER CK kernels compile ONCE and are reused +# across runs. Cold compile is ~15 min for the DeepSeek MoE/FP8 GEMM kernel set; a warm +# boot is ~1 min. Keyed by image ID so a new image starts a fresh cache and does not reuse +# stale-ABI shared objects. Host dir on local NVMe. Override JIT_CACHE_HOST, or set +# JIT_CACHE_PERSIST=0 to disable and fall back to the image empty in-container cache. +# NOTE: this whole section runs inside a single-quoted `srun bash -c '...'`, so avoid +# single quotes here; the image-id hash is extracted with tr, not sed. +if [ "${JIT_CACHE_PERSIST:-1}" = "1" ]; then + _IMG_RAW=$(docker image inspect --format "{{.Id}}" "$DOCKER_IMAGE_NAME" 2>/dev/null); + _IMG_KEY=$(printf "%s" "$_IMG_RAW" | tr -cd "a-f0-9" | cut -c1-12); + _IMG_KEY="${_IMG_KEY:-noimg}"; + _JIT_CACHE_HOST="${JIT_CACHE_HOST:-/mnt/m2m_nobackup/${USER}/vllm_jit_cache/${_IMG_KEY}}"; + mkdir -p "$_JIT_CACHE_HOST"/{aiter_jit,triton,vllm,comgr} 2>/dev/null || true; + _JIT_CACHE_MOUNT="-v ${_JIT_CACHE_HOST}:/opt/vllm_cache"; + echo "[jit-cache] persistent image ${_IMG_KEY}: ${_JIT_CACHE_HOST} to /opt/vllm_cache"; +else + _JIT_CACHE_MOUNT=""; +fi + +# --- Build host RDMA library mounts --- +_RDMA_MOUNTS="" +_LIBDIR=/usr/lib/x86_64-linux-gnu + +for _lib in libibverbs.so libibverbs.so.1 librdmacm.so librdmacm.so.1; do + [ -e "$_LIBDIR/$_lib" ] && _RDMA_MOUNTS="$_RDMA_MOUNTS -v $_LIBDIR/$_lib:$_LIBDIR/$_lib:ro" +done +for _vlib in $_LIBDIR/libibverbs.so.1.* $_LIBDIR/librdmacm.so.1.*; do + [ -e "$_vlib" ] && _RDMA_MOUNTS="$_RDMA_MOUNTS -v $_vlib:$_vlib:ro" +done + +for _pattern in libmlx5.so* libionic*.so* libbnxt_re*.so* libefa.so* libhns.so*; do + for _vlib in $_LIBDIR/${_pattern}; do + # Require a regular file AFTER symlink resolution: these mounts are built on + # ONE node but applied on ALL nodes, and vendor NIC libs (e.g. libionic.so.1) + # can be a DANGLING symlink on some nodes -> bind-mount fails "not a directory" + # -> container create exit 125. `-f` (follows symlink, requires regular file) + # skips those; the fabric in use (mlx5) is still mounted where present. + [ -f "$_vlib" ] && _RDMA_MOUNTS="$_RDMA_MOUNTS -v $_vlib:$_vlib:ro" + done +done + +[ -d "$_LIBDIR/libibverbs" ] && _RDMA_MOUNTS="$_RDMA_MOUNTS -v $_LIBDIR/libibverbs:$_LIBDIR/libibverbs:ro" +[ -d /etc/libibverbs.d ] && _RDMA_MOUNTS="$_RDMA_MOUNTS -v /etc/libibverbs.d:/etc/libibverbs.d:ro" +echo "[host-rdma] mounts: $_RDMA_MOUNTS" + +docker run --rm \ + --device /dev/dri \ + --device /dev/kfd \ + --device /dev/infiniband \ + --network host \ + --ipc host \ + --group-add video \ + --cap-add SYS_PTRACE \ + --security-opt seccomp=unconfined \ + --privileged \ + -v $HOME:$HOME \ + -v /shared_inference:/shared_inference \ + -v /mnt/m2m_nobackup:/mnt/m2m_nobackup \ + -v $HOME/.ssh:/root/.ssh \ + --shm-size ${DOCKER_SHM_SIZE:-256G} \ + --ulimit nofile=524288:524288 \ + --ulimit memlock=-1:-1 \ + -v ${LOG_PATH}:/run_logs \ + -v $NIXL_REPO_DIR:$NIXL_COOKBOOK_PATH \ + -v /tmp/vllm_cache:/tmp/vllm_cache \ + ${_JIT_CACHE_MOUNT} \ + ${GLM_KERNEL_PATCH:+-v ${GLM_KERNEL_PATCH}:/usr/local/lib/python3.12/dist-packages/vllm/v1/attention/ops/rocm_aiter_mla_sparse.py:ro} \ + ${GLM_BACKEND_PATCH:+-v ${GLM_BACKEND_PATCH}:/usr/local/lib/python3.12/dist-packages/vllm/v1/attention/backends/mla/rocm_aiter_mla_sparse.py:ro} \ + $_RDMA_MOUNTS \ + --entrypoint /bin/bash \ + -e SLURM_JOB_ID=$SLURM_JOB_ID \ + -e SLURM_JOB_NODELIST=$SLURM_JOB_NODELIST \ + -e NNODES=$NNODES \ + -e NODE_RANK=$SLURM_PROCID \ + -e MASTER_ADDR=$MASTER_ADDR \ + -e MASTER_PORT=$MASTER_PORT \ + -e MODEL_PATH=$MODEL_PATH \ + -e NIXL_COOKBOOK_PATH=$NIXL_COOKBOOK_PATH \ + -e xP=$xP \ + -e yD=$yD \ + -e USER_NAME=$USER_NAME \ + -e MODEL_NAME=$MODEL_NAME \ + -e BENCHMARK_ITR=$BENCHMARK_ITR \ + -e BENCHMARK_CON="${BENCHMARK_CON}" \ + -e BENCHMARK_COMBINATIONS="${BENCHMARK_COMBINATIONS}" \ + ${BENCHMARK_PORT:+-e BENCHMARK_PORT=$BENCHMARK_PORT} \ + ${PROXY_TYPE:+-e PROXY_TYPE=$PROXY_TYPE} \ + ${ROUTER_PORT:+-e ROUTER_PORT=$ROUTER_PORT} \ + -e IPADDRS=$IPADDRS \ + ${CONNECTOR:+-e CONNECTOR=$CONNECTOR} \ + ${WIDE_EP:+-e WIDE_EP=$WIDE_EP} \ + ${EP_BACKEND:+-e EP_BACKEND=$EP_BACKEND} \ + ${RUN_MORI:+-e RUN_MORI=$RUN_MORI} \ + ${RUN_DEEPEP:+-e RUN_DEEPEP=$RUN_DEEPEP} \ + ${VLLM_ALL2ALL_BACKEND:+-e VLLM_ALL2ALL_BACKEND=$VLLM_ALL2ALL_BACKEND} \ + ${PREFILL_MORI_BACKEND:+-e PREFILL_MORI_BACKEND=$PREFILL_MORI_BACKEND} \ + ${DECODE_MORI_BACKEND:+-e DECODE_MORI_BACKEND=$DECODE_MORI_BACKEND} \ + -e MODELS_YAML_PROTECT="${MODELS_YAML_PROTECT:-}" \ + ${GLM_PERSIST_GATE:+-e GLM_PERSIST_GATE=$GLM_PERSIST_GATE} \ + ${GLM_SKIP_PATCHERS:+-e GLM_SKIP_PATCHERS=$GLM_SKIP_PATCHERS} \ + ${KV_BLOCK_SIZE:+-e KV_BLOCK_SIZE=$KV_BLOCK_SIZE} \ + ${KV_CACHE_MEMORY_BYTES:+-e KV_CACHE_MEMORY_BYTES=$KV_CACHE_MEMORY_BYTES} \ + ${VLLM_ROCM_USE_AITER_MLA:+-e VLLM_ROCM_USE_AITER_MLA=$VLLM_ROCM_USE_AITER_MLA} \ + ${ROUTER_BINARY:+-e ROUTER_BINARY=$ROUTER_BINARY} \ + ${KV_CACHE_DTYPE:+-e KV_CACHE_DTYPE=$KV_CACHE_DTYPE} \ + ${MORIIO_TOY_PROXY:+-e MORIIO_TOY_PROXY=$MORIIO_TOY_PROXY} \ + ${BENCHMARK_SCRIPT_FILE:+-e BENCHMARK_SCRIPT_FILE=$BENCHMARK_SCRIPT_FILE} \ + ${KEEPALIVE_MINS:+-e KEEPALIVE_MINS=$KEEPALIVE_MINS} \ + ${PREFILL_CUDAGRAPH_MODE:+-e PREFILL_CUDAGRAPH_MODE=$PREFILL_CUDAGRAPH_MODE} \ + ${DECODE_CUDAGRAPH_MODE:+-e DECODE_CUDAGRAPH_MODE=$DECODE_CUDAGRAPH_MODE} \ + ${CUDAGRAPH_CAPTURE_SIZES:+-e CUDAGRAPH_CAPTURE_SIZES="$CUDAGRAPH_CAPTURE_SIZES"} \ + ${MORI_RDMA_TC:+-e MORI_RDMA_TC=$MORI_RDMA_TC} \ + ${MORI_RDMA_SL:+-e MORI_RDMA_SL=$MORI_RDMA_SL} \ + ${MORI_SHMEM_HEAP_SIZE:+-e MORI_SHMEM_HEAP_SIZE=$MORI_SHMEM_HEAP_SIZE} \ + ${PREFILL_DEEPEP_BACKEND:+-e PREFILL_DEEPEP_BACKEND=$PREFILL_DEEPEP_BACKEND} \ + ${DECODE_DEEPEP_BACKEND:+-e DECODE_DEEPEP_BACKEND=$DECODE_DEEPEP_BACKEND} \ + ${ENABLE_DBO:+-e ENABLE_DBO=$ENABLE_DBO} \ + ${DBO_COMM_SMS:+-e DBO_COMM_SMS=$DBO_COMM_SMS} \ + ${ENABLE_PROFILING:+-e ENABLE_PROFILING=$ENABLE_PROFILING} \ + ${NCCL_IB_HCA:+-e NCCL_IB_HCA=$NCCL_IB_HCA} \ + ${NCCL_IB_GID_INDEX:+-e NCCL_IB_GID_INDEX=$NCCL_IB_GID_INDEX} \ + ${NCCL_NET_GDR_LEVEL:+-e NCCL_NET_GDR_LEVEL=$NCCL_NET_GDR_LEVEL} \ + ${NCCL_CROSS_NIC:+-e NCCL_CROSS_NIC=$NCCL_CROSS_NIC} \ + ${NCCL_SOCKET_IFNAME:+-e NCCL_SOCKET_IFNAME=$NCCL_SOCKET_IFNAME} \ + ${GLOO_SOCKET_IFNAME:+-e GLOO_SOCKET_IFNAME=$GLOO_SOCKET_IFNAME} \ + -e MORI_SOCKET_IFNAME=${MORI_SOCKET_IFNAME:-eth0} \ + ${MORI_IB_GID_INDEX:+-e MORI_IB_GID_INDEX=$MORI_IB_GID_INDEX} \ + ${MORI_RDMA_DEVICES:+-e MORI_RDMA_DEVICES=$MORI_RDMA_DEVICES} \ + ${MORI_NUM_QP_PER_PE:+-e MORI_NUM_QP_PER_PE=$MORI_NUM_QP_PER_PE} \ + ${VLLM_MORIIO_QP_PER_TRANSFER:+-e VLLM_MORIIO_QP_PER_TRANSFER=$VLLM_MORIIO_QP_PER_TRANSFER} \ + ${VLLM_MORIIO_NUM_WORKERS:+-e VLLM_MORIIO_NUM_WORKERS=$VLLM_MORIIO_NUM_WORKERS} \ + -e GPU_MEMORY_UTILIZATION=${GPU_MEMORY_UTILIZATION:-0.8} \ + -e GPUS_PER_NODE=${GPUS_PER_NODE:-8} \ + ${GPU_MAX_HW_QUEUES:+-e GPU_MAX_HW_QUEUES=$GPU_MAX_HW_QUEUES} \ + ${HIP_FORCE_DEV_KERNARG:+-e HIP_FORCE_DEV_KERNARG=$HIP_FORCE_DEV_KERNARG} \ + ${HSA_NO_SCRATCH_RECLAIM:+-e HSA_NO_SCRATCH_RECLAIM=$HSA_NO_SCRATCH_RECLAIM} \ + ${VLLM_HANDSHAKE_TIMEOUT_MINS:+-e VLLM_HANDSHAKE_TIMEOUT_MINS=$VLLM_HANDSHAKE_TIMEOUT_MINS} \ + ${VLLM_ENGINE_READY_TIMEOUT_S:+-e VLLM_ENGINE_READY_TIMEOUT_S=$VLLM_ENGINE_READY_TIMEOUT_S} \ + ${ROCSHMEM_HEAP_SIZE:+-e ROCSHMEM_HEAP_SIZE=$ROCSHMEM_HEAP_SIZE} \ + ${ROCSHMEM_MAX_NUM_CONTEXTS:+-e ROCSHMEM_MAX_NUM_CONTEXTS=$ROCSHMEM_MAX_NUM_CONTEXTS} \ + ${LOG_WAIT_TIMEOUT_SECONDS:+-e LOG_WAIT_TIMEOUT_SECONDS=$LOG_WAIT_TIMEOUT_SECONDS} \ + ${TRITON_CACHE_DIR:+-e TRITON_CACHE_DIR=$TRITON_CACHE_DIR} \ + ${VLLM_CACHE_ROOT:+-e VLLM_CACHE_ROOT=$VLLM_CACHE_ROOT} \ + ${COMGR_CACHE_DIR:+-e COMGR_CACHE_DIR=$COMGR_CACHE_DIR} \ + ${AITER_JIT_DIR:+-e AITER_JIT_DIR=$AITER_JIT_DIR} \ + -e DISTRIBUTED_TIMEOUT_SECONDS=${DISTRIBUTED_TIMEOUT_SECONDS:-7200} \ + -e VLLM_RPC_TIMEOUT=${VLLM_RPC_TIMEOUT:-300000} \ + -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=${VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS:-3600} \ + ${CONNECTOR_ENV_ARGS} \ + ${VLLM_CUDAGRAPH_MODE:+-e VLLM_CUDAGRAPH_MODE=$VLLM_CUDAGRAPH_MODE} \ + ${CUDAGRAPH_CAPTURE_SIZES:+-e CUDAGRAPH_CAPTURE_SIZES="$CUDAGRAPH_CAPTURE_SIZES"} \ + --name $DOCKER_CONT_NAME \ + $DOCKER_IMAGE_NAME -c " + mkdir -p /run_logs/${SLURM_JOB_ID} + $RUN_FILE_FULL 2>&1 | tee /run_logs/${SLURM_JOB_ID}/pd_vllm_bench_NODE${SLURM_PROCID}.log + " +' +srun --nodelist="$SELECTED_NODELIST_SRUN" bash -c 'docker stop $DOCKER_CONT_NAME 2>/dev/null || true; docker rm $DOCKER_CONT_NAME 2>/dev/null || true' + diff --git a/scripts/vllm_dissag/connectors/moriio.sh b/scripts/vllm_dissag/connectors/moriio.sh index 33d5322d..e1c7a51a 100644 --- a/scripts/vllm_dissag/connectors/moriio.sh +++ b/scripts/vllm_dissag/connectors/moriio.sh @@ -274,24 +274,15 @@ connector_launch_worker() { local _all2all="${PREFILL_MORI_BACKEND}" [[ "$log_prefix" == "decode" ]] && _all2all="${DECODE_MORI_BACKEND}" - # Per-role MoRI EP buffer width. VLLM_MORI_MAX_TOKENS_PER_RANK sizes the - # dispatch/combine buffer; unset (0) it inherits max_num_batched_tokens -- a - # chunked-prefill SCHEDULER knob (8192) -- so a decode instance moves an - # 8192-token-wide buffer every step, per layer: ~302ms vs ~88ms TPOT (3.4x). - # Prefill and decode want OPPOSITE values (prefill genuinely dispatches wide - # batches and must keep the large buffer), but models.yaml env: applies to BOTH - # roles -- so split it here, mirroring PREFILL/DECODE_MORI_BACKEND above. - if [[ "$log_prefix" == "decode" ]]; then - [[ -n "${DECODE_MORI_MAX_TOKENS_PER_RANK:-}" ]] && \ - export VLLM_MORI_MAX_TOKENS_PER_RANK="${DECODE_MORI_MAX_TOKENS_PER_RANK}" - # Recv capacity must still cover vLLM's profiling dummy run (which pushes - # max_num_batched_tokens tokens) even though steady-state dispatch is narrow. - [[ -n "${DECODE_MORI_MAX_TOTAL_RECV_TOKENS:-}" ]] && \ - export VLLM_MORI_MAX_TOTAL_RECV_TOKENS="${DECODE_MORI_MAX_TOTAL_RECV_TOKENS}" - else - export VLLM_MORI_MAX_TOKENS_PER_RANK="${PREFILL_MORI_MAX_TOKENS_PER_RANK:-0}" - export VLLM_MORI_MAX_TOTAL_RECV_TOKENS="${PREFILL_MORI_MAX_TOTAL_RECV_TOKENS:-0}" - fi + # NOTE on MoRI EP buffer width: it is sized from max_num_batched_tokens + # (fused_moe/layer.py -> all2all.py max_num_inp_token_per_rank), so a decode + # instance otherwise runs an 8192-token-wide all2all every step (~302ms vs ~88ms + # TPOT). The fix is per-role `--max-num-batched-tokens` in models.yaml + # (decode.dp), NOT an env knob: mori derives recv capacity from the send width + # (MaxNumTokensToRecvPerRank returns min(ceil(maxTotalRecvTokens/ws), + # maxNumInpTokenPerRank)), so shrinking the width alone under-provisions recv and + # trips a device assert during vLLM's profiling dummy run. See + # skills_vllm_disagg.md for the measurements and the dead ends. local extra_args=() kv_args=() if [[ "$role" == "master" ]]; then diff --git a/scripts/vllm_dissag/models.yaml b/scripts/vllm_dissag/models.yaml index e0017704..064ced6d 100644 --- a/scripts/vllm_dissag/models.yaml +++ b/scripts/vllm_dissag/models.yaml @@ -261,16 +261,8 @@ GLM-5.1-FP8: # decode step moves an 8192-token-wide buffer per layer x78 layers regardless of the # real batch. That is a fixed ~300ms/step floor (~320x this model's HBM-bandwidth # bound). Sizing it for the actual decode batch gives TPOT 302ms -> 88ms (3.4x), - # matching the published EP8 figure. Needs vLLM >= e8c186f71b. - # Prefill is unaffected (it genuinely dispatches wide batches). - # Per-role (moriio.sh -> VLLM_MORI_MAX_TOKENS_PER_RANK). MUST be >= max_num_seqs * - # num_experts_per_token (top-8 here): the value is also the RECV capacity - # (mori MaxNumTokensToRecv = worldSize * this). 512 tripped a device assert - # "Total recv token overflow". 2048 is safe and still 4x smaller than 8192. Decode wants - # the buffer sized to its real batch; prefill must KEEP the wide buffer (0 = inherit - # max_num_batched_tokens) because it genuinely dispatches 8192-token chunks. - # Recv capacity stays wide enough for the profiling dummy run (8192 tokens) and - # bursty top-8 routing: worldSize * ceil(65536/worldSize) >= 8192 for EP8..EP32. + # matching the published EP8 figure. The knob is decode.dp below -- prefill is + # unaffected (it genuinely dispatches wide 8192-token chunked-prefill batches). MORI_SHMEM_HEAP_SIZE: "17179869184" # DSA sparse-indexer logits-buffer cap (crash fix). The indexer prefill computes an # M*N fp32 logits buffer; split_indexer_prefill_chunks only sub-chunks the query dim diff --git a/scripts/vllm_dissag/run_xPyD_models.slurm b/scripts/vllm_dissag/run_xPyD_models.slurm index 52d046a1..365883b8 100755 --- a/scripts/vllm_dissag/run_xPyD_models.slurm +++ b/scripts/vllm_dissag/run_xPyD_models.slurm @@ -76,7 +76,7 @@ echo "Running from: $(pwd)" # model's own recipe — GLM-5.1 DSA needs block=1 + AITER MLA on), while a genuine # submit-time `-e VAR=...` still wins. Precedence: image-baked < models.yaml < submit -e. # Captured HERE (before the slurm sets any defaults) so it reflects user intent only. -_RECIPE_ENV_KEYS="VLLM_USE_V1 DECODE_MORI_MAX_TOTAL_RECV_TOKENS PREFILL_MORI_MAX_TOTAL_RECV_TOKENS DECODE_MORI_MAX_TOKENS_PER_RANK PREFILL_MORI_MAX_TOKENS_PER_RANK VLLM_MORI_MAX_TOKENS_PER_RANK VLLM_MORI_WARP_NUM_PER_BLOCK VLLM_MORI_BLOCK_NUM VLLM_MORI_RDMA_BLOCK_NUM VLLM_USE_LAYERNAME VLLM_ROCM_USE_AITER VLLM_ROCM_USE_AITER_RMSNORM VLLM_ROCM_USE_AITER_MLA KV_BLOCK_SIZE KV_CACHE_DTYPE KV_CACHE_MEMORY_BYTES GPU_MEMORY_UTILIZATION VLLM_CUDAGRAPH_MODE PREFILL_CUDAGRAPH_MODE DECODE_CUDAGRAPH_MODE CUDAGRAPH_CAPTURE_SIZES VLLM_ALL2ALL_BACKEND PREFILL_MORI_BACKEND DECODE_MORI_BACKEND MORI_SHMEM_HEAP_SIZE" +_RECIPE_ENV_KEYS="VLLM_USE_V1 VLLM_USE_LAYERNAME VLLM_ROCM_USE_AITER VLLM_ROCM_USE_AITER_RMSNORM VLLM_ROCM_USE_AITER_MLA KV_BLOCK_SIZE KV_CACHE_DTYPE KV_CACHE_MEMORY_BYTES GPU_MEMORY_UTILIZATION VLLM_CUDAGRAPH_MODE PREFILL_CUDAGRAPH_MODE DECODE_CUDAGRAPH_MODE CUDAGRAPH_CAPTURE_SIZES VLLM_ALL2ALL_BACKEND PREFILL_MORI_BACKEND DECODE_MORI_BACKEND MORI_SHMEM_HEAP_SIZE" MODELS_YAML_PROTECT="" for _k in $_RECIPE_ENV_KEYS; do [ -n "${!_k+x}" ] && MODELS_YAML_PROTECT="${MODELS_YAML_PROTECT} ${_k}" From 69eb062f4a8f20923ff61cf38be3c951947ce856 Mon Sep 17 00:00:00 2001 From: raviguptaamd Date: Sun, 16 Aug 2026 16:21:02 -0700 Subject: [PATCH 09/41] vllm_disagg: remove stray NFS silly-rename artifact scripts/vllm_dissag/.nfs0000000016f44d2b00008188 was committed by accident. It is a stale byte-copy of run_xPyD_models.slurm (same #!/bin/bash, #SBATCH --job-name=vllm-pd, #SBATCH -N 2 header) left behind by NFS when the real file was rewritten while an open handle still referenced it. Nothing reads it and it will drift from the real launcher on every edit. --- .../vllm_dissag/.nfs0000000016f44d2b00008188 | 685 ------------------ 1 file changed, 685 deletions(-) delete mode 100755 scripts/vllm_dissag/.nfs0000000016f44d2b00008188 diff --git a/scripts/vllm_dissag/.nfs0000000016f44d2b00008188 b/scripts/vllm_dissag/.nfs0000000016f44d2b00008188 deleted file mode 100755 index 52d046a1..00000000 --- a/scripts/vllm_dissag/.nfs0000000016f44d2b00008188 +++ /dev/null @@ -1,685 +0,0 @@ -#!/bin/bash -#SBATCH --job-name=vllm-pd # Specify a custom string for your slurm batch job -#SBATCH -N 2 # Default 2 nodes (1P/1D); override with sbatch -N for larger topologies -#SBATCH --ntasks-per-node=1 -#SBATCH --spread-job -#SBATCH --gres=gpu:8 # Request 8 GPUs and 8 NICs (use --gres if specific GPU resources are needed) -#SBATCH --time=24:00:00 # Set a time limit for the job (HH:MM:SS) -#SBATCH --output="/shared_inference/%u/model_blog_logs/slurm-%j.out" -#SBATCH --error="/shared_inference/%u/model_blog_logs/slurm-%j.err" - - -# ------------------------ -# Auto-detect and cd to the script directory so that $(pwd) always -# points to the folder containing the server scripts, regardless of -# where the user called sbatch from. -# -# Priority: -# 1. BASH_SOURCE — works for direct invocation (bash script.sh) -# 2. SLURM_SUBMIT_DIR + script path — when sbatch submits from repo root, -# SLURM_SUBMIT_DIR is the CWD, not the script dir. Append the relative -# path from the SBATCH command to get the actual script directory. -# 3. SLURM_SUBMIT_DIR alone — last resort (assumes sbatch was run from -# the script directory). -# ------------------------ -_resolve_script_dir() { - # Try BASH_SOURCE first (works for direct invocation) - if [[ -n "${BASH_SOURCE[0]:-}" ]]; then - local _d - _d="$(cd "$(dirname "${BASH_SOURCE[0]}")" 2>/dev/null && pwd)" - if [[ -n "$_d" && -f "$_d/vllm_disagg.sh" ]]; then - echo "$_d" - return 0 - fi - fi - - # Try SLURM_SUBMIT_DIR + relative script path - # When madengine does: sbatch scripts/vllm_dissag/run_xPyD_models.slurm - # SLURM_SUBMIT_DIR = repo root, so we need to append the dirname - if [[ -n "${SLURM_SUBMIT_DIR:-}" ]]; then - local _candidate="$SLURM_SUBMIT_DIR/scripts/vllm_dissag" - if [[ -f "$_candidate/vllm_disagg.sh" ]]; then - echo "$_candidate" - return 0 - fi - # Maybe they ran sbatch from the script dir itself - if [[ -f "$SLURM_SUBMIT_DIR/vllm_disagg.sh" ]]; then - echo "$SLURM_SUBMIT_DIR" - return 0 - fi - fi - - # Fallback - echo "." - return 1 -} - -SCRIPT_DIR="$(_resolve_script_dir)" -cd "$SCRIPT_DIR" || { echo "Error: cannot cd to $SCRIPT_DIR" >&2; exit 1; } - -REQUIRED_FILES=("vllm_disagg.sh" "parallelism.sh" "connectors/rixl.sh" "connectors/moriio.sh" "models.yaml" "benchmark_xPyD.sh" "parse_to_csv.py" "socket_barrier.py" "socket_wait.py" "connectors/moriio.env" "connectors/rixl.env") -for f in "${REQUIRED_FILES[@]}"; do - if [[ ! -f "$f" ]]; then - echo "Error: Required file '$f' not found in $(pwd)." >&2 - echo "Please run sbatch from the scripts/vllm_dissag/ directory, e.g.:" >&2 - echo " cd MAD/scripts/vllm_dissag && sbatch run_xPyD_models.slurm" >&2 - exit 1 - fi -done -echo "Running from: $(pwd)" - -# ------------------------------------------------------------------------------ -# models.yaml env precedence: capture which recipe knobs the USER explicitly set -# at submit time. The driver (vllm_disagg.sh) uses this to let models.yaml `env:` -# OVERRIDE image-baked ENV defaults (e.g. a DeepSeek-tuned image bakes -# KV_BLOCK_SIZE=16 / VLLM_ROCM_USE_AITER_MLA=0, which would otherwise shadow a -# model's own recipe — GLM-5.1 DSA needs block=1 + AITER MLA on), while a genuine -# submit-time `-e VAR=...` still wins. Precedence: image-baked < models.yaml < submit -e. -# Captured HERE (before the slurm sets any defaults) so it reflects user intent only. -_RECIPE_ENV_KEYS="VLLM_USE_V1 DECODE_MORI_MAX_TOTAL_RECV_TOKENS PREFILL_MORI_MAX_TOTAL_RECV_TOKENS DECODE_MORI_MAX_TOKENS_PER_RANK PREFILL_MORI_MAX_TOKENS_PER_RANK VLLM_MORI_MAX_TOKENS_PER_RANK VLLM_MORI_WARP_NUM_PER_BLOCK VLLM_MORI_BLOCK_NUM VLLM_MORI_RDMA_BLOCK_NUM VLLM_USE_LAYERNAME VLLM_ROCM_USE_AITER VLLM_ROCM_USE_AITER_RMSNORM VLLM_ROCM_USE_AITER_MLA KV_BLOCK_SIZE KV_CACHE_DTYPE KV_CACHE_MEMORY_BYTES GPU_MEMORY_UTILIZATION VLLM_CUDAGRAPH_MODE PREFILL_CUDAGRAPH_MODE DECODE_CUDAGRAPH_MODE CUDAGRAPH_CAPTURE_SIZES VLLM_ALL2ALL_BACKEND PREFILL_MORI_BACKEND DECODE_MORI_BACKEND MORI_SHMEM_HEAP_SIZE" -MODELS_YAML_PROTECT="" -for _k in $_RECIPE_ENV_KEYS; do - [ -n "${!_k+x}" ] && MODELS_YAML_PROTECT="${MODELS_YAML_PROTECT} ${_k}" -done -export MODELS_YAML_PROTECT="${MODELS_YAML_PROTECT# }" -echo "models.yaml protect-list (submit-time overrides): '${MODELS_YAML_PROTECT}'" - -# ------------------------ -# Print current time in UTC and PST formats -# ------------------------ -echo "=== Job Start Time ===" -echo "UTC Time: $(TZ=UTC date '+%Y-%m-%d %H:%M:%S %Z')" -echo "PST Time: $(TZ=America/Los_Angeles date '+%Y-%m-%d %H:%M:%S %Z')" -echo "=======================" -echo "" - -# Define valid model names (must have a models.yaml entry) -VALID_MODELS=( \ - "Llama-3.1-405B-Instruct-FP8-KV" \ - "amd-Llama-3.3-70B-Instruct-FP8-KV" \ - "DeepSeek-V3" \ - "DeepSeek-V3-5layer" \ - "gpt-oss-120b" \ - "DeepSeek-R1" \ - "Qwen3-32B" \ - "Qwen3-30B-A3B" \ - "GLM-5.1-FP8" \ -) - -# Models allowed for CONNECTOR=moriio WIDE_EP=1 (MoRI-EP; legacy RUN_MORI=1) -MORI_EP_VALID_MODELS=( \ - "DeepSeek-V3" \ - "DeepSeek-V3-5layer" \ - "DeepSeek-R1" \ - "GLM-5.1-FP8" \ -) - -# Models allowed for CONNECTOR=rixl WIDE_EP=1 EP_BACKEND=deepep (legacy RUN_DEEPEP=1) -DEEPEP_VALID_MODELS=( \ - "DeepSeek-V3" \ - "DeepSeek-V3-5layer" \ - "DeepSeek-R1" \ -) - -MODEL_NAME="${MODEL_NAME:-None}" - -validate_model_name() { - local is_valid_model=false - - for model in "${VALID_MODELS[@]}"; do - if [[ "$MODEL_NAME" == "$model" ]]; then - is_valid_model=true - break - fi - done - - if ! $is_valid_model; then - printf "Error: Invalid MODEL_NAME: '%s'\nValid models are:\n" "$MODEL_NAME" - for model in "${VALID_MODELS[@]}"; do - printf " - %s\n" "$model" - done - exit 1 - fi - - echo "MODEL_NAME '$MODEL_NAME' is valid." - return 0 -} - -validate_model_name "${MODEL_NAME}" - -model_allows_mori_ep() { - local m="$1" - for x in "${MORI_EP_VALID_MODELS[@]}"; do - [[ "$m" == "$x" ]] && return 0 - done - return 1 -} - -model_allows_deepep() { - local m="$1" - for x in "${DEEPEP_VALID_MODELS[@]}"; do - [[ "$m" == "$x" ]] && return 0 - done - return 1 -} - -# --------------------------------------------------------------------------- -# Axis selection -> single launcher (vllm_disagg.sh). -# Two axes: CONNECTOR={rixl|moriio} x WIDE_EP={0=TP|1=wideEP}; EP_BACKEND only -# when WIDE_EP=1 (rixl->deepep, moriio->mori). Legacy RUN_MORI / RUN_DEEPEP are -# still honored via a back-compat shim. The launcher itself does the final -# CONNECTOR/WIDE_EP/EP_BACKEND validation; here we resolve them + gate the model -# against the existing allowlists. -# --------------------------------------------------------------------------- -RUN_FILE="vllm_disagg.sh" -_run_mori="${RUN_MORI:-0}" -_run_deepep="${RUN_DEEPEP:-0}" - -if [[ "$_run_mori" == "1" && "$_run_deepep" == "1" ]]; then - echo "Error: Both RUN_MORI and RUN_DEEPEP are set to 1. Set only one." >&2 - exit 1 -fi - -# Back-compat: legacy flags map onto the axes when CONNECTOR is not set explicitly. -if [[ -z "${CONNECTOR:-}" ]]; then - if [[ "$_run_mori" == "1" ]]; then - CONNECTOR=moriio; WIDE_EP="${WIDE_EP:-1}"; EP_BACKEND="${EP_BACKEND:-mori}" - elif [[ "$_run_deepep" == "1" ]]; then - CONNECTOR=rixl; WIDE_EP="${WIDE_EP:-1}"; EP_BACKEND="${EP_BACKEND:-deepep}" - else - # Default keeps the historical "no flags" behavior: rixl + TP. - CONNECTOR=rixl; WIDE_EP="${WIDE_EP:-0}" - fi -fi -WIDE_EP="${WIDE_EP:-0}" - -# Models that ONLY run wideEP (DP/EP), never TP: the DeepSeek family is served with -# the MoRI-EP / DeepEP recipe (block=16, MLA off, per-role cudagraph). Running them -# in TP mode is unsupported — the TP argv would double the model's own -# --compilation-config and drop the mandatory +quant_fp8 op. Reject early. -# GLM-5.1-FP8 (GlmMoeDsaForCausalLM, MLA+DSA) is validated only under MoRI-EP -# wideEP disagg (block=1, AITER sparse MLA on, per-role all2all). The moriio+TP -# ("Stage B") path is untested for DSA, so reject WIDE_EP=0 for it too. -WIDE_EP_ONLY_MODELS=( "DeepSeek-V3" "DeepSeek-V3-5layer" "DeepSeek-R1" "GLM-5.1-FP8" ) -model_is_wide_ep_only() { - local m="$1" - for x in "${WIDE_EP_ONLY_MODELS[@]}"; do [[ "$m" == "$x" ]] && return 0; done - return 1 -} - -# Model allowlist gate (kept, sglang-style). wideEP modes use the per-backend -# allowlists; TP uses the global VALID_MODELS (already validated above) minus the -# wideEP-only models. -if [[ "$WIDE_EP" == "0" ]]; then - if model_is_wide_ep_only "$MODEL_NAME"; then - echo "Error: MODEL_NAME '$MODEL_NAME' is wideEP-only (set WIDE_EP=1). TP mode is not supported for it." >&2 - printf "wideEP-only models:\n"; for m in "${WIDE_EP_ONLY_MODELS[@]}"; do printf " - %s\n" "$m"; done - exit 1 - fi -elif [[ "$WIDE_EP" == "1" && "${CONNECTOR}" == "moriio" ]]; then - if ! model_allows_mori_ep "$MODEL_NAME"; then - echo "Error: CONNECTOR=moriio WIDE_EP=1 but MODEL_NAME '$MODEL_NAME' is not in MORI_EP_VALID_MODELS" >&2 - printf "MoRI EP allowed models:\n"; for m in "${MORI_EP_VALID_MODELS[@]}"; do printf " - %s\n" "$m"; done - exit 1 - fi -elif [[ "$WIDE_EP" == "1" && "${CONNECTOR}" == "rixl" ]]; then - if ! model_allows_deepep "$MODEL_NAME"; then - echo "Error: CONNECTOR=rixl WIDE_EP=1 (deepep) but MODEL_NAME '$MODEL_NAME' is not in DEEPEP_VALID_MODELS" >&2 - printf "DeepEP allowed models:\n"; for m in "${DEEPEP_VALID_MODELS[@]}"; do printf " - %s\n" "$m"; done - exit 1 - fi -fi - -export CONNECTOR WIDE_EP EP_BACKEND -echo "Launcher: $RUN_FILE (CONNECTOR=${CONNECTOR} WIDE_EP=${WIDE_EP} EP_BACKEND=${EP_BACKEND:-}) for model '$MODEL_NAME'" - -# --------------------------------------------------------------------------- -# Connector platform env: per-connector .env holds the ROCm-7.2.3 -# runtime env that MUST reach the container at PID 1 (e.g. expandable_segments:False -# for GPU-RDMA registration). Source the resolved connector's file and collect its -# KEY=VALUE lines into CONNECTOR_ENV_ARGS as `-e KEY=${KEY:-VALUE}` pairs, so a -# submit-time export of the same name still overrides. Forwarded in the docker run. -# --------------------------------------------------------------------------- -CONNECTOR_ENV_FILE="${SCRIPT_DIR}/connectors/${CONNECTOR}.env" -CONNECTOR_ENV_ARGS="" -if [[ -f "$CONNECTOR_ENV_FILE" ]]; then - echo "Loading connector platform env: $CONNECTOR_ENV_FILE" - while IFS= read -r _line; do - [[ "$_line" =~ ^[[:space:]]*# || -z "${_line// }" ]] && continue - _k="${_line%%=*}"; _v="${_line#*=}" - CONNECTOR_ENV_ARGS+=" -e ${_k}=${!_k:-$_v}" # submit-time export of $_k wins - done < "$CONNECTOR_ENV_FILE" -else - echo "WARN: connector env file not found: $CONNECTOR_ENV_FILE" >&2 -fi - -if [[ -z "${DOCKER_IMAGE_NAME:-}" ]]; then - echo "Error: DOCKER_IMAGE_NAME is not set. Please export DOCKER_IMAGE_NAME before running." >&2 - echo " There is no public prebuilt image. Build your own from the provided Dockerfile:" >&2 - echo " docker build -f docker/vllm_disagg_inference.ubuntu.amd.Dockerfile \\" >&2 - echo " -t /vllm-disagg:local . # all connectors (WITH_NIXL=1 default)" >&2 - echo " (add --build-arg WITH_NIXL=0 for a lean MoRI-EP-only image)" >&2 - echo " then: export DOCKER_IMAGE_NAME=/vllm-disagg:local" >&2 - exit 1 -fi -export DOCKER_IMAGE_NAME - -# Set current directory to be REPO directory with all relevant scripts -NIXL_REPO_DIR=$(pwd) -LOG_PATH="${LOG_PATH:-/shared_inference/${USER}/model_blog_logs}" - -xP="${xP:-1}" #-> Number of Prefill Servers -yD="${yD:-1}" #-> Number of Decode Servers - -MODEL_DIR="${MODEL_DIR:-"/shared_inference/models_blog/"}" - - -# ------------------------ -# Model path validation and selection across all nodes -# ------------------------ -echo "Looking for model: $MODEL_NAME" -echo "Checking model availability across all allocated nodes..." - -# Get all allocated nodes -ALL_NODES=$(scontrol show hostnames "$SLURM_JOB_NODELIST") -TOTAL_NODES=$(echo "$ALL_NODES" | wc -l) - -echo "Total allocated nodes: $TOTAL_NODES" -echo "Nodes: $(echo "$ALL_NODES" | tr '\n' ' ')" - -# Function to check model path on all nodes -check_model_path() { - local path=$1 - local check_name=$2 - - echo "Checking $check_name: $path" - - # Run check on all nodes in parallel - srun --nodes=$SLURM_NNODES --ntasks=$SLURM_NNODES /bin/bash -c " - if [ -d '$path' ]; then - echo \"\$(hostname): ✓ Found $path\" - exit 0 - else - echo \"\$(hostname): ✗ Missing $path\" - exit 1 - fi - " - - # Check if all nodes succeeded (exit code 0) - local exit_code=$? - if [ $exit_code -eq 0 ]; then - echo "✓ $check_name available on ALL nodes" - return 0 - else - echo "✗ $check_name NOT available on all nodes" - return 1 - fi -} - -# Check /mnt/m2m_nobackup/models_blog first -MODEL_PATH_1="/mnt/m2m_nobackup/models_blog/$MODEL_NAME" -if check_model_path "$MODEL_PATH_1" "/mnt/m2m_nobackup/models_blog"; then - MODEL_PATH="$MODEL_PATH_1" - echo "" - echo "✓ Selected MODEL_PATH: $MODEL_PATH (available on all nodes)" -# Check /shared-inference/models_blog -elif check_model_path "/shared_inference/models_blog/$MODEL_NAME" "/shared_inference/models_blog"; then - MODEL_PATH="/shared_inference/models_blog/$MODEL_NAME" - echo "" - echo "✓ Selected MODEL_PATH: $MODEL_PATH (available on all nodes)" -elif check_model_path "$MODEL_DIR/$MODEL_NAME" "$MODEL_DIR"; then - MODEL_PATH="$MODEL_DIR/$MODEL_NAME" - echo "" - echo "✓ Selected MODEL_PATH: $MODEL_PATH (available on all nodes)" -else - echo "" - echo "✗ FATAL ERROR: Model '$MODEL_NAME' not found on ALL allocated nodes in either:" - echo " - /mnt/m2m_nobackup/models_blog/$MODEL_NAME" - echo " - /shared_inference/models_blog/$MODEL_NAME" - echo "" - echo "Model must be accessible from all nodes for distributed execution." - echo "Please ensure the model is available on all allocated nodes." - exit 1 -fi - -echo "Final MODEL_PATH: $MODEL_PATH" -echo "" - - -# Calculate NUM_NODES based on xP and yD -NUM_NODES=$((xP + yD)) -echo "Calculated NUM_NODES: $NUM_NODES (xP=$xP + yD=$yD, proxy co-located on prefill master)" - -# DeepEP configuration (only exported when RUN_DEEPEP=1) -if [[ "$_run_deepep" == "1" ]]; then - export PREFILL_DEEPEP_BACKEND="${PREFILL_DEEPEP_BACKEND:-deepep_high_throughput}" - export DECODE_DEEPEP_BACKEND="${DECODE_DEEPEP_BACKEND:-deepep_low_latency}" - export ENABLE_DBO="${ENABLE_DBO:-false}" - export DBO_COMM_SMS="${DBO_COMM_SMS:-}" - export ENABLE_PROFILING="${ENABLE_PROFILING:-false}" - echo "DeepEP config: PREFILL_BACKEND=$PREFILL_DEEPEP_BACKEND DECODE_BACKEND=$DECODE_DEEPEP_BACKEND DBO=$ENABLE_DBO" -fi - -# ------------------------ -# Extract first NUM_NODES from SLURM allocation and update SLURM variables -# ------------------------ -echo "Original SLURM allocation:" -echo "SLURM_JOB_NODELIST: $SLURM_JOB_NODELIST" -echo "SLURM_NNODES: $SLURM_NNODES" -echo "SLURM_NTASKS: $SLURM_NTASKS" - -# Get the full nodelist and extract first NUM_NODES -FULL_NODELIST=$(scontrol show hostnames "$SLURM_JOB_NODELIST") -SELECTED_NODES=$(echo "$FULL_NODELIST" | head -n $NUM_NODES) -NEW_SLURM_NODELIST=$(echo "$SELECTED_NODES" | paste -sd,) - -# Update SLURM environment variables -export SLURM_NNODES=$NUM_NODES -export SLURM_NTASKS=$NUM_NODES -export SLURM_JOB_NUM_NODES=$NUM_NODES -export SLURM_NPROCS=$NUM_NODES -export SLURM_JOB_NODELIST="$NEW_SLURM_NODELIST" -export SLURM_NODELIST="$NEW_SLURM_NODELIST" - -# Keep other SLURM variables as they were or set defaults -export SLURM_TASKS_PER_NODE="1(x$NUM_NODES)" - -export SLURM_CLUSTER_NAME="${SLURM_CLUSTER_NAME}" -export SLURM_JOB_CPUS_PER_NODE="${SLURM_JOB_CPUS_PER_NODE}" -export SLURM_JOB_PARTITION="${SLURM_JOB_PARTITION}" -export SLURM_JOBID="${SLURM_JOBID:-$SLURM_JOB_ID}" -export SLURM_JOB_QOS="${SLURM_JOB_QOS:-normal}" -export SLURM_JOB_ACCOUNT="${SLURM_JOB_ACCOUNT}" -export SLURM_NTASKS_PER_NODE=1 -export SLURM_SUBMIT_HOST="${SLURM_SUBMIT_HOST}" -export SLURM_JOB_ID="${SLURM_JOB_ID}" -export SLURM_CONF="${SLURM_CONF:-/etc/slurm/slurm.conf}" -export SLURM_JOB_NAME="${SLURM_JOB_NAME:-1p1d_bench-serving}" - -echo "" -echo "Updated SLURM Environment Variables:" -echo "SLURM_JOB_ID: $SLURM_JOB_ID" -echo "SLURM_JOB_NODELIST: $SLURM_JOB_NODELIST" -echo "SLURM_NNODES: $SLURM_NNODES" -echo "SLURM_NTASKS: $SLURM_NTASKS" -echo "SLURM_TASKS_PER_NODE: $SLURM_TASKS_PER_NODE" -echo "SLURM_JOB_CPUS_PER_NODE: $SLURM_JOB_CPUS_PER_NODE" -echo "SLURM_JOB_PARTITION: $SLURM_JOB_PARTITION" -echo "SLURM_JOB_NUM_NODES: $SLURM_JOB_NUM_NODES" -echo "SLURM_JOBID: $SLURM_JOBID" -echo "SLURM_JOB_QOS: $SLURM_JOB_QOS" -echo "SLURM_NODELIST: $SLURM_NODELIST" -echo "SLURM_JOB_ACCOUNT: $SLURM_JOB_ACCOUNT" -echo "SLURM_NPROCS: $SLURM_NPROCS" -echo "SLURM_SUBMIT_HOST: $SLURM_SUBMIT_HOST" -echo "SLURM_CONF: $SLURM_CONF" -echo "SLURM_JOB_NAME: $SLURM_JOB_NAME" -echo "SLURM_NTASKS_PER_NODE: $SLURM_NTASKS_PER_NODE" -#echo "SLURM_SUBMIT_DIR: $SLURM_SUBMIT_DIR" -echo "SLURM_CLUSTER_NAME: $SLURM_CLUSTER_NAME" -echo "ulimit: $(ulimit -a)" -echo "" -echo "Selected nodes for execution:" -echo "$SELECTED_NODES" -echo "" - -# Node information -USER_NAME=$(whoami) -MASTER_NODE=$(echo "$SELECTED_NODES" | head -n 1) -# Pick the routable fabric IP, not just hostname -I's first entry. These nodes expose -# multiple NICs (e.g. a 10.224.x overlay listed BEFORE the routable 10.158.x fabric); -# taking $1 blindly can advertise an unreachable addr -> prefill/decode barrier hangs -# "Waiting for nodes" forever. Prefer FABRIC_SUBNET (default 10.158.), fall back to $1. -FABRIC_SUBNET="${FABRIC_SUBNET:-10.158.}" -# From a "hostname -I" line, return the first IP on FABRIC_SUBNET, else the first IP. -_pick_fabric_ip() { - awk -v pfx="$FABRIC_SUBNET" '{f=$1; for(i=1;i<=NF;i++) if(index($i,pfx)==1){f=$i; break} print f}' -} -MASTER_ADDR=$(srun --nodes=1 --ntasks=1 --time=00:20:00 --nodelist="$MASTER_NODE" bash -c 'hostname -I' | _pick_fabric_ip) -MASTER_PORT=39566 # Choose an open port - -IPS=() - -for NODE in $SELECTED_NODES; do - IP=$(srun --nodes=1 --ntasks=1 --time=00:20:00 --nodelist="$NODE" bash -c 'hostname -I' | _pick_fabric_ip) - IPS+=("$IP") -done - -echo "Selected node IPs: ${IPS[*]}" | sed 's/ /,/g' - -NIXL_COOKBOOK_PATH="/opt/nixl-vllm-cookbook" -BENCHMARK_ITR="${BENCHMARK_ITR:-1}" -BENCHMARK_CON="${BENCHMARK_CON:-}" -BENCHMARK_COMBINATIONS="${BENCHMARK_COMBINATIONS:-}" - -# Benchmark script selector: BENCHMARK_SCRIPT tag -> file run by the launcher. -# sweep (default) -> benchmark_xPyD.sh (general concurrency sweep) -# long_context -> benchmark_long_context.sh (per-shape warmup, c=1-first) -# keepalive -> keepalive_bench.sh (hold server up KEEPALIVE_MINS -# for external accuracy probes) -BENCHMARK_SCRIPT="${BENCHMARK_SCRIPT:-sweep}" -case "$BENCHMARK_SCRIPT" in - sweep) BENCHMARK_SCRIPT_FILE="benchmark_xPyD.sh" ;; - long_context) BENCHMARK_SCRIPT_FILE="benchmark_long_context.sh" ;; - keepalive) BENCHMARK_SCRIPT_FILE="keepalive_bench.sh" ;; - *) echo "Error: invalid BENCHMARK_SCRIPT='$BENCHMARK_SCRIPT' (valid: sweep, long_context, keepalive)" >&2; exit 1 ;; -esac -if [[ ! -f "$BENCHMARK_SCRIPT_FILE" ]]; then - echo "Error: selected benchmark script '$BENCHMARK_SCRIPT_FILE' not found in $(pwd)." >&2 - exit 1 -fi -echo "BENCHMARK_SCRIPT='$BENCHMARK_SCRIPT' -> $BENCHMARK_SCRIPT_FILE" - -NNODES=$NUM_NODES - -echo "MASTER_NODE is ${MASTER_NODE}" -echo "MASTER_ADDR is ${MASTER_ADDR}" -echo "MASTER_PORT is ${MASTER_PORT}" -echo "NNODES is ${NNODES}" -echo "REPO Directory is ${NIXL_REPO_DIR}" - -if [ ! -d "$LOG_PATH" ]; then - mkdir -p "$LOG_PATH" - echo "Created directory: $LOG_PATH" -else - echo "Directory already exists: $LOG_PATH" -fi - -export CONNECTOR_ENV_ARGS="$CONNECTOR_ENV_ARGS" -export LOG_PATH=$LOG_PATH -export NIXL_REPO_DIR=$NIXL_REPO_DIR -export NIXL_COOKBOOK_PATH=$NIXL_COOKBOOK_PATH -export NNODES=$NNODES -export MASTER_ADDR=$MASTER_ADDR -export MASTER_PORT=$MASTER_PORT -export MODEL_PATH=$MODEL_PATH -export xP=$xP -export yD=$yD -export MODEL_NAME=$MODEL_NAME -export USER_NAME=$USER_NAME -export IPADDRS="$(echo "${IPS[*]}" | sed 's/ /,/g')" -export BENCHMARK_ITR=$BENCHMARK_ITR -export BENCHMARK_CON="${BENCHMARK_CON}" -export BENCHMARK_COMBINATIONS="${BENCHMARK_COMBINATIONS}" -export BENCHMARK_SCRIPT_FILE="${BENCHMARK_SCRIPT_FILE}" - -export DOCKER_CONT_NAME="container_${MODEL_NAME}_${SLURM_JOB_ID}" -export RUN_FILE_FULL="$NIXL_COOKBOOK_PATH/${RUN_FILE}" - -# Use only the selected nodes for srun execution -SELECTED_NODELIST_SRUN=$(echo "$SELECTED_NODES" | paste -sd,) - -srun --nodelist="$SELECTED_NODELIST_SRUN" bash -c ' -echo "Rank $SLURM_PROCID on $(hostname)"; -docker ps -q | xargs --no-run-if-empty docker stop; -docker rm -f $DOCKER_CONT_NAME 2>/dev/null || true; -fuser -k 5000/tcp 2>/dev/null || true; -fuser -k 2222/tcp 2>/dev/null || true; -fuser -k 15000/tcp 2>/dev/null || true; -sleep 2; -docker pull $DOCKER_IMAGE_NAME 2>/dev/null || true; - -# --- Create host-local compilation cache dirs (ext4, survives container restarts) --- -mkdir -p /tmp/vllm_cache/{aiter_jit,triton,vllm,comgr} 2>/dev/null || true; - -# --- Persistent JIT cache mount --- -# The image points AITER_JIT_DIR/TRITON_CACHE_DIR/VLLM_CACHE_ROOT/COMGR_CACHE_DIR at -# /opt/vllm_cache. Mount a host dir there so AITER CK kernels compile ONCE and are reused -# across runs. Cold compile is ~15 min for the DeepSeek MoE/FP8 GEMM kernel set; a warm -# boot is ~1 min. Keyed by image ID so a new image starts a fresh cache and does not reuse -# stale-ABI shared objects. Host dir on local NVMe. Override JIT_CACHE_HOST, or set -# JIT_CACHE_PERSIST=0 to disable and fall back to the image empty in-container cache. -# NOTE: this whole section runs inside a single-quoted `srun bash -c '...'`, so avoid -# single quotes here; the image-id hash is extracted with tr, not sed. -if [ "${JIT_CACHE_PERSIST:-1}" = "1" ]; then - _IMG_RAW=$(docker image inspect --format "{{.Id}}" "$DOCKER_IMAGE_NAME" 2>/dev/null); - _IMG_KEY=$(printf "%s" "$_IMG_RAW" | tr -cd "a-f0-9" | cut -c1-12); - _IMG_KEY="${_IMG_KEY:-noimg}"; - _JIT_CACHE_HOST="${JIT_CACHE_HOST:-/mnt/m2m_nobackup/${USER}/vllm_jit_cache/${_IMG_KEY}}"; - mkdir -p "$_JIT_CACHE_HOST"/{aiter_jit,triton,vllm,comgr} 2>/dev/null || true; - _JIT_CACHE_MOUNT="-v ${_JIT_CACHE_HOST}:/opt/vllm_cache"; - echo "[jit-cache] persistent image ${_IMG_KEY}: ${_JIT_CACHE_HOST} to /opt/vllm_cache"; -else - _JIT_CACHE_MOUNT=""; -fi - -# --- Build host RDMA library mounts --- -_RDMA_MOUNTS="" -_LIBDIR=/usr/lib/x86_64-linux-gnu - -for _lib in libibverbs.so libibverbs.so.1 librdmacm.so librdmacm.so.1; do - [ -e "$_LIBDIR/$_lib" ] && _RDMA_MOUNTS="$_RDMA_MOUNTS -v $_LIBDIR/$_lib:$_LIBDIR/$_lib:ro" -done -for _vlib in $_LIBDIR/libibverbs.so.1.* $_LIBDIR/librdmacm.so.1.*; do - [ -e "$_vlib" ] && _RDMA_MOUNTS="$_RDMA_MOUNTS -v $_vlib:$_vlib:ro" -done - -for _pattern in libmlx5.so* libionic*.so* libbnxt_re*.so* libefa.so* libhns.so*; do - for _vlib in $_LIBDIR/${_pattern}; do - # Require a regular file AFTER symlink resolution: these mounts are built on - # ONE node but applied on ALL nodes, and vendor NIC libs (e.g. libionic.so.1) - # can be a DANGLING symlink on some nodes -> bind-mount fails "not a directory" - # -> container create exit 125. `-f` (follows symlink, requires regular file) - # skips those; the fabric in use (mlx5) is still mounted where present. - [ -f "$_vlib" ] && _RDMA_MOUNTS="$_RDMA_MOUNTS -v $_vlib:$_vlib:ro" - done -done - -[ -d "$_LIBDIR/libibverbs" ] && _RDMA_MOUNTS="$_RDMA_MOUNTS -v $_LIBDIR/libibverbs:$_LIBDIR/libibverbs:ro" -[ -d /etc/libibverbs.d ] && _RDMA_MOUNTS="$_RDMA_MOUNTS -v /etc/libibverbs.d:/etc/libibverbs.d:ro" -echo "[host-rdma] mounts: $_RDMA_MOUNTS" - -docker run --rm \ - --device /dev/dri \ - --device /dev/kfd \ - --device /dev/infiniband \ - --network host \ - --ipc host \ - --group-add video \ - --cap-add SYS_PTRACE \ - --security-opt seccomp=unconfined \ - --privileged \ - -v $HOME:$HOME \ - -v /shared_inference:/shared_inference \ - -v /mnt/m2m_nobackup:/mnt/m2m_nobackup \ - -v $HOME/.ssh:/root/.ssh \ - --shm-size ${DOCKER_SHM_SIZE:-256G} \ - --ulimit nofile=524288:524288 \ - --ulimit memlock=-1:-1 \ - -v ${LOG_PATH}:/run_logs \ - -v $NIXL_REPO_DIR:$NIXL_COOKBOOK_PATH \ - -v /tmp/vllm_cache:/tmp/vllm_cache \ - ${_JIT_CACHE_MOUNT} \ - ${GLM_KERNEL_PATCH:+-v ${GLM_KERNEL_PATCH}:/usr/local/lib/python3.12/dist-packages/vllm/v1/attention/ops/rocm_aiter_mla_sparse.py:ro} \ - ${GLM_BACKEND_PATCH:+-v ${GLM_BACKEND_PATCH}:/usr/local/lib/python3.12/dist-packages/vllm/v1/attention/backends/mla/rocm_aiter_mla_sparse.py:ro} \ - $_RDMA_MOUNTS \ - --entrypoint /bin/bash \ - -e SLURM_JOB_ID=$SLURM_JOB_ID \ - -e SLURM_JOB_NODELIST=$SLURM_JOB_NODELIST \ - -e NNODES=$NNODES \ - -e NODE_RANK=$SLURM_PROCID \ - -e MASTER_ADDR=$MASTER_ADDR \ - -e MASTER_PORT=$MASTER_PORT \ - -e MODEL_PATH=$MODEL_PATH \ - -e NIXL_COOKBOOK_PATH=$NIXL_COOKBOOK_PATH \ - -e xP=$xP \ - -e yD=$yD \ - -e USER_NAME=$USER_NAME \ - -e MODEL_NAME=$MODEL_NAME \ - -e BENCHMARK_ITR=$BENCHMARK_ITR \ - -e BENCHMARK_CON="${BENCHMARK_CON}" \ - -e BENCHMARK_COMBINATIONS="${BENCHMARK_COMBINATIONS}" \ - ${BENCHMARK_PORT:+-e BENCHMARK_PORT=$BENCHMARK_PORT} \ - ${PROXY_TYPE:+-e PROXY_TYPE=$PROXY_TYPE} \ - ${ROUTER_PORT:+-e ROUTER_PORT=$ROUTER_PORT} \ - -e IPADDRS=$IPADDRS \ - ${CONNECTOR:+-e CONNECTOR=$CONNECTOR} \ - ${WIDE_EP:+-e WIDE_EP=$WIDE_EP} \ - ${EP_BACKEND:+-e EP_BACKEND=$EP_BACKEND} \ - ${RUN_MORI:+-e RUN_MORI=$RUN_MORI} \ - ${RUN_DEEPEP:+-e RUN_DEEPEP=$RUN_DEEPEP} \ - ${VLLM_ALL2ALL_BACKEND:+-e VLLM_ALL2ALL_BACKEND=$VLLM_ALL2ALL_BACKEND} \ - ${PREFILL_MORI_BACKEND:+-e PREFILL_MORI_BACKEND=$PREFILL_MORI_BACKEND} \ - ${DECODE_MORI_BACKEND:+-e DECODE_MORI_BACKEND=$DECODE_MORI_BACKEND} \ - -e MODELS_YAML_PROTECT="${MODELS_YAML_PROTECT:-}" \ - ${GLM_PERSIST_GATE:+-e GLM_PERSIST_GATE=$GLM_PERSIST_GATE} \ - ${GLM_SKIP_PATCHERS:+-e GLM_SKIP_PATCHERS=$GLM_SKIP_PATCHERS} \ - ${KV_BLOCK_SIZE:+-e KV_BLOCK_SIZE=$KV_BLOCK_SIZE} \ - ${KV_CACHE_MEMORY_BYTES:+-e KV_CACHE_MEMORY_BYTES=$KV_CACHE_MEMORY_BYTES} \ - ${VLLM_ROCM_USE_AITER_MLA:+-e VLLM_ROCM_USE_AITER_MLA=$VLLM_ROCM_USE_AITER_MLA} \ - ${ROUTER_BINARY:+-e ROUTER_BINARY=$ROUTER_BINARY} \ - ${KV_CACHE_DTYPE:+-e KV_CACHE_DTYPE=$KV_CACHE_DTYPE} \ - ${MORIIO_TOY_PROXY:+-e MORIIO_TOY_PROXY=$MORIIO_TOY_PROXY} \ - ${BENCHMARK_SCRIPT_FILE:+-e BENCHMARK_SCRIPT_FILE=$BENCHMARK_SCRIPT_FILE} \ - ${KEEPALIVE_MINS:+-e KEEPALIVE_MINS=$KEEPALIVE_MINS} \ - ${PREFILL_CUDAGRAPH_MODE:+-e PREFILL_CUDAGRAPH_MODE=$PREFILL_CUDAGRAPH_MODE} \ - ${DECODE_CUDAGRAPH_MODE:+-e DECODE_CUDAGRAPH_MODE=$DECODE_CUDAGRAPH_MODE} \ - ${CUDAGRAPH_CAPTURE_SIZES:+-e CUDAGRAPH_CAPTURE_SIZES="$CUDAGRAPH_CAPTURE_SIZES"} \ - ${MORI_RDMA_TC:+-e MORI_RDMA_TC=$MORI_RDMA_TC} \ - ${MORI_RDMA_SL:+-e MORI_RDMA_SL=$MORI_RDMA_SL} \ - ${MORI_SHMEM_HEAP_SIZE:+-e MORI_SHMEM_HEAP_SIZE=$MORI_SHMEM_HEAP_SIZE} \ - ${PREFILL_DEEPEP_BACKEND:+-e PREFILL_DEEPEP_BACKEND=$PREFILL_DEEPEP_BACKEND} \ - ${DECODE_DEEPEP_BACKEND:+-e DECODE_DEEPEP_BACKEND=$DECODE_DEEPEP_BACKEND} \ - ${ENABLE_DBO:+-e ENABLE_DBO=$ENABLE_DBO} \ - ${DBO_COMM_SMS:+-e DBO_COMM_SMS=$DBO_COMM_SMS} \ - ${ENABLE_PROFILING:+-e ENABLE_PROFILING=$ENABLE_PROFILING} \ - ${NCCL_IB_HCA:+-e NCCL_IB_HCA=$NCCL_IB_HCA} \ - ${NCCL_IB_GID_INDEX:+-e NCCL_IB_GID_INDEX=$NCCL_IB_GID_INDEX} \ - ${NCCL_NET_GDR_LEVEL:+-e NCCL_NET_GDR_LEVEL=$NCCL_NET_GDR_LEVEL} \ - ${NCCL_CROSS_NIC:+-e NCCL_CROSS_NIC=$NCCL_CROSS_NIC} \ - ${NCCL_SOCKET_IFNAME:+-e NCCL_SOCKET_IFNAME=$NCCL_SOCKET_IFNAME} \ - ${GLOO_SOCKET_IFNAME:+-e GLOO_SOCKET_IFNAME=$GLOO_SOCKET_IFNAME} \ - -e MORI_SOCKET_IFNAME=${MORI_SOCKET_IFNAME:-eth0} \ - ${MORI_IB_GID_INDEX:+-e MORI_IB_GID_INDEX=$MORI_IB_GID_INDEX} \ - ${MORI_RDMA_DEVICES:+-e MORI_RDMA_DEVICES=$MORI_RDMA_DEVICES} \ - ${MORI_NUM_QP_PER_PE:+-e MORI_NUM_QP_PER_PE=$MORI_NUM_QP_PER_PE} \ - ${VLLM_MORIIO_QP_PER_TRANSFER:+-e VLLM_MORIIO_QP_PER_TRANSFER=$VLLM_MORIIO_QP_PER_TRANSFER} \ - ${VLLM_MORIIO_NUM_WORKERS:+-e VLLM_MORIIO_NUM_WORKERS=$VLLM_MORIIO_NUM_WORKERS} \ - -e GPU_MEMORY_UTILIZATION=${GPU_MEMORY_UTILIZATION:-0.8} \ - -e GPUS_PER_NODE=${GPUS_PER_NODE:-8} \ - ${GPU_MAX_HW_QUEUES:+-e GPU_MAX_HW_QUEUES=$GPU_MAX_HW_QUEUES} \ - ${HIP_FORCE_DEV_KERNARG:+-e HIP_FORCE_DEV_KERNARG=$HIP_FORCE_DEV_KERNARG} \ - ${HSA_NO_SCRATCH_RECLAIM:+-e HSA_NO_SCRATCH_RECLAIM=$HSA_NO_SCRATCH_RECLAIM} \ - ${VLLM_HANDSHAKE_TIMEOUT_MINS:+-e VLLM_HANDSHAKE_TIMEOUT_MINS=$VLLM_HANDSHAKE_TIMEOUT_MINS} \ - ${VLLM_ENGINE_READY_TIMEOUT_S:+-e VLLM_ENGINE_READY_TIMEOUT_S=$VLLM_ENGINE_READY_TIMEOUT_S} \ - ${ROCSHMEM_HEAP_SIZE:+-e ROCSHMEM_HEAP_SIZE=$ROCSHMEM_HEAP_SIZE} \ - ${ROCSHMEM_MAX_NUM_CONTEXTS:+-e ROCSHMEM_MAX_NUM_CONTEXTS=$ROCSHMEM_MAX_NUM_CONTEXTS} \ - ${LOG_WAIT_TIMEOUT_SECONDS:+-e LOG_WAIT_TIMEOUT_SECONDS=$LOG_WAIT_TIMEOUT_SECONDS} \ - ${TRITON_CACHE_DIR:+-e TRITON_CACHE_DIR=$TRITON_CACHE_DIR} \ - ${VLLM_CACHE_ROOT:+-e VLLM_CACHE_ROOT=$VLLM_CACHE_ROOT} \ - ${COMGR_CACHE_DIR:+-e COMGR_CACHE_DIR=$COMGR_CACHE_DIR} \ - ${AITER_JIT_DIR:+-e AITER_JIT_DIR=$AITER_JIT_DIR} \ - -e DISTRIBUTED_TIMEOUT_SECONDS=${DISTRIBUTED_TIMEOUT_SECONDS:-7200} \ - -e VLLM_RPC_TIMEOUT=${VLLM_RPC_TIMEOUT:-300000} \ - -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=${VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS:-3600} \ - ${CONNECTOR_ENV_ARGS} \ - ${VLLM_CUDAGRAPH_MODE:+-e VLLM_CUDAGRAPH_MODE=$VLLM_CUDAGRAPH_MODE} \ - ${CUDAGRAPH_CAPTURE_SIZES:+-e CUDAGRAPH_CAPTURE_SIZES="$CUDAGRAPH_CAPTURE_SIZES"} \ - --name $DOCKER_CONT_NAME \ - $DOCKER_IMAGE_NAME -c " - mkdir -p /run_logs/${SLURM_JOB_ID} - $RUN_FILE_FULL 2>&1 | tee /run_logs/${SLURM_JOB_ID}/pd_vllm_bench_NODE${SLURM_PROCID}.log - " -' -srun --nodelist="$SELECTED_NODELIST_SRUN" bash -c 'docker stop $DOCKER_CONT_NAME 2>/dev/null || true; docker rm $DOCKER_CONT_NAME 2>/dev/null || true' - From 0f660fe71ec1e30931012a9bb599263f6860e1f8 Mon Sep 17 00:00:00 2001 From: raviguptaamd Date: Sun, 16 Aug 2026 16:21:09 -0700 Subject: [PATCH 10/41] vllm_disagg: build the GLM image for gfx950 (MI355X) + ionic, and fix a silent wrong-arch bug Targets MI355X and the Pensando Ionic AI NIC: GFX_COMPILATION_ARCH gfx942 -> gfx950 NIC_COMPILATION_ARCH cx7 -> ionic MORI_GPU_ARCHS gfx942 -> gfx950 MORI_DEVICE_NIC (new) = ionic MORI_DEVICE_NIC is pinned rather than left to auto-detect. Detection does resolve correctly here (8x rocep*s0 -> driver readlink -> ionic, with libionic.so.1.1.54.0-184 present), but pinning keeps the device-side IBGDA JIT path deterministic across nodes. Renames two build ARGs, which is the load-bearing part of this commit: ARG PYTORCH_ROCM_ARCH -> ARG BUILD_ROCM_ARCH ARG MAX_JOBS -> ARG BUILD_MAX_JOBS (both assigned through to the original ENV names) The base image (rocm/vllm-dev:ci_base-*) exports PYTORCH_ROCM_ARCH and MAX_JOBS as ENV, and an inherited ENV overrides a same-named ARG in the child Dockerfile. So `ARG PYTORCH_ROCM_ARCH="gfx950"` was being silently ignored and the build used the base's `gfx90a;gfx942;gfx950` instead. This surfaced only by luck: the base also sets MAX_JOBS to the empty string, so setup.py::compute_num_jobs crashed on int(''). Had it been non-empty the build would have "succeeded" while compiling for the wrong arch list, and every downstream perf number would have been measured on an unintended binary with nothing in the logs to say so. Confirmed fixed by the rebuild log: "84 warnings generated when compiling for gfx950". Also replaces the post-install verification heredoc with an equivalent `python3 -c "$(printf ...)"` form, and reads the vLLM version via importlib.metadata instead of `import vllm` (importing pulls torch -> amdsmi -> libamd_smi.so, which is not loadable in the no-GPU build sandbox). --- ...gg_inference.glmv5.1.ubuntu.amd.Dockerfile | 53 ++++++++++--------- 1 file changed, 28 insertions(+), 25 deletions(-) diff --git a/docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile b/docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile index edd4d797..5955db3d 100644 --- a/docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile +++ b/docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile @@ -89,16 +89,16 @@ FROM ${BASE_IMAGE} ENTRYPOINT [] WORKDIR /app -ARG GFX_COMPILATION_ARCH="gfx942" -ARG PYTORCH_ROCM_ARCH="gfx942" -ARG MAX_JOBS=32 +ARG GFX_COMPILATION_ARCH="gfx950" +ARG BUILD_ROCM_ARCH="gfx950" +ARG BUILD_MAX_JOBS=32 # NIXL/RIXL transport for the rixl connector. GLM-5.1 is served over MoRI-EP + MoRI-IO, # so the UCX/RIXL/rocSHMEM/DeepEP stack is dead weight here: it lengthens the build and # ships transports this recipe never selects. Default 0 => lean MoRI-EP-only image, which # is also exactly how the validated image (glm5.1-vllm027-b8) was built. Set # --build-arg WITH_NIXL=1 only if you need the rixl connector from this same Dockerfile. ARG WITH_NIXL=0 -ARG NIC_COMPILATION_ARCH="cx7" +ARG NIC_COMPILATION_ARCH="ionic" # ----------------------------------------------------------------------------- # 1. MoRI: replace the base's bundled MoRI with the validated ROCm/MoRI @ v1.2.1 @@ -119,7 +119,11 @@ ARG MORI_REPO=https://github.com/ROCm/mori.git # base's bundled mori for debugging. ARG WITH_MORI_BUILD=1 ARG MORI_REF=42e895472b08 -ENV MORI_GPU_ARCHS=gfx942 +ENV MORI_GPU_ARCHS=gfx950 +# AAC MI355X: Pensando Ionic (AINIC). Pin device-side IBGDA dispatch to the ionic +# provider. Auto-detect would also work (8x rocep*s0 -> driver readlink -> ionic, +# libionic.so.1.1.54.0-184 present), but pin it so the JIT path is deterministic. +ENV MORI_DEVICE_NIC=ionic # Newer MoRI added the UMBP subsystem which requires gRPC (grpcpp/grpcpp.h) not # present in this base; UMBP is unrelated to the EP dispatch/combine kernels, so # disable it to avoid pulling in a gRPC build dependency. @@ -196,35 +200,34 @@ ARG VLLM_REPO=https://github.com/raviguptaamd/vllm.git # disagg long-ctx). NIAH-validated 1P/1D + 2P/1D + 1P/2D, 2k-35k, decode PIECEWISE. ARG VLLM_REF=glm5.1-dsa-wideEP_on_vllm-v0.27 ENV VLLM_TARGET_DEVICE=rocm \ - PYTORCH_ROCM_ARCH=${PYTORCH_ROCM_ARCH} \ - MAX_JOBS=${MAX_JOBS} + PYTORCH_ROCM_ARCH=${BUILD_ROCM_ARCH} \ + MAX_JOBS=${BUILD_MAX_JOBS} RUN rm -rf /tmp/vllm-src && \ git clone "${VLLM_REPO}" /tmp/vllm-src && \ cd /tmp/vllm-src && git checkout "${VLLM_REF}" && \ echo "VLLM_REF=${VLLM_REF}@$(git rev-parse HEAD)" >> /app/versions.txt && \ pip uninstall -y vllm 2>/dev/null || true && \ pip install --no-deps --no-build-isolation -v . && \ - python3 -c "import vllm; print('vLLM', vllm.__version__, 'from', vllm.__file__)" && \ + python3 -c "import importlib.metadata as m, pathlib; d=m.distribution('vllm'); print('vLLM', d.version, 'from', pathlib.Path(d.locate_file('vllm')))" && \ rm -rf /tmp/vllm-src # Cross-check MoRI + AITER survived the vLLM install (no silent downgrade). -RUN python3 - <<'PYEOF' -from importlib.metadata import version as v, PackageNotFoundError -def get(names): - for n in names: - try: return v(n) - except PackageNotFoundError: pass - return None -av = get(("amd-aiter", "amd_aiter", "aiter")) -# Verify the aiter install survived the vLLM install (present, not silently downgraded -# to a base-bundled wheel). We pin aiter by commit (e03fa6040), whose reported version -# string varies by build, so assert presence rather than a hardcoded commit substring. Do NOT -# `import aiter` here: it pulls torch->amdsmi->libamd_smi.so, not loadable in the no-GPU -# build sandbox (same reason the Stage-2 verify reads mla.py from disk instead). -assert av, "AITER missing after vLLM install (expected bundled 0.1.19 or source-built ref)" -import mori, mori.io, mori.ops -print("Post-vLLM check OK: AITER", av, "present + MoRI importable") -PYEOF +RUN python3 -c "$(printf '%s\n' \ + 'from importlib.metadata import version as v, PackageNotFoundError' \ + 'def get(names):' \ + ' for n in names:' \ + ' try: return v(n)' \ + ' except PackageNotFoundError: pass' \ + ' return None' \ + 'av = get(("amd-aiter", "amd_aiter", "aiter"))' \ + '# Verify the aiter install survived the vLLM install (present, not silently downgraded' \ + '# to a base-bundled wheel). We pin aiter by commit (e03fa6040), whose reported version' \ + '# string varies by build, so assert presence rather than a hardcoded commit substring. Do NOT' \ + '# `import aiter` here: it pulls torch->amdsmi->libamd_smi.so, not loadable in the no-GPU' \ + '# build sandbox (same reason the Stage-2 verify reads mla.py from disk instead).' \ + 'assert av, "AITER missing after vLLM install (expected bundled 0.1.19 or source-built ref)"' \ + 'import mori, mori.io, mori.ops' \ + 'print("Post-vLLM check OK: AITER", av, "present + MoRI importable")')" # ----------------------------------------------------------------------------- # 4. vllm-router (DP-rank round-robin + MoRIIO connector) — built in, so NO From 500353fd923faeeed327a865c34be065449595ff Mon Sep 17 00:00:00 2001 From: raviguptaamd Date: Sun, 16 Aug 2026 16:22:12 -0700 Subject: [PATCH 11/41] vllm_disagg/moriio: pass RDMA tuning via extra_config, widen GLM gate, disable 2 harmful patchers Three independent fixes to the MoRIIO connector wiring. 1. qp_per_transfer / num_workers / post_batch_size now go through kv_connector_extra_config instead of VLLM_MORIIO_* env vars, which the connector no longer reads. moriio_common.py:227-229 holds a rename map and _warn_deprecated_env_vars() only warns; moriio_engine.py reads no env vars at all; vllm/envs.py has no MORIIO entries. The values are consumed solely at moriio_common.py:342-344 via extra_config.get(), defaulting to 1/1/-1. So the env vars set in moriio.env were silently inert, and the live log read: Using MoRIIO backend: RDMA (qp_per_transfer=1, post_batch_size=-1, num_workers=1) i.e. one queue pair and one worker thread for the whole prefill->decode KV handoff on a node with 8 RDMA rails. Grep that line in the decode log to confirm the values now change. Also passes host_ip explicitly. The connector otherwise derives it from get_ip(), which on this cluster returns the non-routable public address, so decode cannot notify prefill of block allocation and the two deadlock. 2. The DSA patcher gate goes from an exact "GLM-5.1-FP8" match to the whole GLM DSA family (GLM-5.1-FP8 | GLM-5.2-FP8 | GLM-5.2-MXFP4). These are the same architecture -- GlmMoeDsaForCausalLM, 78 layers, 256+1 experts, index_topk=2048, verified by diffing config.json -- so they need identical patchers. Under the old exact match GLM-5.2 silently skipped all of them. 3. Two patchers become opt-in (default OFF). Both are actively harmful on this image; each retains a full rationale in-line, summarised here: apply_glm_dsa_persistent_kernel_gate_fix.py (GLM_PERSIST_GATE=1 to restore) Sets work_meta_data=None on chunked-prefill continuations, routing to a non-persistent kernel that does not exist for this cell: gqa_ratio=64 + fp8/fp8 hits AITER_CHECK(false) at asm_mla.cu:945 -> C++ abort, no Python unwind, server dies. Explains the reproducible 2000-ok / 8000-dies boundary (8192 = prefill max_num_batched_tokens). The image already carries aiter#3921 and vllm#47766, which superseded #47567 and whose diff is the literal inverse of this patcher. apply_glm_dsa_kernel_fix.py (GLM_DSA_SENTINEL_FIX=1 to restore) Flips the DSA invalid-token sentinel 0 -> -1 per vllm#45324. This image ships 0 deliberately: aiter's decode-only mla_decode_fwd sparse kernel dereferences paged_kv_indices, so -1 becomes kv_cache + (-1)*stride -> memory access fault, and it only bites under disagg, which is what we run. Applying it killed decode DP0 with hipErrorIllegalAddress on a 5-token warmup curl. With index_topk=2048 a short prompt is the worst case, not the safest. --- scripts/vllm_dissag/connectors/moriio.sh | 90 +++++++++++++++++++++--- 1 file changed, 82 insertions(+), 8 deletions(-) diff --git a/scripts/vllm_dissag/connectors/moriio.sh b/scripts/vllm_dissag/connectors/moriio.sh index e1c7a51a..2105ab25 100644 --- a/scripts/vllm_dissag/connectors/moriio.sh +++ b/scripts/vllm_dissag/connectors/moriio.sh @@ -118,9 +118,20 @@ connector_setup_env() { export MORI_SHMEM_HEAP_SIZE="${MORI_SHMEM_HEAP_SIZE:-17179869184}" } +# NOTE: qp_per_transfer / num_workers / post_batch_size MUST be passed through +# kv_connector_extra_config, NOT as VLLM_MORIIO_* env vars. Those env vars were +# deprecated out of the connector: moriio_common.py:227-229 holds an explicit +# rename map and _warn_deprecated_env_vars() only warns. moriio_engine.py reads +# NO env vars at all, and vllm/envs.py has no MORIIO entries. The values are read +# ONLY at moriio_common.py:342-344 via extra_config.get(...), defaulting to +# 1 / 1 / -1. Before this, the live log read: +# Using MoRIIO backend: RDMA (qp_per_transfer=1, post_batch_size=-1, num_workers=1) +# i.e. one queue pair and one worker thread for the entire prefill->decode KV +# handoff, on a node with 8 RDMA rails. Verify the fix by grepping decode log for +# that same line and confirming the numbers changed. _moriio_build_kv_transfer_config() { local kv_role="$1" - echo '{"kv_connector":"MoRIIOConnector","kv_role":"'"${kv_role}"'","kv_port":"'"${KV_PORT}"'","kv_connector_extra_config":{"proxy_ip":"'"${MASTER_ADDR}"'","proxy_port":"'"${PROXY_PORT}"'","proxy_ping_port":"'"${PROXY_PING_PORT}"'","http_port":"'"${SERVE_PORT}"'","local_ping_port":"'"${LOCAL_PING_PORT}"'","handshake_port":"'"${HANDSHAKE_PORT}"'","notify_port":"'"${NOTIFY_PORT}"'"}}' + echo '{"kv_connector":"MoRIIOConnector","kv_role":"'"${kv_role}"'","kv_port":"'"${KV_PORT}"'","kv_connector_extra_config":{"host_ip":"'"${host_ip}"'","proxy_ip":"'"${MASTER_ADDR}"'","proxy_port":"'"${PROXY_PORT}"'","proxy_ping_port":"'"${PROXY_PING_PORT}"'","http_port":"'"${SERVE_PORT}"'","local_ping_port":"'"${LOCAL_PING_PORT}"'","handshake_port":"'"${HANDSHAKE_PORT}"'","notify_port":"'"${NOTIFY_PORT}"'","qp_per_transfer":'"${VLLM_MORIIO_QP_PER_TRANSFER:-1}"',"num_workers":'"${VLLM_MORIIO_NUM_WORKERS:-1}"',"post_batch_size":'"${VLLM_MORIIO_POST_BATCH_SIZE:--1}"'}}' } connector_runtime_patch() { @@ -140,7 +151,15 @@ connector_runtime_patch() { # The MoRI version is pinned by the Dockerfile MORI_REF (post-1.2.1 main with the # large-transfer notify/mapping fixes #424/#436/#432 baked in); if a newer MoRI is # needed, update MORI_REF and rebuild the image — no runtime library swap here. - [ "${MODEL_NAME:-}" = "GLM-5.1-FP8" ] || return 0 + # Gate widened from an exact "GLM-5.1-FP8" match to the whole GLM DSA family: + # GLM-5.2-FP8 / GLM-5.2-MXFP4 are the SAME architecture (GlmMoeDsaForCausalLM, + # 78 layers, 256+1 experts, index_topk=2048 -- verified by diffing config.json), + # so they need the identical DSA patchers. Under the old exact match they would + # silently skip all of them and emit garbage or stall the KV transfer. + case "${MODEL_NAME:-}" in + GLM-5.1-FP8|GLM-5.2-FP8|GLM-5.2-MXFP4) ;; + *) return 0 ;; + esac _glm_dsa_runtime_patch } @@ -166,16 +185,71 @@ _glm_dsa_runtime_patch() { echo "Error: [glm] cannot locate vLLM install dir for DSA patchers. Aborting." >&2 exit 1 fi - echo "[glm] MODEL_NAME=GLM-5.1-FP8: applying DSA runtime patchers against ${_vllm_dir}" + echo "[glm] MODEL_NAME=${MODEL_NAME}: applying DSA runtime patchers against ${_vllm_dir}" # Ordered list of REQUIRED patchers (all abort on hard failure). - # GLM_PERSIST_GATE=0 skips the persistent-MLA accuracy gate (debug only: to test - # whether the non-persistent kernel it routes to is what crashes disagg at >=8k). - local _gate_patcher="apply_glm_dsa_persistent_kernel_gate_fix.py" - [ "${GLM_PERSIST_GATE:-1}" = "0" ] && _gate_patcher="" + # apply_glm_dsa_persistent_kernel_gate_fix.py is now OFF BY DEFAULT (was ON). + # + # That patcher ports vLLM #47567: for chunked-prefill CONTINUATION batches it + # sets work_meta_data=None to dodge a numerically-wrong persistent sparse-MLA + # kernel. Its own docstring targets aiter 0.1.16.post3, "before the aiter-side + # kernel fix #3921". On THIS image it is fatal, not merely redundant: + # + # work_meta_data=None + # -> aiter/mla.py:247 persistent_mode = False + # -> non-persistent branch passes a real num_kv_splits_indptr + # -> asm_mla.cu:680 persistent = (num_kv_splits_indptr == nullptr) -> false + # -> asm_mla.cu:945 gqa_ratio=64 + fp8/fp8 + !persistent -> AITER_CHECK(false) + # "fp8/fp8 with gqa_ratio=64 only supports persistent mode" + # -> C++ abort, no Python unwind, worker exit code None, server dies. + # + # GLM-5.2 is 64 heads / 1 latent KV head -> gqa_ratio 64, fp8 weights AND fp8 KV + # -> fp8/fp8. That is the ONLY cell in asm_mla.cu with no non-persistent kernel + # (bf16/gqa64 has one; fp8/gqa32 has one; gqa8 has both). The gate routes to a + # kernel that was never built. It bites at >=8k words -- where a prompt first + # exceeds prefill max_num_batched_tokens=8192 and becomes a chunked + # continuation -- matching the observed 2000-ok / 8000-dies boundary exactly, + # reproduced three times. Reported independently upstream from gfx942 hardware + # (vllm-project/vllm#49649) quoting the same asm_mla.cu check. + # + # This image carries BOTH the fixes that made the gate obsolete: aiter#3921 + # (merged 2026-06-26) and vllm#47766 (merged 2026-07-08) are ancestors of our + # aiter 0.1.17.dev195+ge03fa6040 and vLLM v0.16.0rc2.dev5996 -- verified + # behind_by=0 via the GitHub compare API. #47766 SUPERSEDED #47567 and its diff + # is the literal inverse of this patcher: it deletes the + # is_chunked_continuation/use_persistent block and restores work_meta_data + # unconditionally. Its real repair was the metadata fingerprint (the old 4-field + # key collided across batches with different chunk boundaries, so + # get_mla_metadata_v1 was skipped and the kernel ran on a stale work schedule). + # + # Set GLM_PERSIST_GATE=1 only for an image that genuinely predates #47766 AND + # does not hit the crash cell (bf16, or gqa_ratio != 64). + local _gate_patcher="" + [ "${GLM_PERSIST_GATE:-0}" = "1" ] && _gate_patcher="apply_glm_dsa_persistent_kernel_gate_fix.py" + # apply_glm_dsa_kernel_fix.py is OPT-IN (default OFF). It implements vllm #45324, + # flipping the DSA invalid-token sentinel 0 -> -1 in + # _convert_req_index_to_global_index_kernel (rocm_aiter_mla_sparse.py). On THIS + # image that patch CAUSES a decode crash: the image ships 0 deliberately, with the + # verbatim comment above the tl.where -- + # "output 0 (NOT -1): the downstream aiter mla_decode_fwd sparse kernel + # dereferences paged_kv_indices, so a -1 becomes kv_cache + (-1)*stride -> + # page-aligned GPU memory access fault. Only bites at disagg" + # We run disagg. With the patch applied, decode DP0 died with + # hipErrorIllegalAddress on the router's 5-token warm-up curl, while PREFILL ran + # the identical DSA indexer JIT chain and survived -- because aiter's + # mla_decode_fwd (mla_a8w8_qh64_qseqlen1_*) sparse kernel is DECODE-ONLY; prefill + # scores with _gluon_fp8_mqa_logits_kernel and never dereferences these indices. + # index_topk=2048 means a 5-token prompt has ~all 2048 slots set to the sentinel, + # so a SHORT prompt is the worst case for this bug, not the safest. + # The patcher cannot self-skip: it keys on the literal 0 and treats it as "the + # known bug" by definition, and the function docstring 30 lines below is STALE + # (still describes outputting -1). Set GLM_DSA_SENTINEL_FIX=1 only for an older + # image that genuinely ships the #45324 bug and lacks the aiter dereference. + local _dsa_sentinel_patcher="" + [ "${GLM_DSA_SENTINEL_FIX:-0}" = "1" ] && _dsa_sentinel_patcher="apply_glm_dsa_kernel_fix.py" local _p for _p in \ - apply_glm_dsa_kernel_fix.py \ + ${_dsa_sentinel_patcher} \ apply_glm_dsa_moriio_dualkv_fix.py \ apply_glm_dsa_moriio_engine_fix.py \ apply_glm_dsa_moriio_gate_fix.py \ From d40c5e3823a185f49b5dc4e583cc847892cea369 Mon Sep 17 00:00:00 2001 From: raviguptaamd Date: Sun, 16 Aug 2026 16:22:50 -0700 Subject: [PATCH 12/41] vllm_disagg/models.yaml: add GLM-5.2-FP8 and GLM-5.2-MXFP4 recipes for MI355X Adds the two GLM-5.2 recipes and corrects two stale values inherited from the GLM-5.1/MI300X recipe they were derived from. New: GLM-5.2-FP8 (zai-org/GLM-5.2-FP8) and GLM-5.2-MXFP4 (amd/GLM-5.2-MXFP4), 1P/1D EP8 over MoRI-IO. Measured on 2x8 MI355X at ISL/OSL 28672/1024: TPOT 33.7 / 36.4 / 40.2 ms at concurrency 16 / 32 / 64, against a 50 ms target. GLM-5.2-MXFP4 is CONFIG-ONLY AND HAS NEVER BEEN BOOTED. It is a copy of the FP8 recipe with the two values that provably do not transfer re-derived. Treat its first run as bring-up, not benchmark. Two values carry the recipe; both are noted in full in-file. DECODE_CUDAGRAPH_MODE: FULL_AND_PIECEWISE (was PIECEWISE) The single change that fixed TPOT: 104 -> 34-40 ms. PIECEWISE splits the decode graph on the 3 DSA ops x 78 layers, ~234 launch boundaries per step. The cost is batch-INDEPENDENT, so it presents as a latency floor that does not move with concurrency, batch size or fabric -- which is why it survived so much tuning before being found. Co-requisite: use_inductor_graph_partition must stay ON and no bare --enforce-eager, else boot and warmup both pass and the first real MLA decode dies. prefill.dp: --gpu-memory-utilization 0.72 Prefill only. The MLA chunked-prefill workspace is sized from a hardcoded 64k-token clamp (determine_chunked_prefill_workspace_size, mla_attention.py:1935), NOT from max_num_batched_tokens, so it allocates 65536*64*(192+256)*2 = 3.50 GiB on top of the KV pool regardless of scheduler batch size -- and lazily, on the first long prefill, i.e. after boot and warmup have both reported healthy. At 0.80 this OOM'd at mla_attention.py:739. Corrections to the shared GLM section: VLLM_SPARSE_INDEXER_MAX_LOGITS_MB: 64 -> 2048 64MB was an MI300X crash workaround, and the crash cannot occur here: the workaround kernel is behind a hard arch gate (rocm_aiter_mla_sparse.py:591 dispatches fp8_mqa_logits_gfx942 only if _ON_GFX942). Verified live in-image: _ON_GFX942=False, arch=gfx950. The module self-describes as a "Temporary gfx942 fallback" for that part's 64 KiB LDS budget. Keeping 64MB cost 50 indexer sub-chunks at ISL 28672 where 2048MB needs 2, each a full kernel chain over 78 layers. 2048 not 4096 deliberately: nearly all the reduction at half the peak memory (2.00 GiB, against ~31 GiB/GPU idle). KV_BLOCK_SIZE comment corrected (value unchanged at 1) It claimed the DSA sparse indexer "REQUIRES" block-size 1. That inverts the real constraint: rocm_aiter_mla_sparse.py:283 returns [1, 64] and indexer.py:137 returns `[1, 64] if current_platform.is_rocm() else [64]` -- ROCm is the PERMISSIVE branch and it is non-ROCm that is locked to 64. Measured 1 vs 64: <2% either way, so 1 stays as the validated default, but the comment would have stopped anyone from trying 64. The same section's "decode PIECEWISE" prose is updated to match the new value. --- scripts/vllm_dissag/models.yaml | 362 ++++++++++++++++++++++++++++++-- 1 file changed, 348 insertions(+), 14 deletions(-) diff --git a/scripts/vllm_dissag/models.yaml b/scripts/vllm_dissag/models.yaml index 064ced6d..9648c723 100644 --- a/scripts/vllm_dissag/models.yaml +++ b/scripts/vllm_dissag/models.yaml @@ -201,12 +201,21 @@ DeepSeek-R1: # # GLM differs from DeepSeek in 3 recipe-defining ways (both are MLA MoE, but GLM # is DSA-sparse): -# - KV_BLOCK_SIZE=1 (DSA sparse indexer REQUIRES block-size 1; DS uses 16) +# - KV_BLOCK_SIZE=1 (DS uses 16. NOT a hard DSA requirement on ROCm -- an +# earlier revision of this comment claimed the sparse indexer "REQUIRES" 1, which +# inverts the actual constraint: rocm_aiter_mla_sparse.py:283 returns [1, 64] and +# indexer.py:137 returns `[1, 64] if current_platform.is_rocm() else [64]`, i.e. +# ROCm is the PERMISSIVE branch and it is non-ROCm that is locked to 64. Measured +# 1 vs 64 on this stack: <2% either way, so 1 is kept as the validated default, +# but 64 is a legal experiment -- if you try it, see the block_size assert in +# moriio_connector.py, which may need to be relaxed to an override.) # - VLLM_ROCM_USE_AITER_MLA=1 (GLM MLA path ON via AITER sparse; DS sets 0) -# - prefill EAGER (NONE), decode PIECEWISE (prefill cudagraph capture deadlocks -# on this stack; decode PIECEWISE captures cleanly on DSA and is a ~3.7x ITL win -# (validated: ~69ms vs ~264ms eager). Global VLLM_CUDAGRAPH_MODE=NONE as the -# safe floor; per-role PREFILL=NONE / DECODE=PIECEWISE override it.) +# - prefill EAGER (NONE), decode FULL_AND_PIECEWISE (prefill cudagraph capture +# deadlocks on this stack. Decode was PIECEWISE historically, but PIECEWISE splits +# the decode graph on the 3 DSA ops x 78 layers (~234 launch boundaries/step) -- +# a batch-INDEPENDENT ~64ms floor. FULL_AND_PIECEWISE measured TPOT 104ms -> 34-40ms +# at concurrency 16/32/64. Global VLLM_CUDAGRAPH_MODE=NONE as the safe floor; +# per-role PREFILL=NONE / DECODE=FULL_AND_PIECEWISE override it.) # Note: block=1 + AITER_MLA=1 are ALREADY the moriio.sh connector defaults — GLM # keeps them; DeepSeek is the one that overrides them off. Set here explicitly so # the recipe is self-documenting and robust to connector default changes. @@ -264,15 +273,39 @@ GLM-5.1-FP8: # matching the published EP8 figure. The knob is decode.dp below -- prefill is # unaffected (it genuinely dispatches wide 8192-token chunked-prefill batches). MORI_SHMEM_HEAP_SIZE: "17179869184" - # DSA sparse-indexer logits-buffer cap (crash fix). The indexer prefill computes an - # M*N fp32 logits buffer; split_indexer_prefill_chunks only sub-chunks the query dim - # when M*N*4 > this budget. Default 512MB lets an 8192-token prefill build a single - # 268MB buffer + launch the fp8_mqa_logits kernel at grid=(8192,), which HARD-FAULTS - # the worker on gfx942 (silent GPU fault -> DP group collapse -> 503 at >=8k prompts). - # Capping at 64MB forces M-dim sub-chunking (~2k tokens/chunk) so the buffer and the - # kernel launch stay bounded. Root cause: vllm/v1/attention/ops/triton_fp8_mqa_logits.py - # fp8_mqa_logits_gfx942; chunking logic: mla/indexer.py split_indexer_prefill_chunks. - VLLM_SPARSE_INDEXER_MAX_LOGITS_MB: "64" + # DSA sparse-indexer logits-buffer cap. The indexer prefill computes an M*N fp32 + # logits buffer; split_indexer_prefill_chunks only sub-chunks the query dim when + # M*N*4 > this budget. + # + # The 64MB value was a gfx942 (MI300X) CRASH WORKAROUND: there, a single 8192-token + # prefill built a 268MB buffer and launched fp8_mqa_logits at grid=(8192,), which + # hard-faulted the worker (silent GPU fault -> DP group collapse -> 503 at >=8k). + # + # THAT CRASH CANNOT HAPPEN ON gfx950 (MI355X). The vendored workaround kernel is + # behind a hard architecture gate: rocm_aiter_mla_sparse.py:591 dispatches + # fp8_mqa_logits_gfx942 only `if _ON_GFX942`. Verified live in-image on mi355-gpu-45: + # _ON_GFX942 = False, arch = gfx950:sramecc+:xnack-. We take AITER's mainline kernel. + # The gfx942 module even self-describes as "Temporary gfx942 fallback" for that + # part's 64 KiB LDS budget -- an MI300X constraint, not an MI355X one. + # + # Cost of keeping 64MB here: it is 8x TIGHTER than upstream's 512MB default, and the + # sub-chunk count is ISL-dependent. Simulated with vLLM's own + # split_indexer_prefill_chunks at ISL=28672 (1 request): + # 64MB -> 50 sub-chunks 512MB -> 7 2048MB -> 2 4096MB -> 1 + # Each sub-chunk is a separate kernel chain over all 78 layers, so 64MB pays ~50x the + # per-chunk launch/setup overhead. Same failure CLASS as the decode-cudagraph finding + # (DECODE_CUDAGRAPH_MODE, section 44): work needlessly split into many small pieces, + # each paying fixed cost, showing up as a batch-independent latency floor. + # + # 2048MB -> 2 sub-chunks, peak logits buffer 2.00 GiB. We have ~31 GiB/GPU HBM idle + # (engine suggests --kv-cache-memory=126.47GiB vs 95.08GiB in use), so this is well + # inside headroom rather than scraping the last drop. NOT raised to 4096MB (1 chunk) + # deliberately: 2048 captures nearly all of the 50x reduction at half the peak memory. + # + # If a >=8k prefill ever faults after this change, this line is the first suspect -- + # revert to "64" and re-test. Chunking logic: mla/indexer.py + # split_indexer_prefill_chunks; dispatch gate: ops/rocm_aiter_mla_sparse.py:591. + VLLM_SPARSE_INDEXER_MAX_LOGITS_MB: "2048" # NCCL heartbeat watchdog: at long context (>~8k) a DP rank's sparse-MLA/MoE # all2all collective can exceed the default HeartbeatMonitor timeout -> # ProcessGroupNCCL::HeartbeatMonitor::runLoop() declares the rank dead and @@ -300,3 +333,304 @@ GLM-5.1-FP8: # Value must stay >= typical prompt length: 512 gave 87.9ms TPOT but 13.1s TTFT # (a 1024-token prompt could not be admitted in one step). 2048 keeps TTFT healthy. dp: "--max-num-batched-tokens 2048" + + +# ============================================================================= +# GLM-5.2 (AIMODELS-1198). Weights are pre-staged on the AAC shared FS -- no HF +# download: /shared/data/amd_int/models/GLM-5.2-FP8 (704G, 141 shards) +# /shared/data/amd_int/models/GLM-5.2-MXFP4 (408G, 282 shards) +# +# GLM-5.2 == GLM-5.1 architecturally (verified by diffing config.json): +# GlmMoeDsaForCausalLM, 78 layers (3 dense + 75 MoE), 256 routed + 1 shared expert, +# top-8, sigmoid+noaux_tc scale 2.5, index_topk=2048, first_k_dense_replace=3, +# max_position_embeddings=1048576. +# Therefore the entire validated GLM-5.1-FP8 recipe below transfers unchanged. Both +# entries are wideEP-only (add to WIDE_EP_ONLY_MODELS in run_xPyD_models.slurm). +# +# CARRY-OVER RISK to re-verify on the gfx950 image, NOT assumed fixed: the GLM-5.1 +# note above records ROCM_AITER_MLA_SPARSE prefill corruption beyond ~16-18k tokens on +# the OLD mori-v1.2.1 / gfx942 image. Our 28k/1k target sits ABOVE that threshold. The +# new image is a different vLLM ref (glm5.1-dsa-wideEP_on_vllm-v0.27) on gfx950, so the +# bug may or may not be present -- this is precisely what the NIAH sweep must answer +# BEFORE any 28k perf number is meaningful. Run NIAH first; a clean 20k+35k needle is +# the gate for trusting the 28k/1k throughput run. +# ============================================================================= +GLM-5.2-FP8: + env: + VLLM_USE_V1: "1" + # v0.27: layer_name is wrapped in a torch OpaqueBase (LayerName) and passed through + # the unified_mla_kv_cache_update / unified_mla_attention custom ops. On this ROCm + # torch 2.12 build the opaque-type boxing FAILS -> "RuntimeError: unknown parameter + # type" at torch/_ops.py on the first real MLA decode forward (the fake/compile path + # returns early, so it passes boot+warmup then crashes on first request -> DP gloo + # cascade -> 0 prefill/0 decode). VLLM_USE_LAYERNAME=0 makes LayerNameType=str (the + # pre-2.11 path), so a plain string passes through the op. No image rebuild needed. + VLLM_USE_LAYERNAME: "0" + VLLM_ROCM_USE_AITER: "1" + VLLM_ROCM_USE_AITER_RMSNORM: "1" + VLLM_ROCM_USE_AITER_MLA: "1" + KV_BLOCK_SIZE: "1" + KV_CACHE_DTYPE: "fp8" + GPU_MEMORY_UTILIZATION: "0.80" + VLLM_CUDAGRAPH_MODE: "NONE" + PREFILL_CUDAGRAPH_MODE: "NONE" + # PERF A/B 2026-08-16: PIECEWISE cuts the graph at every op in splitting_ops. + # For DSA that list holds THREE hot ops (unified_mla_attention_with_output, + # sparse_attn_indexer, rocm_aiter_sparse_attn_indexer) -> ~3 segments x 78 layers + # ~= 234 launch boundaries per decode step. That is a FIXED per-step cost, which + # matches the measured signature: TPOT flat at 103.8/104.3/104.9 ms across + # con=16/32/64, and only ~40ms of the ~104ms step is accounted for by compute. + # Both loaded backends (ROCMAiterMLASparseBackend rocm_aiter_mla_sparse.py:364 and + # the DSA indexer mla/indexer.py:462) declare AttentionCGSupport.UNIFORM_BATCH, + # which is exactly the level compilation.py requires for full DECODE-ONLY graphs; + # the downgrade cascade only forces PIECEWISE when support is NEVER. There is no + # mori guard in compilation.py (deepep_high_throughput IS force-disabled, mori is + # not), and mori dispatch/combine are sync-free (the .item()/.cpu()/synchronize() + # calls live in get_dispatch_src_token_pos, a debug helper). + # Verified in-image: resolve_cudagraph_mode_and_sizes(min_cg_support=UNIFORM_BATCH, + # backend=ROCMAiterMLASparseBackend, all2all=mori_low_latency, dp=8) returns + # FULL_AND_PIECEWISE UNCHANGED -> decode=FULL, mixed=PIECEWISE. Chosen over + # FULL_DECODE_ONLY because that one sets mixed=NONE (eager mixed batches); + # FULL_AND_PIECEWISE is a strict superset of today behaviour: decode gets full + # graphs, mixed keeps the piecewise graphs it already had. It is also the vLLM v1 + # default. use_inductor_graph_partition MUST stay on -- it is what keeps the + # STABLE-ABI concat_and_cache_mla out of the compiled graph; without it the first + # real MLA decode dies with "unknown parameter type". Never express "no graphs" as + # bare --enforce-eager: that drops +quant_fp8 and hits an AITER + # dynamic_per_token_scaled_quant signature mismatch at engine init. + # Fallback order if capture fails or regresses: FULL_DECODE_ONLY, then PIECEWISE. + DECODE_CUDAGRAPH_MODE: "FULL_AND_PIECEWISE" + CUDAGRAPH_CAPTURE_SIZES: "1 2 4 8 16 32 64 128 256" + VLLM_ALL2ALL_BACKEND: "mori_high_throughput" + PREFILL_MORI_BACKEND: "mori_high_throughput" + DECODE_MORI_BACKEND: "mori_low_latency" + # MoRI EP dispatch/combine buffer width. Without this it inherits + # max_num_batched_tokens (8192) -- a chunked-prefill SCHEDULER setting -- so every + # decode step moves an 8192-token-wide buffer per layer x78 layers regardless of the + # real batch. That is a fixed ~300ms/step floor (~320x this model's HBM-bandwidth + # bound). Sizing it for the actual decode batch gives TPOT 302ms -> 88ms (3.4x), + # matching the published EP8 figure. The knob is decode.dp below -- prefill is + # unaffected (it genuinely dispatches wide 8192-token chunked-prefill batches). + MORI_SHMEM_HEAP_SIZE: "17179869184" + # DSA sparse-indexer logits-buffer cap. The indexer prefill computes an M*N fp32 + # logits buffer; split_indexer_prefill_chunks only sub-chunks the query dim when + # M*N*4 > this budget. + # + # The 64MB value was a gfx942 (MI300X) CRASH WORKAROUND: there, a single 8192-token + # prefill built a 268MB buffer and launched fp8_mqa_logits at grid=(8192,), which + # hard-faulted the worker (silent GPU fault -> DP group collapse -> 503 at >=8k). + # + # THAT CRASH CANNOT HAPPEN ON gfx950 (MI355X). The vendored workaround kernel is + # behind a hard architecture gate: rocm_aiter_mla_sparse.py:591 dispatches + # fp8_mqa_logits_gfx942 only `if _ON_GFX942`. Verified live in-image on mi355-gpu-45: + # _ON_GFX942 = False, arch = gfx950:sramecc+:xnack-. We take AITER's mainline kernel. + # The gfx942 module even self-describes as "Temporary gfx942 fallback" for that + # part's 64 KiB LDS budget -- an MI300X constraint, not an MI355X one. + # + # Cost of keeping 64MB here: it is 8x TIGHTER than upstream's 512MB default, and the + # sub-chunk count is ISL-dependent. Simulated with vLLM's own + # split_indexer_prefill_chunks at ISL=28672 (1 request): + # 64MB -> 50 sub-chunks 512MB -> 7 2048MB -> 2 4096MB -> 1 + # Each sub-chunk is a separate kernel chain over all 78 layers, so 64MB pays ~50x the + # per-chunk launch/setup overhead. Same failure CLASS as the decode-cudagraph finding + # (DECODE_CUDAGRAPH_MODE, section 44): work needlessly split into many small pieces, + # each paying fixed cost, showing up as a batch-independent latency floor. + # + # 2048MB -> 2 sub-chunks, peak logits buffer 2.00 GiB. We have ~31 GiB/GPU HBM idle + # (engine suggests --kv-cache-memory=126.47GiB vs 95.08GiB in use), so this is well + # inside headroom rather than scraping the last drop. NOT raised to 4096MB (1 chunk) + # deliberately: 2048 captures nearly all of the 50x reduction at half the peak memory. + # + # If a >=8k prefill ever faults after this change, this line is the first suspect -- + # revert to "64" and re-test. Chunking logic: mla/indexer.py + # split_indexer_prefill_chunks; dispatch gate: ops/rocm_aiter_mla_sparse.py:591. + VLLM_SPARSE_INDEXER_MAX_LOGITS_MB: "2048" + # NCCL heartbeat watchdog: at long context (>~8k) a DP rank's sparse-MLA/MoE + # all2all collective can exceed the default HeartbeatMonitor timeout -> + # ProcessGroupNCCL::HeartbeatMonitor::runLoop() declares the rank dead and + # tears down the whole process group -> prefill EngineCore crashes -> 503. + # (Confirmed root cause of the 8k+ prefill crash; #338 EP-landmine.) Disable + # the monitor-triggered teardown and extend timeouts so long-ctx collectives + # complete instead of being watchdog-killed. + TORCH_NCCL_ENABLE_MONITORING: "0" + TORCH_NCCL_HEARTBEAT_TIMEOUT_SEC: "1800" + TORCH_NCCL_DUMP_ON_TIMEOUT: "0" + TORCH_NCCL_BLOCKING_WAIT: "0" + TORCH_NCCL_ASYNC_ERROR_HANDLING: "1" + NCCL_IB_TIMEOUT: "22" + dp_flags: "--tool-call-parser glm47 --reasoning-parser glm45 --enable-auto-tool-choice --chat-template-content-format string" + prefill: + # --gpu-memory-utilization here OVERRIDES the global one that + # connectors/moriio.sh emits at line 344, because model_args is appended LAST + # (line 355) and argparse takes the last occurrence. It is per-role + # (MODEL_CONFIG_PREFILL vs MODEL_CONFIG_DECODE, moriio.sh:275), so decode + # keeps 0.80. run_xPyD_models.slurm has no PREFILL_GPU_MEMORY_UTILIZATION. + # + # WHY: prefill OOMed in compile_or_warm_up_model -- AFTER the KV cache was + # already resident -- on the MLA chunked-prefill warm-up workspace + # (mla_attention.py:739): + # 65536 * 64 heads * (192 qk_nope + 256 v) * 2 B = 3.50 GiB + # Free at that moment: 3.39 GiB. Short by ~110 MiB, with 5.53 GiB already + # reserved-but-unallocated (fragmented), so no contiguous 3.50 GiB block. + # + # The 65536 is HARDCODED, not a batch-size effect: + # determine_chunked_prefill_workspace_size (mla_attention.py:1935) = + # min(max(8*max_model_len, 4*max_num_seqs*block_size), 64*1024) + # max_model_len=1048576 -> first term 8M -> min clamps to 64k. + # So --max-num-batched-tokens CANNOT shrink it (and prefill is already capped + # at 8192 by default: "Chunked prefill is enabled with + # max_num_batched_tokens=8192"). The only lever is leaving room for it. + # + # 0.80 -> 0.72 reclaims 0.08 * 287.98 = 23.0 GiB. KV goes 76.58 -> ~54 GiB + # (~1.2M tokens at block_size 1), which is ample for 28k/1k at concurrency 64 + # in a 1P/1D split where prefill streams KV out over MoRIIO instead of holding + # it. Decode stays 0.80 -- it boots clean and its pool is what bounds + # concurrency here. + # + # Prefill needs a lower budget than decode because its profiled activation + # peak is higher (8192-token chunked-prefill batches vs decode's 2048 cap): + # same 0.80 gave prefill 76.58 GiB KV but decode 95.08 GiB. ~23 GiB also sits + # outside torch accounting (MoRI's 16 GiB pinned symmetric heap + + # mori_high_throughput buffers): 253.43 GiB allocated vs a 230.4 GiB budget. + # + # NOT using PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True (which the OOM + # message suggests and would probably work): it remaps GPU VA ranges while + # MoRIIO has the KV cache registered with the ionic NIC and MoRI holds a + # pinned symmetric heap. Silent KV corruption there would poison the NIAH + # accuracy run that comes next. Revisit only with a clean baseline in hand. + # + # If prefill-side preemption shows up later, raise toward 0.76 rather than + # reverting to 0.80. + dp: "--gpu-memory-utilization 0.72" + decode: + # PERF: the MoRI EP dispatch buffer width is max_num_batched_tokens (via + # FusedMoEConfig.max_num_tokens -> all2all.py max_num_inp_token_per_rank), so the + # decode role otherwise runs an 8192-token-wide all2all every step: ~302ms TPOT. + # mori bounds recv capacity BY the send width (MaxNumTokensToRecvPerRank returns + # min(ceil(maxTotalRecvTokens/ws), maxNumInpTokenPerRank)), so the buffer must still + # cover vLLM's profiling dummy run -- it cannot be shrunk via env alone. Lowering + # this knob on the DECODE role lowers both consistently. + # Value must stay >= typical prompt length: 512 gave 87.9ms TPOT but 13.1s TTFT + # (a 1024-token prompt could not be admitted in one step). 2048 keeps TTFT healthy. + dp: "--max-num-batched-tokens 2048" + + +# GLM-5.2-MXFP4 (amd/GLM-5.2-MXFP4, Quark). OCP MXFP4 per_group g=32 with e8m0 scales -- +# native in gfx950/CDNA4 matrix cores, which is why this variant is a MI355X-specific +# play (MI308X/gfx942 has no MXFP4 hardware path). 408G vs 704G leaves ~300G more HBM +# for KV, i.e. materially more room at 80K-200K context. +# Deltas vs the FP8 entry: +# - KV_CACHE_DTYPE stays fp8: MXFP4 quantizes WEIGHTS; the KV cache is separate and +# fp8 KV is the validated DSA path. Do not set kv-cache-dtype to fp4. +# - GPU_MEMORY_UTILIZATION raised 0.80 -> 0.90: weights are 42% smaller so the +# same fraction wastes ~120G/node. Back off to 0.80 if it OOMs during capture. +# Everything else is identical -- same architecture, same DSA constraints. +# STATUS: unvalidated on this stack. vLLM #51915 (open) is the MXFP4-on-deepseek_v32 +# enablement PR; if MXFP4 fails to load, that PR is the first place to look. +GLM-5.2-MXFP4: + env: + VLLM_USE_V1: "1" + # v0.27: layer_name is wrapped in a torch OpaqueBase (LayerName) and passed through + # the unified_mla_kv_cache_update / unified_mla_attention custom ops. On this ROCm + # torch 2.12 build the opaque-type boxing FAILS -> "RuntimeError: unknown parameter + # type" at torch/_ops.py on the first real MLA decode forward (the fake/compile path + # returns early, so it passes boot+warmup then crashes on first request -> DP gloo + # cascade -> 0 prefill/0 decode). VLLM_USE_LAYERNAME=0 makes LayerNameType=str (the + # pre-2.11 path), so a plain string passes through the op. No image rebuild needed. + VLLM_USE_LAYERNAME: "0" + VLLM_ROCM_USE_AITER: "1" + VLLM_ROCM_USE_AITER_RMSNORM: "1" + VLLM_ROCM_USE_AITER_MLA: "1" + KV_BLOCK_SIZE: "1" + KV_CACHE_DTYPE: "fp8" + GPU_MEMORY_UTILIZATION: "0.90" + VLLM_CUDAGRAPH_MODE: "NONE" + PREFILL_CUDAGRAPH_MODE: "NONE" + # Same value and same reason as GLM-5.2-FP8 above: PIECEWISE splits the decode graph on + # the 3 DSA ops x 78 layers (~234 launch boundaries/step), a batch-INDEPENDENT ~64ms + # floor. FULL_AND_PIECEWISE measured TPOT 104ms -> 34-40ms on FP8. Quantisation does not + # change the decode graph topology, so the fix carries over unchanged. + # HARD CO-REQUISITE: use_inductor_graph_partition must stay ON (see the FP8 section) -- + # otherwise boot and warmup both PASS and the first real MLA decode dies. + DECODE_CUDAGRAPH_MODE: "FULL_AND_PIECEWISE" + CUDAGRAPH_CAPTURE_SIZES: "1 2 4 8 16 32 64 128 256" + VLLM_ALL2ALL_BACKEND: "mori_high_throughput" + PREFILL_MORI_BACKEND: "mori_high_throughput" + DECODE_MORI_BACKEND: "mori_low_latency" + # MoRI EP dispatch/combine buffer width. Without this it inherits + # max_num_batched_tokens (8192) -- a chunked-prefill SCHEDULER setting -- so every + # decode step moves an 8192-token-wide buffer per layer x78 layers regardless of the + # real batch. That is a fixed ~300ms/step floor (~320x this model's HBM-bandwidth + # bound). Sizing it for the actual decode batch gives TPOT 302ms -> 88ms (3.4x), + # matching the published EP8 figure. The knob is decode.dp below -- prefill is + # unaffected (it genuinely dispatches wide 8192-token chunked-prefill batches). + MORI_SHMEM_HEAP_SIZE: "17179869184" + # DSA sparse-indexer logits-buffer cap. The indexer prefill computes an M*N fp32 + # logits buffer; split_indexer_prefill_chunks only sub-chunks the query dim when + # M*N*4 > this budget. + # + # The 64MB value was a gfx942 (MI300X) CRASH WORKAROUND: there, a single 8192-token + # prefill built a 268MB buffer and launched fp8_mqa_logits at grid=(8192,), which + # hard-faulted the worker (silent GPU fault -> DP group collapse -> 503 at >=8k). + # + # THAT CRASH CANNOT HAPPEN ON gfx950 (MI355X). The vendored workaround kernel is + # behind a hard architecture gate: rocm_aiter_mla_sparse.py:591 dispatches + # fp8_mqa_logits_gfx942 only `if _ON_GFX942`. Verified live in-image on mi355-gpu-45: + # _ON_GFX942 = False, arch = gfx950:sramecc+:xnack-. We take AITER's mainline kernel. + # The gfx942 module even self-describes as "Temporary gfx942 fallback" for that + # part's 64 KiB LDS budget -- an MI300X constraint, not an MI355X one. + # + # Cost of keeping 64MB here: it is 8x TIGHTER than upstream's 512MB default, and the + # sub-chunk count is ISL-dependent. Simulated with vLLM's own + # split_indexer_prefill_chunks at ISL=28672 (1 request): + # 64MB -> 50 sub-chunks 512MB -> 7 2048MB -> 2 4096MB -> 1 + # Each sub-chunk is a separate kernel chain over all 78 layers, so 64MB pays ~50x the + # per-chunk launch/setup overhead. Same failure CLASS as the decode-cudagraph finding + # (DECODE_CUDAGRAPH_MODE, section 44): work needlessly split into many small pieces, + # each paying fixed cost, showing up as a batch-independent latency floor. + # + # 2048MB -> 2 sub-chunks, peak logits buffer 2.00 GiB. We have ~31 GiB/GPU HBM idle + # (engine suggests --kv-cache-memory=126.47GiB vs 95.08GiB in use), so this is well + # inside headroom rather than scraping the last drop. NOT raised to 4096MB (1 chunk) + # deliberately: 2048 captures nearly all of the 50x reduction at half the peak memory. + # + # If a >=8k prefill ever faults after this change, this line is the first suspect -- + # revert to "64" and re-test. Chunking logic: mla/indexer.py + # split_indexer_prefill_chunks; dispatch gate: ops/rocm_aiter_mla_sparse.py:591. + VLLM_SPARSE_INDEXER_MAX_LOGITS_MB: "2048" + # NCCL heartbeat watchdog: at long context (>~8k) a DP rank's sparse-MLA/MoE + # all2all collective can exceed the default HeartbeatMonitor timeout -> + # ProcessGroupNCCL::HeartbeatMonitor::runLoop() declares the rank dead and + # tears down the whole process group -> prefill EngineCore crashes -> 503. + # (Confirmed root cause of the 8k+ prefill crash; #338 EP-landmine.) Disable + # the monitor-triggered teardown and extend timeouts so long-ctx collectives + # complete instead of being watchdog-killed. + TORCH_NCCL_ENABLE_MONITORING: "0" + TORCH_NCCL_HEARTBEAT_TIMEOUT_SEC: "1800" + TORCH_NCCL_DUMP_ON_TIMEOUT: "0" + TORCH_NCCL_BLOCKING_WAIT: "0" + TORCH_NCCL_ASYNC_ERROR_HANDLING: "1" + NCCL_IB_TIMEOUT: "22" + dp_flags: "--tool-call-parser glm47 --reasoning-parser glm45 --enable-auto-tool-choice --chat-template-content-format string" + prefill: + # PREFILL ONLY. The MLA chunked-prefill workspace is sized from a hardcoded 64k-token + # clamp (determine_chunked_prefill_workspace_size, mla_attention.py:1935), NOT from + # max_num_batched_tokens, so it allocates 65536*64*(192+256)*2 = 3.50 GiB on top of the + # KV pool no matter how small the scheduler batch is -- and lazily, on the first real + # long prefill, i.e. AFTER boot and warmup have both reported healthy. + # On FP8 at 0.80 this OOM'd at mla_attention.py:739; 0.72 was the validated headroom. + # This variant runs GPU_MEMORY_UTILIZATION 0.90 (weights are 42% smaller), so it has + # LESS free HBM at the moment that 3.50 GiB lands, not more. Same 0.72 prefill clamp. + # NOTE: inferred from the FP8 failure, not measured here -- MXFP4 has not been booted. + dp: "--gpu-memory-utilization 0.72" + decode: + # PERF: the MoRI EP dispatch buffer width is max_num_batched_tokens (via + # FusedMoEConfig.max_num_tokens -> all2all.py max_num_inp_token_per_rank), so the + # decode role otherwise runs an 8192-token-wide all2all every step: ~302ms TPOT. + # mori bounds recv capacity BY the send width (MaxNumTokensToRecvPerRank returns + # min(ceil(maxTotalRecvTokens/ws), maxNumInpTokenPerRank)), so the buffer must still + # cover vLLM's profiling dummy run -- it cannot be shrunk via env alone. Lowering + # this knob on the DECODE role lowers both consistently. + # Value must stay >= typical prompt length: 512 gave 87.9ms TPOT but 13.1s TTFT + # (a 1024-token prompt could not be admitted in one step). 2048 keeps TTFT healthy. + dp: "--max-num-batched-tokens 2048" From 24b5c9cc9fb8409bef67c8ea6c96f2fc12cb4bf4 Mon Sep 17 00:00:00 2001 From: raviguptaamd Date: Sun, 16 Aug 2026 16:23:30 -0700 Subject: [PATCH 13/41] vllm_disagg: make the launcher portable across sites, and add the AAC/ionic env file Everything needed to run this recipe on a second cluster without editing the launcher. Previously several site facts were baked in as defaults from the MI300X cluster it was written on. FABRIC_SUBNET is the important one. It defaults to 10.158. (the MI300X fabric); AAC is 10.2.80.. When it does not match, the awk picker falls through to the first address on the line -- a 10.224.x overlay -- MASTER_ADDR is advertised as something the peer cannot reach, and the run hangs at "Waiting for nodes" forever with no error. So a fresh clone of this recipe would have hung on AAC. Unlike every other key in the connector .env files, FABRIC_SUBNET is consumed by the launcher on the HOST rather than inside the container, and the .env loader only built `-e KEY=VAL` docker args. Setting it in the .env file would therefore have been silently inert. The loader now has an explicit one-key host-export allow-list. Deliberately an allow-list and not a blanket export: the other keys are HSA/RDMA runtime settings that should not leak into srun and other host-side tooling. Submit-time `export FABRIC_SUBNET=` still wins over the file, as before. Also in the launcher: - CONNECTOR_ENV_FILE is overridable, so a site can supply its own platform env (e.g. connectors/moriio.env.aac) without touching the tracked default. - CONTAINER_CLI auto-selects docker or podman by probing for a usable docker daemon. Compute nodes here deny the daemon socket and only have podman. The two differ in ways that matter: podman gets --group-add keep-groups (video is not a valid group there) and no --shm-size (it does not accept it in this configuration). - MODEL_DIR defaults to /shared/data/amd_int/models when that exists. - Optional site mounts (/shared/data, /shared/apps, /media/NVME, ...) are mounted only when present, instead of failing on absence. - The JIT cache host path is chosen from a candidate list keyed by image ID, and falls back with a loud warning rather than silently paying a cold compile every run. - GLM-5.2-FP8 and GLM-5.2-MXFP4 added to the three model allow-lists. - New niah / niah_perf benchmark cases. connectors/moriio.env.aac (new) is the AAC MI355X + Pensando Ionic platform env: 8 ionic rails (rocep{9,25,105,121,137,153,233,249}s0) at MORI_IB_GID_INDEX=1, control/bootstrap on the mgmt NIC via MORI_SOCKET_IFNAME=enp193s0f1np1, and FABRIC_SUBNET=10.2.80.. The interface choice is not the obvious one and the file records why: the public/default-route NIC answers ICMP but has TCP firewalled node-to-node, so "the one that pings" picks the wrong link. Test the actual port -- measured gpu-44 -> gpu-45, /dev/tcp//22 FAILS while /dev/tcp/10.2.80.16/22 succeeds. MORI_SOCKET_IFNAME left undefined fell through to a hardcoded eth0 that does not exist here, ShmemGetUniqueId failed, and the null shmem context then segfaulted inside EpDispatchCombineHandle -> ShmemNumQpPerPe(). --- scripts/vllm_dissag/connectors/moriio.env.aac | 185 ++++++++++++++++++ scripts/vllm_dissag/run_xPyD_models.slurm | 141 +++++++++++-- 2 files changed, 310 insertions(+), 16 deletions(-) create mode 100644 scripts/vllm_dissag/connectors/moriio.env.aac diff --git a/scripts/vllm_dissag/connectors/moriio.env.aac b/scripts/vllm_dissag/connectors/moriio.env.aac new file mode 100644 index 00000000..9e8e90ca --- /dev/null +++ b/scripts/vllm_dissag/connectors/moriio.env.aac @@ -0,0 +1,185 @@ +# --------------------------------------------------------------------------- +# AAC MI355X (gfx950 / CDNA4) + Pensando Ionic "AINIC" RoCEv2 site overrides. +# +# Use INSTEAD of connectors/moriio.env on the AAC cluster: +# CONNECTOR_ENV_FILE=.../moriio.env.aac (or copy over moriio.env) +# Same format: KEY=VALUE, one per line; the slurm turns each into +# `-e KEY=${KEY:-VALUE}` so a submit-time export still wins. +# +# Every value below that differs from the stock (OCI MI300X + CX7) file is marked +# CHANGED with the measurement that justifies it. Measured on mi355-gpu-44/45, +# job 5635, 2026-08-16. +# --------------------------------------------------------------------------- + +# --- Platform-mandatory (unchanged from stock; see moriio.env for the rationale: +# ROCm cannot dmabuf-export HIP-VMM memory, so expandable_segments must be False on +# both allocators or MoRI RegisterRdmaMemoryRegion EFAULTs on the first disagg WRITE). +PYTORCH_ALLOC_CONF=expandable_segments:False +PYTORCH_HIP_ALLOC_CONF=expandable_segments:False +HSA_ENABLE_IPC_MODE_LEGACY=0 +HSA_NO_SCRATCH_RECLAIM=1 + +# CHANGED gfx942 -> gfx950. MI355X is CDNA4. Matches the image's baked MORI_GPU_ARCHS. +MORI_GPU_ARCHS=gfx950 + +# CHANGED (new). Pin MoRI's device-NIC dispatch to the ionic provider. Auto-detect +# WOULD land here anyway, but only via a fallback: MoRI's sysfs check matches device +# NAME prefixes (^mlx5, ^bnxt_re, ^ionic) and our devices are named rocep*s0, so the +# name match misses and it falls through to `readlink device/driver` -> ionic. Inside a +# container /sys/class/infiniband can be masked, which would silently drop it to the +# mlx5 default. Pin it so the JIT path is deterministic. +MORI_DEVICE_NIC=ionic + +# --- RDMA fabric. All values below are cluster-specific and were MEASURED here. +# +# CHANGED 3 -> 1. GID index 1 is the RoCEv2 entry on Pensando Ionic. Proven by +# ib_write_bw between gpu-44/45: idx 1 succeeds, idx 0 fails to connect. Independently +# corroborated by MAD commit 61dd42c, which set MORI_IB_GID_INDEX=1 for this fabric. +MORI_IB_GID_INDEX=1 +NCCL_IB_GID_INDEX=1 + +# CHANGED mlx5_* -> the 8 ionic rails. Measured on both nodes: 8 x rocep{9,25,105,121, +# 137,153,233,249}s0, driver=ionic, 400000 Mb/s, state ACTIVE -- these are the GPU +# data-plane NICs (1 per GPU). EXCLUDED: rocep193s0f0/f1, which are driver=bnxt_en at +# 200000 Mb/s and carry the management/default route. Without this allowlist MoRI +# enumerates all 10 ibv devices and tries to raise QPs over the mgmt pair -> QP +# connection timeout at the prefill->decode KV transfer. +MORI_RDMA_DEVICES=rocep9s0,rocep25s0,rocep105s0,rocep121s0,rocep137s0,rocep153s0,rocep233s0,rocep249s0 +NCCL_IB_HCA=rocep9s0,rocep25s0,rocep105s0,rocep121s0,rocep137s0,rocep153s0,rocep233s0,rocep249s0 + +# CHANGED eth0 -> enp193s0f1np1 (was briefly enp193s0f0np0 -- wrong, see below). +# There is no eth0 on AAC. Both nodes expose the same two named interfaces: +# enp193s0f0np0 216.128.144.179 / 104.238.162.159 public, default route +# enp193s0f1np1 10.2.80.11 / 10.2.80.16 management +# The obvious pick is the default-route interface, but the PUBLIC IPs answer ICMP +# while their TCP is FIREWALLED between the nodes. Measured gpu-44 -> gpu-45: +# /dev/tcp/104.238.162.159/22 -> FAIL (public) +# /dev/tcp/10.2.80.16/22 -> OK (mgmt) +# so every control-plane socket must use the mgmt interface. This is the same +# firewall that hung the launch barrier until FABRIC_SUBNET=10.2.80. was set. +# +# ...which is set right here. run_xPyD_models.slurm defaults FABRIC_SUBNET to the +# MI300X cluster's 10.158. prefix; without this line the awk picker at :467 falls +# through to $1 = the 10.224.x overlay, MASTER_ADDR is advertised as an address the +# peer cannot reach, and the run hangs at "Waiting for nodes" forever with no error. +# Unlike every other key in this file, FABRIC_SUBNET is consumed by the launcher on +# the HOST, not inside the container -- the .env loader has an explicit host-export +# allow-list for exactly this key (see run_xPyD_models.slurm:255-263). +FABRIC_SUBNET=10.2.80. +# +# MORI_SOCKET_IFNAME is what MoRI shmem uses to bootstrap ShmemGetUniqueId. Left +# undefined it fell through to the hardcoded eth0 default in +# run_xPyD_models.slurm:721, ShmemGetUniqueId failed, and the null shmem context +# then SEGFAULTED inside EpDispatchCombineHandle -> ShmemNumQpPerPe(). Defining it +# here overrides that default because CONNECTOR_ENV_ARGS is expanded at line 744, +# after 721, and the last -e wins. +# +# Control/bootstrap plane only -- bulk EP and KV traffic still ride the 8 ionic +# rails above via MORI_IB_GID_INDEX=1, so the mgmt link speed does not matter here. +MORI_SOCKET_IFNAME=enp193s0f1np1 +NCCL_SOCKET_IFNAME=enp193s0f1np1 +GLOO_SOCKET_IFNAME=enp193s0f1np1 + +# --- Tuning carried over from the validated MI300X+CX7 recipe. These are NOT +# AINIC-tuned; treat them as a starting point to sweep once 1P/1D is functionally +# green. TC/SL in particular are fabric-QoS dependent and may not map on Pensando. +MORI_RDMA_TC=41 +MORI_RDMA_SL=0 +MORI_IO_SL=1 +MORI_IB_ENABLE_RELAXED_ORDERING=1 +MORI_NUM_QP_PER_PE=8 +VLLM_MORIIO_QP_PER_TRANSFER=8 +VLLM_MORIIO_NUM_WORKERS=8 +# -1 = let MoRI choose the WR post batch size. +VLLM_MORIIO_POST_BATCH_SIZE=-1 +HSA_FORCE_FINE_GRAIN_PCIE=1 +HSA_ENABLE_SDMA=1 +# --- RDMA fork safety ----------------------------------------------------------- +# MoRI registers a 16 GiB pinned symmetric heap (MORI_SHMEM_HEAP_SIZE) with the +# ionic NIC. Immediately afterwards the AITER fp8 BMM precompile walks batch sizes +# 1..1024 across 78 layers, and Triton forks a compiler subprocess per unseen +# kernel variant -- the highest fork rate in the whole engine init. Without fork +# safety, copy-on-write can remap the registered pages in the PARENT while the NIC +# still DMAs to the old physical pages, producing a host-side SIGSEGV with no GPU +# fault. That is exactly the observed failure: decode segfaults at launchKernel and +# prefill then trips gloo nfds != -1 (epoll on a peer that vanished). +# +# RDMAV_FORK_SAFE makes libibverbs madvise(MADV_DONTFORK) the registered regions so +# the child never inherits them. IBV_FORK_SAFE is the older synonym; set both for +# driver-version portability. No cost when nothing forks. +# +# The BMM precompile itself is innocent -- reproduced standalone at real GLM-5.2 +# dims, at 8 concurrent ranks on one JIT cache, and under 200 GiB of ballast: +# all passed all 78 layers. The pinned heap is the only variable those repros lacked. +RDMAV_FORK_SAFE=1 +IBV_FORK_SAFE=1 + +# --- AITER fp8 BMM precompile: DISABLED pending root-cause ----------------------- +# Both nodes SIGSEGV inside this precompile, top frame launchKernel, ~20 layers in. +# The loop itself is innocent: reproduced standalone at real GLM-5.2 dims, at 8 +# concurrent ranks on one JIT cache, and under 200 GiB ballast -- all passed all 78 +# layers. Note the progress bars complete at ~1300 it/s (sub-second, cache-warm), so +# Triton is NOT compiling here; it is replaying cached kernels. That means the fault +# is in the LAUNCH path, not the compile path. +# +# Setting this to 0 removes ~160k kernel launches from engine init -- that is the +# reason it is off, and it is a good one. If the engine boots with this off, the +# fault is confined to the precompile launch storm. If it still dies, the corruption +# lives in MoRI EP setup and the precompile was merely the first heavy GPU work to +# touch it. +# +# NOT a warm-up-only knob (an earlier version of this comment said it was). It is a +# permanent hot-path kernel selector, read once at layer construction and branched on +# every forward. Traced in vLLM fe1c317 (this image's code): +# envs.py:147 defaults True UPSTREAM +# _aiter_ops.py:1626 "Controls FP8 batched matrix multiply." +# _aiter_ops.py:1918 is_fp8bmm_enabled() +# mla_attention.py:611 self.is_aiter_triton_fp8_bmm_enabled = is_fp8bmm_enabled() +# mla_attention.py:898 ON -> triton_fp8_bmm(q, W_K, ...) +# mla_attention.py:921 OFF -> torch.bmm(mqa_q_nope, W_UK_T, ...) +# mla_attention.py:1186 ON -> triton_fp8_bmm(x, W_V, ...) +# mla_attention.py:1193 OFF -> torch.bmm(x, self.W_UV, ...) +# It also picks which weights are materialised at load (:1081-1126): ON keeps fp8 +# W_K/W_V + scales, OFF keeps bf16 W_UV/W_UK_T. The settings never converge -- the +# aiter kernels are not "compiled later", they are never used. +# +# PERF: negligible, so do NOT chase this for the TPOT floor. GLM-5.2 at TP8 touches +# 78 layers * 8 heads * (192*512 + 512*256) = 143M weight elements per decode step in +# these two BMMs: 0.29 GB/step bf16 vs 0.14 GB/step fp8, i.e. ~0.04 vs ~0.02 ms at +# MI355X's ~8 TB/s -- about 0.1% of a 144 ms step, and memory-bound (2.3 GFLOP at +# batch 8). The decode floor is the MoRI EP all2all buffer width; see the GLM-5.2-FP8 +# decode block in models.yaml, where it is already sized. +# +# Correctness is unaffected either way (same math, different kernel), but because it +# IS a real kernel + weight-dtype change, this is a legitimate A/B if numerics ever +# look wrong -- which the old "pure warm-up optimisation" wording would have talked +# you out of. +VLLM_ROCM_USE_AITER_FP8BMM=0 + +# --- disable the flydsl GEMM backend (version skew inside the image) ------------ +# aiter/ops/flydsl/kernels/tensor_shim.py:34 calls flydsl.compiler.from_c_void_p, +# which the installed flydsl 0.1.8 does not expose (it has from_dlpack instead). +# Decode therefore died during KV-cache profiling with +# RuntimeError: module flydsl.compiler has no attribute from_c_void_p +# reached via tuned_gemm.py:554 flydsl_gemm -> gemm_kernels.py:941 flydsl_hgemm. +# +# The version guard in aiter/ops/flydsl/__init__.py does not catch this: it +# requires >=0.1.8 and the wheel reports exactly 0.1.8, so the floor passes while +# the API is still stale. +# +# is_flydsl_available() (aiter/ops/flydsl/utils.py:72) returns False when the arch +# is outside flydsl's SMEM_CAPACITY_MAP (keys: gfx942 gfx950 gfx1201 gfx1250), and +# flydsl's get_rocm_arch() honours FLYDSL_GPU_ARCH first. So this uses the module's +# own documented "report unavailable rather than fail the import" path. +# +# Verified in-image with GPUs attached: FLYDSL_AVAIL False, flydsl_hgemm no longer +# exported, while aiter's own get_gfx() still returns gfx950 and all 8 GPUs are +# visible. FLYDSL_GPU_ARCH is read only by flydsl and by the offline aiter/aot +# build scripts -- nothing on the runtime path. Do NOT use HSA_OVERRIDE_GFX_VERSION +# for this: the ROCm runtime reads that one and it would mis-target every kernel. +# +# Fallback is explicit, not a raise (tuned_gemm.py:231-245): config = None -> +# continue -> next candidate -> hipblaslt/asm -> default_config. PERF NOTE: some +# GEMM shapes will run an untuned kernel. Revisit by shipping a matching flydsl +# wheel if GEMM is hot in the 28k/1k profiles. +FLYDSL_GPU_ARCH=gfx000 diff --git a/scripts/vllm_dissag/run_xPyD_models.slurm b/scripts/vllm_dissag/run_xPyD_models.slurm index 365883b8..02d54730 100755 --- a/scripts/vllm_dissag/run_xPyD_models.slurm +++ b/scripts/vllm_dissag/run_xPyD_models.slurm @@ -104,6 +104,8 @@ VALID_MODELS=( \ "Qwen3-32B" \ "Qwen3-30B-A3B" \ "GLM-5.1-FP8" \ + "GLM-5.2-FP8" \ + "GLM-5.2-MXFP4" \ ) # Models allowed for CONNECTOR=moriio WIDE_EP=1 (MoRI-EP; legacy RUN_MORI=1) @@ -112,6 +114,8 @@ MORI_EP_VALID_MODELS=( \ "DeepSeek-V3-5layer" \ "DeepSeek-R1" \ "GLM-5.1-FP8" \ + "GLM-5.2-FP8" \ + "GLM-5.2-MXFP4" \ ) # Models allowed for CONNECTOR=rixl WIDE_EP=1 EP_BACKEND=deepep (legacy RUN_DEEPEP=1) @@ -200,7 +204,7 @@ WIDE_EP="${WIDE_EP:-0}" # GLM-5.1-FP8 (GlmMoeDsaForCausalLM, MLA+DSA) is validated only under MoRI-EP # wideEP disagg (block=1, AITER sparse MLA on, per-role all2all). The moriio+TP # ("Stage B") path is untested for DSA, so reject WIDE_EP=0 for it too. -WIDE_EP_ONLY_MODELS=( "DeepSeek-V3" "DeepSeek-V3-5layer" "DeepSeek-R1" "GLM-5.1-FP8" ) +WIDE_EP_ONLY_MODELS=( "DeepSeek-V3" "DeepSeek-V3-5layer" "DeepSeek-R1" "GLM-5.1-FP8" "GLM-5.2-FP8" "GLM-5.2-MXFP4" ) model_is_wide_ep_only() { local m="$1" for x in "${WIDE_EP_ONLY_MODELS[@]}"; do [[ "$m" == "$x" ]] && return 0; done @@ -240,7 +244,7 @@ echo "Launcher: $RUN_FILE (CONNECTOR=${CONNECTOR} WIDE_EP=${WIDE_EP} EP_BACKEND # KEY=VALUE lines into CONNECTOR_ENV_ARGS as `-e KEY=${KEY:-VALUE}` pairs, so a # submit-time export of the same name still overrides. Forwarded in the docker run. # --------------------------------------------------------------------------- -CONNECTOR_ENV_FILE="${SCRIPT_DIR}/connectors/${CONNECTOR}.env" +CONNECTOR_ENV_FILE="${CONNECTOR_ENV_FILE:-${SCRIPT_DIR}/connectors/${CONNECTOR}.env}" CONNECTOR_ENV_ARGS="" if [[ -f "$CONNECTOR_ENV_FILE" ]]; then echo "Loading connector platform env: $CONNECTOR_ENV_FILE" @@ -248,6 +252,14 @@ if [[ -f "$CONNECTOR_ENV_FILE" ]]; then [[ "$_line" =~ ^[[:space:]]*# || -z "${_line// }" ]] && continue _k="${_line%%=*}"; _v="${_line#*=}" CONNECTOR_ENV_ARGS+=" -e ${_k}=${!_k:-$_v}" # submit-time export of $_k wins + # Most keys here are container-runtime only, so `-e` is enough. A few are needed + # by THIS script, on the HOST, before any container starts -- those must also be + # exported into the host shell or setting them in the .env file is silently inert. + # Deliberately an allow-list: blanket-exporting the RDMA/HSA vars onto the host + # would also leak them into srun and every other host-side tool. + case "$_k" in + FABRIC_SUBNET) export "$_k=${!_k:-$_v}" ;; + esac done < "$CONNECTOR_ENV_FILE" else echo "WARN: connector env file not found: $CONNECTOR_ENV_FILE" >&2 @@ -271,6 +283,12 @@ LOG_PATH="${LOG_PATH:-/shared_inference/${USER}/model_blog_logs}" xP="${xP:-1}" #-> Number of Prefill Servers yD="${yD:-1}" #-> Number of Decode Servers +# AAC MI355X: weights are pre-staged on the 100T /shared/data NFS export (visible on +# every node), not on the OCI-style models_blog paths. The two checks above will miss, +# and this third one resolves. Override MODEL_DIR at submit time for other sites. +if [ -z "${MODEL_DIR:-}" ] && [ -d /shared/data/amd_int/models ]; then + MODEL_DIR=/shared/data/amd_int/models +fi MODEL_DIR="${MODEL_DIR:-"/shared_inference/models_blog/"}" @@ -430,6 +448,17 @@ MASTER_NODE=$(echo "$SELECTED_NODES" | head -n 1) # multiple NICs (e.g. a 10.224.x overlay listed BEFORE the routable 10.158.x fabric); # taking $1 blindly can advertise an unreachable addr -> prefill/decode barrier hangs # "Waiting for nodes" forever. Prefer FABRIC_SUBNET (default 10.158.), fall back to $1. +# +# THIS DEFAULT IS CLUSTER-SPECIFIC AND YOU PROBABLY NEED TO OVERRIDE IT. +# Known values: +# 10.158. MI300X cluster this recipe was originally developed on (the default) +# 10.2.80. AAC MI355X (gpu-44/45); set in connectors/moriio.env.aac +# If it does not match, MASTER_ADDR silently becomes an unreachable address and the +# launch barrier hangs at "Waiting for nodes" with no error. Check with: +# srun -N1 -w hostname -I +# and pick the prefix of the interface that is routable BETWEEN nodes -- note that on +# AAC the public/default-route NIC answers ICMP but has TCP firewalled node-to-node, +# so "the one that pings" is not a safe test; test the actual port. FABRIC_SUBNET="${FABRIC_SUBNET:-10.158.}" # From a "hostname -I" line, return the first IP on FABRIC_SUBNET, else the first IP. _pick_fabric_ip() { @@ -462,6 +491,16 @@ case "$BENCHMARK_SCRIPT" in sweep) BENCHMARK_SCRIPT_FILE="benchmark_xPyD.sh" ;; long_context) BENCHMARK_SCRIPT_FILE="benchmark_long_context.sh" ;; keepalive) BENCHMARK_SCRIPT_FILE="keepalive_bench.sh" ;; + # NOTE: this case OVERWRITES any inherited BENCHMARK_SCRIPT_FILE (plain `=`, + # not `${VAR:-}`), so exporting BENCHMARK_SCRIPT_FILE=benchmark_niah.sh from a + # driver script is silently discarded -- several runs executed the default + # throughput sweep instead of NIAH before this was spotted. Select via + # BENCHMARK_SCRIPT instead; that is the knob this case actually keys on. + niah) BENCHMARK_SCRIPT_FILE="benchmark_niah.sh" ;; + # NIAH accuracy first, then the ISL/OSL perf sweep, against ONE booted server. + # Perf is SKIPPED if NIAH fails: a throughput number from a server that cannot + # retrieve a needle is not a result worth reporting. + niah_perf) BENCHMARK_SCRIPT_FILE="benchmark_niah_perf.sh" ;; *) echo "Error: invalid BENCHMARK_SCRIPT='$BENCHMARK_SCRIPT' (valid: sweep, long_context, keepalive)" >&2; exit 1 ;; esac if [[ ! -f "$BENCHMARK_SCRIPT_FILE" ]]; then @@ -509,15 +548,27 @@ export RUN_FILE_FULL="$NIXL_COOKBOOK_PATH/${RUN_FILE}" # Use only the selected nodes for srun execution SELECTED_NODELIST_SRUN=$(echo "$SELECTED_NODES" | paste -sd,) +# Container runtime selection. AAC compute nodes deny the docker socket to normal +# users (no `docker` group) but ship podman 4.7.1, which is CLI-compatible for every +# flag used below (run/ps/rm/pull/image inspect/stop). Prefer docker when actually +# usable, else podman. Override with CONTAINER_CLI=podman|docker at submit time. +if [ -z "${CONTAINER_CLI:-}" ]; then + if docker info >/dev/null 2>&1; then CONTAINER_CLI=docker + elif command -v podman >/dev/null 2>&1; then CONTAINER_CLI=podman + else echo "Error: neither a usable docker daemon nor podman found." >&2; exit 1; fi +fi +export CONTAINER_CLI +echo "[runtime] CONTAINER_CLI=$CONTAINER_CLI" + srun --nodelist="$SELECTED_NODELIST_SRUN" bash -c ' echo "Rank $SLURM_PROCID on $(hostname)"; -docker ps -q | xargs --no-run-if-empty docker stop; -docker rm -f $DOCKER_CONT_NAME 2>/dev/null || true; +$CONTAINER_CLI ps -q | xargs --no-run-if-empty $CONTAINER_CLI stop; +$CONTAINER_CLI rm -f $DOCKER_CONT_NAME 2>/dev/null || true; fuser -k 5000/tcp 2>/dev/null || true; fuser -k 2222/tcp 2>/dev/null || true; fuser -k 15000/tcp 2>/dev/null || true; sleep 2; -docker pull $DOCKER_IMAGE_NAME 2>/dev/null || true; +$CONTAINER_CLI pull $DOCKER_IMAGE_NAME 2>/dev/null || true; # --- Create host-local compilation cache dirs (ext4, survives container restarts) --- mkdir -p /tmp/vllm_cache/{aiter_jit,triton,vllm,comgr} 2>/dev/null || true; @@ -532,13 +583,30 @@ mkdir -p /tmp/vllm_cache/{aiter_jit,triton,vllm,comgr} 2>/dev/null || true; # NOTE: this whole section runs inside a single-quoted `srun bash -c '...'`, so avoid # single quotes here; the image-id hash is extracted with tr, not sed. if [ "${JIT_CACHE_PERSIST:-1}" = "1" ]; then - _IMG_RAW=$(docker image inspect --format "{{.Id}}" "$DOCKER_IMAGE_NAME" 2>/dev/null); + _IMG_RAW=$($CONTAINER_CLI image inspect --format "{{.Id}}" "$DOCKER_IMAGE_NAME" 2>/dev/null); _IMG_KEY=$(printf "%s" "$_IMG_RAW" | tr -cd "a-f0-9" | cut -c1-12); _IMG_KEY="${_IMG_KEY:-noimg}"; - _JIT_CACHE_HOST="${JIT_CACHE_HOST:-/mnt/m2m_nobackup/${USER}/vllm_jit_cache/${_IMG_KEY}}"; - mkdir -p "$_JIT_CACHE_HOST"/{aiter_jit,triton,vllm,comgr} 2>/dev/null || true; - _JIT_CACHE_MOUNT="-v ${_JIT_CACHE_HOST}:/opt/vllm_cache"; - echo "[jit-cache] persistent image ${_IMG_KEY}: ${_JIT_CACHE_HOST} to /opt/vllm_cache"; + _JIT_CACHE_HOST=""; + if [ -n "${JIT_CACHE_HOST:-}" ]; then + _JIT_CACHE_CANDS="${JIT_CACHE_HOST}"; + elif [ -n "${JIT_CACHE_ROOT:-}" ]; then + _JIT_CACHE_CANDS="${JIT_CACHE_ROOT}/${USER}/vllm_jit_cache/${_IMG_KEY}"; + else + _JIT_CACHE_CANDS="/mnt/m2m_nobackup/${USER}/vllm_jit_cache/${_IMG_KEY} /media/NVME/${USER}/vllm_jit_cache/${_IMG_KEY} /tmp/${USER}/vllm_jit_cache/${_IMG_KEY} ${HOME}/.cache/vllm_jit/$(hostname -s)/${_IMG_KEY}"; + fi + for _c in $_JIT_CACHE_CANDS; do + if mkdir -p "$_c"/aiter_jit "$_c"/triton "$_c"/vllm "$_c"/comgr 2>/dev/null && [ -w "$_c" ]; then + _JIT_CACHE_HOST="$_c"; break; + fi + done + if [ -n "$_JIT_CACHE_HOST" ]; then + _JIT_CACHE_MOUNT="-v ${_JIT_CACHE_HOST}:/opt/vllm_cache"; + echo "[jit-cache] persistent image ${_IMG_KEY}: ${_JIT_CACHE_HOST} to /opt/vllm_cache"; + else + _JIT_CACHE_MOUNT=""; + echo "[jit-cache] WARNING: no writable cache root among: $_JIT_CACHE_CANDS"; + echo "[jit-cache] falling back to in-container cache -- expect a slow cold compile"; + fi else _JIT_CACHE_MOUNT=""; fi @@ -569,21 +637,50 @@ done [ -d /etc/libibverbs.d ] && _RDMA_MOUNTS="$_RDMA_MOUNTS -v /etc/libibverbs.d:/etc/libibverbs.d:ro" echo "[host-rdma] mounts: $_RDMA_MOUNTS" -docker run --rm \ +# Optional site mounts: only bind paths that EXIST on this node. podman fails container +# creation outright on a missing bind source (docker would auto-create a dir), and the +# AAC cluster has neither /shared_inference nor /mnt/m2m_nobackup -- its weights are on +# /shared/data and its scratch on /media/NVME. +# Rootless podman drops supplementary groups, so a bind mount whose PARENT dir is +# group-readable only (here /shared/data/amd_int is 0770 amd:QLE) fails to stat. +# keep-groups retains them. Docker runs its daemon as root and does not need this. +# keep-groups cannot be combined with any other --group-add, and it makes +# --group-add video redundant (the user is already in video and render). +if [ "${CONTAINER_CLI}" = "podman" ]; then + _GROUP_ADD_ARGS="--group-add keep-groups" +else + _GROUP_ADD_ARGS="--group-add video" +fi + +# --shm-size is incompatible with --ipc host under podman (docker ignores the +# redundant pair, podman errors). With --ipc host the container uses the host +# /dev/shm, so the shm-size request is a no-op anyway -- drop it for podman only. +if [ "${CONTAINER_CLI}" = "podman" ]; then + _SHM_SIZE_ARG="" +else + _SHM_SIZE_ARG="--shm-size ${DOCKER_SHM_SIZE:-256G}" +fi + +_OPT_MOUNTS="" +for _d in /shared_inference /mnt/m2m_nobackup /shared/data /shared/apps /media/NVME /it-share; do + [ -d "$_d" ] && _OPT_MOUNTS="$_OPT_MOUNTS -v $_d:$_d" +done +echo "[site-mounts] $_OPT_MOUNTS" + +$CONTAINER_CLI run --rm \ --device /dev/dri \ --device /dev/kfd \ --device /dev/infiniband \ --network host \ --ipc host \ - --group-add video \ + ${_GROUP_ADD_ARGS} \ --cap-add SYS_PTRACE \ --security-opt seccomp=unconfined \ --privileged \ -v $HOME:$HOME \ - -v /shared_inference:/shared_inference \ - -v /mnt/m2m_nobackup:/mnt/m2m_nobackup \ + ${_OPT_MOUNTS} \ -v $HOME/.ssh:/root/.ssh \ - --shm-size ${DOCKER_SHM_SIZE:-256G} \ + ${_SHM_SIZE_ARG} \ --ulimit nofile=524288:524288 \ --ulimit memlock=-1:-1 \ -v ${LOG_PATH}:/run_logs \ @@ -609,10 +706,21 @@ docker run --rm \ -e BENCHMARK_ITR=$BENCHMARK_ITR \ -e BENCHMARK_CON="${BENCHMARK_CON}" \ -e BENCHMARK_COMBINATIONS="${BENCHMARK_COMBINATIONS}" \ + ${NIAH_WORDS:+-e NIAH_WORDS="$NIAH_WORDS"} \ + ${NIAH_WARMUP:+-e NIAH_WARMUP=$NIAH_WARMUP} \ + ${NIAH_SEEDS:+-e NIAH_SEEDS=$NIAH_SEEDS} \ + ${NIAH_MAXTOK:+-e NIAH_MAXTOK=$NIAH_MAXTOK} \ + ${NIAH_TIMEOUT:+-e NIAH_TIMEOUT=$NIAH_TIMEOUT} \ + ${NIAH_MIN_SCORE:+-e NIAH_MIN_SCORE=$NIAH_MIN_SCORE} \ + ${NIAH_GATE:+-e NIAH_GATE=$NIAH_GATE} \ + ${PERF_CON:+-e PERF_CON="$PERF_CON"} \ + ${PERF_COMBINATIONS:+-e PERF_COMBINATIONS="$PERF_COMBINATIONS"} \ + ${SHAPE_WARMUP:+-e SHAPE_WARMUP=$SHAPE_WARMUP} \ ${BENCHMARK_PORT:+-e BENCHMARK_PORT=$BENCHMARK_PORT} \ ${PROXY_TYPE:+-e PROXY_TYPE=$PROXY_TYPE} \ ${ROUTER_PORT:+-e ROUTER_PORT=$ROUTER_PORT} \ -e IPADDRS=$IPADDRS \ + ${FABRIC_SUBNET:+-e FABRIC_SUBNET=$FABRIC_SUBNET} \ ${CONNECTOR:+-e CONNECTOR=$CONNECTOR} \ ${WIDE_EP:+-e WIDE_EP=$WIDE_EP} \ ${EP_BACKEND:+-e EP_BACKEND=$EP_BACKEND} \ @@ -655,6 +763,7 @@ docker run --rm \ ${MORI_NUM_QP_PER_PE:+-e MORI_NUM_QP_PER_PE=$MORI_NUM_QP_PER_PE} \ ${VLLM_MORIIO_QP_PER_TRANSFER:+-e VLLM_MORIIO_QP_PER_TRANSFER=$VLLM_MORIIO_QP_PER_TRANSFER} \ ${VLLM_MORIIO_NUM_WORKERS:+-e VLLM_MORIIO_NUM_WORKERS=$VLLM_MORIIO_NUM_WORKERS} \ + ${VLLM_MORIIO_POST_BATCH_SIZE:+-e VLLM_MORIIO_POST_BATCH_SIZE=$VLLM_MORIIO_POST_BATCH_SIZE} \ -e GPU_MEMORY_UTILIZATION=${GPU_MEMORY_UTILIZATION:-0.8} \ -e GPUS_PER_NODE=${GPUS_PER_NODE:-8} \ ${GPU_MAX_HW_QUEUES:+-e GPU_MAX_HW_QUEUES=$GPU_MAX_HW_QUEUES} \ @@ -681,5 +790,5 @@ docker run --rm \ $RUN_FILE_FULL 2>&1 | tee /run_logs/${SLURM_JOB_ID}/pd_vllm_bench_NODE${SLURM_PROCID}.log " ' -srun --nodelist="$SELECTED_NODELIST_SRUN" bash -c 'docker stop $DOCKER_CONT_NAME 2>/dev/null || true; docker rm $DOCKER_CONT_NAME 2>/dev/null || true' +srun --nodelist="$SELECTED_NODELIST_SRUN" bash -c '$CONTAINER_CLI stop $DOCKER_CONT_NAME 2>/dev/null || true; $CONTAINER_CLI rm $DOCKER_CONT_NAME 2>/dev/null || true' From adc163cd4c6b204d15c57fc9606e2d16a4a4572c Mon Sep 17 00:00:00 2001 From: raviguptaamd Date: Sun, 16 Aug 2026 16:23:57 -0700 Subject: [PATCH 14/41] vllm_disagg/niah: make the accuracy check a real pass/fail gate, jitter needle placement benchmark_niah.py printed a score table and always exited 0, so an accuracy regression could not fail a run -- it had to be spotted by eye. It now exits non-zero: 3 when a context length produced no usable result, 4 when mean retrieval is below NIAH_MIN_SCORE (default 8.0/10), 0 otherwise, with an explicit "NIAH VERDICT: PASS/FAIL" line and a per-row marker on the offending lengths. Needle placement was fully deterministic: every needle sat at exactly (i+1)*step, the same offsets for every seed, so the documented NIAH_SEEDS sweep was re-testing one identical layout and its mean/min/max spread across seeds could not show what it claimed to. Needles are now jittered within their slot (seed 0 keeps the exact old placement, so existing baselines stay comparable). Adds benchmark_niah_perf.sh: the same NIAH retrieval scoring driven at concurrency so accuracy is checked under load rather than only single-stream. Known gap, not addressed here: needles are placed at (i+1)*step for i in 0..len(ANIMALS)-1, so nothing is ever placed in the final step-width of the context. The last chunk -- often the interesting one for a long-context regression -- is never probed. --- scripts/vllm_dissag/benchmark_niah.py | 52 ++++++++++++-- scripts/vllm_dissag/benchmark_niah_perf.sh | 79 ++++++++++++++++++++++ 2 files changed, 125 insertions(+), 6 deletions(-) create mode 100755 scripts/vllm_dissag/benchmark_niah_perf.sh diff --git a/scripts/vllm_dissag/benchmark_niah.py b/scripts/vllm_dissag/benchmark_niah.py index 43d81850..9fd13c1d 100755 --- a/scripts/vllm_dissag/benchmark_niah.py +++ b/scripts/vllm_dissag/benchmark_niah.py @@ -8,8 +8,13 @@ # NIAH_MODEL model name/tag the server serves (required — the served path) # NIAH_WORDS comma list of context sizes in words (default 2000,8000,20000,35000) # NIAH_MAXTOK max_tokens for the answer (default 2048) -# NIAH_SEEDS comma list of needle-layout seeds (default 0,1,2); summary reports -# mean/min/max across seeds to separate real accuracy from variance +# NIAH_SEEDS comma list of layout seeds (default 0,1,2); summary reports +# mean/min/max across seeds to separate real accuracy from variance. +# Each seed varies BOTH the filler words and the needle offsets +# (seed 0 = the historical evenly-spaced layout, kept for +# comparability). Varying offsets is what gives coverage of +# position-dependent failures -- lost-in-the-middle, RoPE +# extrapolation, and corruption at prefill CHUNK BOUNDARIES. # NIAH_TIMEOUT per-request timeout seconds (default 1800) # NIAH_WARMUP 1 (default) = send one throwaway request per context length BEFORE # scoring, so the first-hit JIT/kernel-autotune compile happens outside @@ -53,7 +58,19 @@ def make_haystack(n_words, seed=0): words = [rng.choice(FILLER) for _ in range(n_words)] step = max(n_words // (len(ANIMALS) + 1), 1) for i, animal in enumerate(ANIMALS): - words[min((i + 1) * step, len(words) - 1)] = animal + base = (i + 1) * step + # Needle offsets must vary with the seed, or the seeds only reshuffle + # filler and every seed probes the SAME 10 positions. That hides exactly + # the failure mode this stack is suspected of: prefill splits at + # max_num_batched_tokens, and a needle pinned on a bad chunk boundary is + # then missed identically by all seeds and reported as a confident mean. + # + # seed 0 reproduces the historical evenly-spaced layout bit-for-bit so + # earlier seed-0 results stay comparable. Jitter is < step and slot i+1 + # starts at (i+2)*step, so needles can never collide or reorder; all 10 + # always survive. + off = 0 if seed == 0 else rng.randrange(step) + words[min(base + off, len(words) - 1)] = animal return " ".join(words) @@ -123,19 +140,42 @@ def main(): for n in WORDS: results[n] = [run(n, s) for s in SEEDS] print("=== NIAH summary (mean/min/max across %d seed(s)) ===" % len(SEEDS), flush=True) + # EXIT-CODE CONTRACT (benchmark_niah_perf.sh:64 gates the perf phase on this): + # 3 = some length had NO usable result at all (dead server / timeout) + # 4 = every length scored, but some mean < NIAH_MIN_SCORE (retrieval broken) + # 0 = all lengths scored >= NIAH_MIN_SCORE + # Previously main() always fell through to an implicit `return None` -> exit 0, + # so the gate could never fire and a dead server produced "0.00 tok/s SUCCESS". + min_score = float(os.environ.get("NIAH_MIN_SCORE", "8.0")) + any_noresult = False + any_lowscore = False for n in WORDS: scored = results[n] vals = [v for v in scored if v is not None] n_to = sum(1 for v in scored if v is None) # timeouts/errors, excluded from mean if not vals: + any_noresult = True print(" words=%6d NO-RESULT (%d/%d timed out or errored — likely cold compile; " "raise NIAH_TIMEOUT or keep NIAH_WARMUP=1)" % (n, n_to, len(scored)), flush=True) continue mean = sum(vals) / len(vals) extra = (" [%d timeout/err excluded]" % n_to) if n_to else "" - print(" words=%6d mean=%.1f/10 min=%d max=%d (n=%d)%s" - % (n, mean, min(vals), max(vals), len(vals), extra), flush=True) + verdict = "" if mean >= min_score else " <-- BELOW NIAH_MIN_SCORE=%.1f" % min_score + if mean < min_score: + any_lowscore = True + print(" words=%6d mean=%.1f/10 min=%d max=%d (n=%d)%s%s" + % (n, mean, min(vals), max(vals), len(vals), extra, verdict), flush=True) + if any_noresult: + print("=== NIAH VERDICT: FAIL (no usable result at one or more lengths) ===", + flush=True) + return 3 + if any_lowscore: + print("=== NIAH VERDICT: FAIL (retrieval below NIAH_MIN_SCORE=%.1f) ===" % min_score, + flush=True) + return 4 + print("=== NIAH VERDICT: PASS (all lengths >= %.1f/10) ===" % min_score, flush=True) + return 0 if __name__ == "__main__": - main() + sys.exit(main()) diff --git a/scripts/vllm_dissag/benchmark_niah_perf.sh b/scripts/vllm_dissag/benchmark_niah_perf.sh new file mode 100755 index 00000000..fe31ab76 --- /dev/null +++ b/scripts/vllm_dissag/benchmark_niah_perf.sh @@ -0,0 +1,79 @@ +#!/bin/bash +# NIAH accuracy first, then the ISL/OSL perf sweep -- both against the SAME live +# server, in that order. Selected with BENCHMARK_SCRIPT=niah_perf. +# +# WHY COMBINED: engine boot is ~20 min (weights + MoRI symmetric heap + DSA indexer +# JIT + cudagraph capture). Two launches pay that twice AND measure accuracy and +# throughput on two different boots, so a perf regression could never be attributed. +# +# WHY GATED: if NIAH cannot retrieve a needle, the 28k/1k throughput number is +# measuring a broken server. Reporting it would be worse than reporting nothing, so +# the perf half is skipped unless NIAH passes. Override with NIAH_GATE=0. +set -u +timestamp=$(date "+%Y%m%d_%H%M%S") +BENCHMARK_PORT="${BENCHMARK_PORT:-30000}" +DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +LOGDIR="/run_logs/${SLURM_JOB_ID}" +NIAH_LOG="${LOGDIR}/niah_${SLURM_JOB_ID}_${timestamp}_xP${xP}_yD${yD}_${MODEL_NAME}.log" + +echo "==== PHASE 1/2: NIAH long-context retrieval ====" +echo "port=${BENCHMARK_PORT} model=${MODEL_PATH} sizes=${NIAH_WORDS:-2000,8000,20000,35000}" + +# Poll until the router serves. On a fresh boot it registers a few seconds after the +# workers report ready, so a blind sleep either wastes time or starts too early. +# Poll with a REAL completion, not /v1/models. That endpoint is wrong in BOTH +# directions on this router: it returned 200 while every request 503'd (the 11:55 +# run: 20 perf cells of 0.00 tok/s recorded as SUCCESS), and it returns 503 +# "No prefill servers available" right now while completions succeed end to end. +# Its listing path and its forwarding path consult different state. A completion +# is the only probe that confirms router + backend + P->D KV transfer together. +_ready=0 +for _i in $(seq 1 60); do + _probe=$(curl -s --max-time 120 "http://127.0.0.1:${BENCHMARK_PORT}/v1/completions" \ + -H 'Content-Type: application/json' \ + -d "{\"model\":\"${MODEL_PATH}\",\"prompt\":\"The CEO of AMD is\",\"max_tokens\":8,\"temperature\":0}" 2>&1) + if echo "$_probe" | grep -q '"text"'; then + _ready=1; echo "[niah] end-to-end probe OK after ~$((_i*10))s"; break + fi + sleep 10 +done +if [ "$_ready" != 1 ]; then + echo "[niah] FATAL: no successful completion in 600s -- not benchmarking a dead server." + echo "[niah] last response: ${_probe:0:400}" + exit 1 +fi +echo "[niah] probe: $(echo "$_probe" | head -c 200)" + + +# Invoked directly, NOT via benchmark_niah.sh: that wrapper ends in `| tee`, whose +# exit status is tee's (always 0), so a FAILING NIAH would report success and the +# gate below would be inert. Pipe to tee here but recover the real status. +set -o pipefail +NIAH_URL="http://127.0.0.1:${BENCHMARK_PORT}/v1/chat/completions" \ +NIAH_MODEL="${MODEL_PATH}" \ +NIAH_WORDS="${NIAH_WORDS:-2000,8000,20000,35000}" \ +NIAH_MAXTOK="${NIAH_MAXTOK:-2048}" \ +NIAH_TIMEOUT="${NIAH_TIMEOUT:-1800}" \ +NIAH_WARMUP="${NIAH_WARMUP:-1}" \ + python3 "${DIR}/benchmark_niah.py" 2>&1 | tee -a "${NIAH_LOG}" +_niah_rc=${PIPESTATUS[0]} +set +o pipefail + +echo "==== NIAH exit=${_niah_rc} -> ${NIAH_LOG} ====" + +if [ "${NIAH_GATE:-1}" = "1" ] && [ "${_niah_rc}" != "0" ]; then + echo "==== PHASE 2/2 SKIPPED: NIAH did not pass (exit ${_niah_rc}) ====" + echo "A 28k/1k throughput number from a server that cannot retrieve a needle is" + echo "not a result. Fix accuracy first, or re-run with NIAH_GATE=0 to force perf." + exit "${_niah_rc}" +fi + +echo "==== PHASE 2/2: ISL/OSL perf sweep ====" +# Defaults are the user's spec: 28k in / 1k out at concurrency 16, 32, 64. +# SHAPE_WARMUP=1 matters here -- benchmark_xPyD.sh's own comment records that without +# it the first measured cell absorbs residual JIT and reports 302 ms TPOT against an +# ~89 ms steady state, i.e. a 3.4x error on the headline metric. +export BENCHMARK_COMBINATIONS="${PERF_COMBINATIONS:-28672/1024}" +export BENCHMARK_CON="${PERF_CON:-16 32 64}" +export SHAPE_WARMUP="${SHAPE_WARMUP:-1}" +bash "${DIR}/benchmark_xPyD.sh" From 801e536164ef316c0a5c9b2679e5e2e8cb4a0d0f Mon Sep 17 00:00:00 2001 From: raviguptaamd Date: Sun, 16 Aug 2026 16:24:09 -0700 Subject: [PATCH 15/41] vllm_disagg/parse_to_csv: stop dropping the first benchmark cell; fix backend and arch tags Three reporting bugs. None affected the served workload, all affected the numbers that get filed. 1. The first (lowest-concurrency) cell of every run was silently discarded. The parse loop guarded its config lookup with `if i > 1`, which reads like a bounds check but is not one: enumerate starts at 1, so i-1 is 0, a valid index -- sections[0] is the preamble that holds the first cell's [RUNNING] line. A 21-cell sweep produced 20 rows. Regression-checked with a synthetic 3-cell log. 2. Rows with no parseable config or throughput were dropped silently, which made this parser bug look like a benchmark-loop bug. They now print a warning naming the result index. 3. Two hardcoded tags: backend fell back to 'nixl'. The real selector is CONNECTOR={rixl|moriio}, resolved and exported by run_xPyD_models.slurm:188-237; RUN_MORI/RUN_DEEPEP are the legacy flags mapped onto it there. So every run driven the current way -- including all the MoRIIO ones -- was filed as nixl. Now reads CONNECTOR, defaulting to moriio, with the legacy flags still taking precedence. gpu_architecture was hardcoded 'gfx942', so MI355X runs were filed as MI300X. Now from GPU_ARCHITECTURE, defaulting to gfx950 (this recipe's target). It is not introspected because this parser usually runs on the head node, which has no GPU. --- scripts/vllm_dissag/parse_to_csv.py | 71 ++++++++++++++++++----------- 1 file changed, 45 insertions(+), 26 deletions(-) diff --git a/scripts/vllm_dissag/parse_to_csv.py b/scripts/vllm_dissag/parse_to_csv.py index e772e394..770a2dfa 100644 --- a/scripts/vllm_dissag/parse_to_csv.py +++ b/scripts/vllm_dissag/parse_to_csv.py @@ -44,35 +44,45 @@ def parse_benchmark_log(log_file: str) -> Dict[Tuple[int, int, int], Dict]: current_concurrency = None for i, section in enumerate(sections[1:], 1): # Skip first empty section - # Look for configuration in previous sections (from [RUNNING] line) - if i > 1: - prev_section = sections[i-1] - - # vllm format: [RUNNING] prompts isl osl con - config_match = re.search( - r'\[RUNNING\]\s+prompts\s+\d+\s+isl\s+(\d+)\s+osl\s+(\d+)\s+con\s+(\d+)', - prev_section - ) - # Fallback: extract from Namespace(...) in vllm bench serve output - if not config_match: - isl_m = re.search(r'random_input_len=(\d+)', prev_section) - osl_m = re.search(r'random_output_len=(\d+)', prev_section) - con_m = re.search(r'max_concurrency=(\d+)', prev_section) - if isl_m and osl_m and con_m: - config_match = type('Match', (), { - 'group': lambda self, n: [None, isl_m.group(1), osl_m.group(1), con_m.group(1)][n] - })() - if config_match: - current_input_seq_len = int(config_match.group(1)) - current_output_seq_len = int(config_match.group(2)) - current_concurrency = int(config_match.group(3)) + # Config for result i lives in sections[i-1]. That is TRUE FOR i==1 TOO: + # sections[0] is the preamble before the first result banner, and it holds + # the FIRST cell's [RUNNING] line. The old code guarded this with `if i > 1` + # -- which reads like a bounds check but is not one, since enumerate starts + # at 1 so i-1 is 0, a valid index. The effect was that every run silently + # lost its first (lowest-concurrency) cell: 21 [RUNNING] cells -> 20 rows. + prev_section = sections[i-1] + + # vllm format: [RUNNING] prompts isl osl con + config_match = re.search( + r'\[RUNNING\]\s+prompts\s+\d+\s+isl\s+(\d+)\s+osl\s+(\d+)\s+con\s+(\d+)', + prev_section + ) + # Fallback: extract from Namespace(...) in vllm bench serve output + if not config_match: + isl_m = re.search(r'random_input_len=(\d+)', prev_section) + osl_m = re.search(r'random_output_len=(\d+)', prev_section) + con_m = re.search(r'max_concurrency=(\d+)', prev_section) + if isl_m and osl_m and con_m: + config_match = type('Match', (), { + 'group': lambda self, n: [None, isl_m.group(1), osl_m.group(1), con_m.group(1)][n] + })() + if config_match: + current_input_seq_len = int(config_match.group(1)) + current_output_seq_len = int(config_match.group(2)) + current_concurrency = int(config_match.group(3)) # Extract Total token throughput (tok/s) from benchmark result section throughput_match = re.search(r'Total token throughput \(tok/s\):\s+([\d.]+)', section) throughput = float(throughput_match.group(1)) if throughput_match else None # Only process if we have a valid configuration from [RUNNING] line and throughput - if current_input_seq_len and current_output_seq_len and current_concurrency and throughput is not None: + if not (current_input_seq_len and current_output_seq_len and current_concurrency + and throughput is not None): + # Never drop a measured result without saying so. The old silent drop + # made a parser bug look like a benchmark-loop bug. + print("Warning: benchmark result #%d has no parseable config/throughput " + "-- row dropped" % i) + else: config_key = (current_input_seq_len, current_output_seq_len, current_concurrency) results[config_key]['concurrency'] = current_concurrency @@ -121,13 +131,17 @@ def _get_run_metadata(pipeline: str = "vllm"): run_deepep = os.environ.get('RUN_DEEPEP', '0') gpus = os.environ.get('GPUS_PER_NODE', '8') - # Determine backend tag + # Determine backend tag. RUN_MORI/RUN_DEEPEP are the LEGACY flags; the current + # axes are CONNECTOR={rixl|moriio} x WIDE_EP x EP_BACKEND, resolved and exported + # by run_xPyD_models.slurm:188-237 (legacy flags are mapped onto them there). + # The fallback used to be hardcoded 'nixl', which mislabelled every run driven by + # CONNECTOR= rather than the legacy flags -- including all of the MoRIIO ones. if run_mori == '1': backend = 'mori' elif run_deepep == '1': backend = 'deepep' else: - backend = 'nixl' + backend = os.environ.get('CONNECTOR', 'moriio').lower() return { 'pipeline': pipeline, @@ -139,7 +153,12 @@ def _get_run_metadata(pipeline: str = "vllm"): 'docker_image': os.environ.get('DOCKER_IMAGE_NAME', ''), 'machine_name': os.environ.get('SLURM_JOB_NODELIST', ''), 'launcher': 'slurm_multi', - 'gpu_architecture': 'gfx942', + # Was hardcoded 'gfx942' (MI300X), so every MI355X run was filed under the + # wrong arch in perf.csv. This parser usually runs on the head/login node, + # which has no GPU, so it cannot reliably introspect the arch -- take it from + # the environment, defaulting to the arch this recipe targets. + # gfx942 = MI300X/MI325X, gfx950 = MI355X. + 'gpu_architecture': os.environ.get('GPU_ARCHITECTURE', 'gfx950'), } From bef784f3e851204516f5b81c9e70356afd3cd8e6 Mon Sep 17 00:00:00 2001 From: raviguptaamd Date: Sun, 16 Aug 2026 16:24:22 -0700 Subject: [PATCH 16/41] vllm_disagg: document the GLM-5.2 MI355X recipe Adds GLM52_MI355X.md (invocation, measured numbers, the two load-bearing config values, the cluster-specific traps) and extends skills_vllm_disagg.md with the perf-tuning findings from this bring-up. Records the results honestly, including what is not solved: TPOT 33.7 / 36.4 / 40.2 ms at concurrency 16 / 32 / 64 -- target 50 ms, met. Prefill is 6,135-8,359 tok/s/rank against a 34,000 target, ~4-5x short, and is now the binding constraint. Profiled at 8192 tokens: MLA sparse attention ~82%, all GEMM 17.8%, of which MoE 2.9%. At 80K ISL the fitted single-request TTFT is ~13.7 s on one rank before any concurrency, against a <7 s SLO. TTFT at concurrency is drain time (458,752 tok / 39,340 tok/s = 11.66 s, an identity that reproduces the measurement), so more prefill nodes fix concurrent TTFT but do nothing for the single-request number. That needs tuned MLA sparse-attention kernels for gfx950 or intra-request sharding. Also records the levers closed BY MEASUREMENT so they are not re-litigated: MoE and AITER tuning (2.9% of prefill), indexer sub-chunking (<=2.5%), max_num_batched_tokens (<=4%), RDMA QP/worker count (-1.9%, inside +-0.4% noise), FP8BMM (~0.7%), block_size 1->64 (<2%). And why PCP is not the near-term answer: config/parallel.py:528 is DP-incompatible, rocm_aiter_mla_sparse.py has no PCP plumbing at all, and platforms/rocm.py:894 force-sets cudagraph_mode=PIECEWISE, which would silently undo the TPOT fix. Two constraints called out because both fail in a way that looks like success: use_inductor_graph_partition must stay ON and no bare --enforce-eager (boot and warmup pass, the first real MLA decode dies), and the MXFP4 recipe has never been booted. --- scripts/vllm_dissag/GLM52_MI355X.md | 137 +++++++++++++++ scripts/vllm_dissag/skills_vllm_disagg.md | 204 +++++++++++++++++++++- 2 files changed, 340 insertions(+), 1 deletion(-) create mode 100644 scripts/vllm_dissag/GLM52_MI355X.md diff --git a/scripts/vllm_dissag/GLM52_MI355X.md b/scripts/vllm_dissag/GLM52_MI355X.md new file mode 100644 index 00000000..cec72ac1 --- /dev/null +++ b/scripts/vllm_dissag/GLM52_MI355X.md @@ -0,0 +1,137 @@ +# GLM-5.2 on MI355X (gfx950) + AMD AI NIC — 1P/1D EP8 + +Recipe notes for `GLM-5.2-FP8` and `GLM-5.2-MXFP4` in `models.yaml`, disaggregated +prefill/decode over MoRI-IO. Sibling of the GLM-5.1 wideEP recipe; this file records only +what is specific to GLM-5.2 and to gfx950 + ionic NICs. + +Measured on AAC, 2 nodes x 8 x MI355X (gpu-44 / gpu-45), ROCm 7.2.3. + +--- + +## 1. Invocation + +```bash +export DOCKER_IMAGE_NAME=/vllm-disagg:glm52-gfx950-ionic +export MODEL_DIR=/shared/data/amd_int/models # holds the HF snapshots +export MODEL_NAME=GLM-5.2-FP8 # or GLM-5.2-MXFP4 +export CONNECTOR=moriio WIDE_EP=1 EP_BACKEND=mori +export CONNECTOR_ENV_FILE=$PWD/connectors/moriio.env.aac # AAC/ionic platform env +export xP=1 yD=1 +export BENCHMARK_COMBINATIONS=28672/1024 +export BENCHMARK_CON="16 32 64" + +srun --jobid $SLURM_JOB_ID --overlap -N2 -w "$NODES" --ntasks-per-node=1 \ + bash ./run_xPyD_models.slurm +``` + +`CONNECTOR_ENV_FILE` is the only cluster-specific knob. On a different cluster, copy +`connectors/moriio.env.aac`, then re-check the three things that are physically per-site: +the RDMA device list, `MORI_SOCKET_IFNAME`, and `FABRIC_SUBNET` (§4). + +--- + +## 2. Measured + +FP8, ISL/OSL 28,672/1,024, EP8, 1P/1D: + +| concurrency | TPOT (ms) | note | +|---|---|---| +| 16 | 33.7 | | +| 32 | 36.4 | | +| 64 | 40.2 | | + +TPOT target for AIMODELS-1198 is 50 ms avg — **met with margin at all three points.** + +Prefill throughput measures **6,135–8,359 tok/s/rank** against a 34,000 tok/s/rank target, +i.e. **~4–5x short**, and that is the binding constraint. See §5. + +`GLM-5.2-MXFP4` is **config-only and has never been booted.** Its recipe is a copy of the +FP8 one with the two values that provably do not transfer re-derived (§3). Treat its first +run as a bring-up, not a benchmark. + +--- + +## 3. The two values that carry the recipe + +**`DECODE_CUDAGRAPH_MODE: FULL_AND_PIECEWISE`** — the single change that fixed TPOT, +104 ms -> 34–40 ms. `PIECEWISE` splits the decode graph on the 3 DSA ops x 78 layers, +~234 launch boundaries per step. That cost is *batch-independent*, so it presents as a +latency floor that does not move when you change concurrency, batch size, or the fabric — +which is exactly why it survived so much tuning before being found. + +> **Hard co-requisite:** `use_inductor_graph_partition` must stay ON, and never pass a bare +> `--enforce-eager`. With the partitioner off, boot and warmup both PASS and the first real +> MLA decode dies. + +**`prefill.dp: --gpu-memory-utilization 0.72`** — prefill only. The MLA chunked-prefill +workspace is sized from a hardcoded 64k-token clamp +(`determine_chunked_prefill_workspace_size`, `mla_attention.py:1935`), *not* from +`max_num_batched_tokens`. It therefore allocates `65536*64*(192+256)*2` = **3.50 GiB** on +top of the KV pool regardless of how small the scheduler batch is — and lazily, on the +first long prefill, i.e. after boot and warmup have both reported healthy. At 0.80 this +OOM'd in `mla_attention.py:739`. + +MXFP4 inherits the same 0.72 clamp. Its weights are ~42% smaller and it runs +`GPU_MEMORY_UTILIZATION 0.90`, so it has *less* free HBM at the moment those 3.50 GiB land, +not more. That value is inferred from the FP8 failure, not measured. + +--- + +## 4. Cluster-specific: what will bite you elsewhere + +**`FABRIC_SUBNET`.** The launcher defaults to `10.158.` (the MI300X cluster). AAC is +`10.2.80.`, set in `moriio.env.aac`. Wrong value => `MASTER_ADDR` is advertised as an +address the peer cannot reach => the run hangs at `Waiting for nodes` forever, with no +error. This one key is host-side, so the `.env` loader has an explicit host-export +allow-list for it (`run_xPyD_models.slurm:255-263`); every other key in that file is +container-only. + +**Pick the interface by testing TCP, not ICMP.** On AAC the public/default-route NIC +answers ping but has TCP firewalled node-to-node. `/dev/tcp//22` is the honest test. + +**8 ionic rails**, `rocep{9,25,105,121,137,153,233,249}s0`, with `MORI_IB_GID_INDEX=1`. +Control/bootstrap traffic rides `MORI_SOCKET_IFNAME=enp193s0f1np1` (mgmt); bulk EP and KV +traffic ride the rails. + +**Two patchers are opt-in OFF on this image** (`GLM_PERSIST_GATE`, `GLM_DSA_SENTINEL_FIX`) +because both are actively harmful here: a C++ abort at `asm_mla.cu:945`, and a +`hipErrorIllegalAddress` from the sentinel 0 -> -1 change respectively. Both are correct on +older images — the gate is by image, not by model. + +--- + +## 5. What this recipe does **not** fix + +Prefill. Profiled at 8,192 tokens (~1,314 ms/step): + +| component | share | +|---|---| +| MLA sparse attention (indexer + core + kv gather) | **~82%** | +| all GEMM | 17.8% | +| — of which MoE | **2.9%** | + +Two structural facts, each verified two independent ways: + +1. **Attention is linear in context, not quadratic.** `index_topk=2048` caps keys per + query, so the attention core is *flat* at 0.791 ms across L = 7,168 / 14,336 / 28,672 / + 57,344. DSA works as designed — good news for the 256K/1M context targets. +2. **Only the DSA indexer grows** (it scores all L keys to pick 2,048): 1.00 / 1.66 / 3.07 + / 5.78x over that same 8x range. End-to-end this reads as + `TTFT ~= 200ms + ~160us/token`. + +**Levers closed by measurement — do not re-litigate these without new evidence:** MoE/AITER +tuning (2.9% of prefill), indexer sub-chunking (<=2.5%), `max_num_batched_tokens` (<=4%), +RDMA QP/worker count (-1.9%, within +-0.4% noise), FP8BMM (~0.7%), `block_size` 1 -> 64 +(<2%). + +**Consequence for the SLO.** At 80K ISL the fitted single-request TTFT is **~13.7 s on one +rank, before any concurrency**, against a <7 s target. TTFT at concurrency is drain time — +`458,752 tok / 39,340 tok/s = 11.66 s`, an identity that reproduces the measurement — so +adding prefill nodes (xP) fixes *concurrent* TTFT but does **nothing** for the +single-request number. The only structural fixes are (a) fused/tuned MLA sparse-attention +kernels for gfx950 or (b) intra-request sharding. + +**PCP is blocked three ways**, so it is not the near-term answer: `config/parallel.py:528` +is incompatible with DP; `rocm_aiter_mla_sparse.py` has zero PCP plumbing (uniquely among +the attention backends); and `platforms/rocm.py:894` force-sets `cudagraph_mode=PIECEWISE`, +which would silently undo the TPOT fix in §3. diff --git a/scripts/vllm_dissag/skills_vllm_disagg.md b/scripts/vllm_dissag/skills_vllm_disagg.md index ee778064..85bfe0f6 100644 --- a/scripts/vllm_dissag/skills_vllm_disagg.md +++ b/scripts/vllm_dissag/skills_vllm_disagg.md @@ -207,7 +207,12 @@ existing `PREFILL_MORI_BACKEND` / `DECODE_MORI_BACKEND`). Verify it landed by re server process, not the container shell, so `docker exec env` shows nothing. ### 5.7 RDMA fabric -- GID index **3** = RoCEv2 IPv4 (check `show_gids` / `sysfs .../gid_attrs/types`). +- GID index is **site-dependent — enumerate it, do not copy it**. `show_gids` / + `sysfs .../gid_attrs/types`. Index 3 = RoCEv2 IPv4 on the clusters the stock + `moriio.env` was written for, but **on AAC (Pensando/Ionic) index 3 does NOT exist** + (`type=none`) and index 0 is link-local `fe80::` (not routable cross-node). AAC must + run `MORI_IB_GID_INDEX=1`; the stock `GID_INDEX=3` selects a nonexistent entry. + See `run_niah_glm52.sh:47-49`. - Restrict `MORI_RDMA_DEVICES` / `NCCL_IB_HCA` to the 8 GPU-local NICs; leave the mgmt NICs out or QPs try to form over a non-routable fabric -> `ibverbs.cpp:189 Connection timed out`. - NCCL/GLOO control sockets on `eth0` (mgmt); KV data on the RDMA NICs. @@ -243,3 +248,200 @@ serially — parallel `docker pull` across 7 nodes is far faster and removes the 2. **Perf:** the MoE all-to-all buffer is sized from a chunked-prefill *scheduler* knob, so decode moves an 8192-wide buffer every step. Lower `--max-num-batched-tokens` on the decode role: **302 ms -> 88 ms TPOT**. + +--- + +## 8. Perf tuning: the decode TPOT budget (MI355X gfx950 + Pensando/Ionic AINIC) + +Added 2026-08-16 from job 5642 (GLM-5.2-FP8, 1P/1D EP8, AAC mi355-gpu-45/46). This +section is the standing perf-tuning reference: **measure the budget before tuning a +kernel**. Everything here is measured on the real deployment shape, not theorised. + +### 8.1 Measure where the step actually goes BEFORE tuning anything + +Microbenchmarks on idle MI355X GPUs, in the serving image, at the real EP8 decode +shape (78 layers, decode batch ~32, 28k context): + +| component (x78 layers) | ms/step | +|---|---| +| fused MoE (untuned, real EP shape) | 24-38 | +| dense GEMMs (q_proj, kv_a, kv_b, o_proj, indexer_q) | ~9.7 | +| DSA sparse indexer scoring (28k ctx, top-2048) | ~6.4 | +| **compute subtotal** | **~40** | +| **measured TPOT** | **~104** | +| **unaccounted: all2all + per-layer launch overhead** | **~64 (62%)** | + +Per-op detail (fused_moe MD=6144 ID=2048 E=32 TOPK=8; x78 layers): + + tokens= 8 0.4873 ms/call -> 38.01 ms q_proj 6144->16384 0.0398 -> 3.10 + tokens= 16 0.3203 ms/call -> 24.98 ms kv_a 6144->576 0.0133 -> 1.04 + tokens= 32 0.3039 ms/call -> 23.71 ms kv_b 512->28672 0.0131 -> 1.02 + tokens= 64 0.4646 ms/call -> 36.24 ms o_proj 16384->6144 0.0425 -> 3.32 + indexer_q 6144->4096 0.0162 -> 1.26 + DSA indexer tok=32 ctx=28672: 0.0818 ms/call -> 6.38 ms + +**Consequence: MoE kernel tuning cannot reach a 60 ms TPOT SLO.** Even a *perfect* +MoE kernel saves <=30 ms; realistic 10-20% gains yield 3-7 ms against a 44 ms deficit. +Do not spend allocation time on AITER MoE tuning until the non-compute term is closed. + +Three independent signatures agree the bottleneck is fixed-cost dispatch, not batch work: +- TPOT flat at 103.8 / 104.3 / 104.9 ms across con=16/32/64 (**4x** concurrency). +- `max_num_batched_tokens` curve flat below 2048 (512 -> 87.9 ms, 2048 -> 88.0 ms). +- The ROCm MI300X + **CX7** blog hits the same ~89 ms on completely different NICs. + +### 8.2 Read AITER's real fused-MoE lookup key — never guess it from source + +AITER logs its own lookup decision. One grep gives the exact runtime key: + + grep -o "\[fused_moe\] using .* for (.*)" decode_NODE1.log | sort -u + +Live GLM-5.2-FP8 EP8 decode printed exactly one key, always "default" (= untuned): + + [fused_moe] using 1stage default for + ('gfx950', 256, 16384, 6144, 2048, 32, 7, Silu, 'torch.bfloat16', + 'torch.float8_e4m3fn', 'torch.float8_e4m3fn', 'QuantType.per_1x128', True, False) + key order: (gfx, cu_num, token, model_dim, inter_dim, expert, topk, ...) + +**EP and TP shard the MoE in opposite dimensions** — this is the trap: +- **EP**: experts shard (256 routed / 8 ranks = **32**); `inter_dim` does **NOT** shard (**2048**). +- **TP**: `inter_dim` shards (2048/8 = **256**); expert count stays whole (256 + 1 fake = **257**). + +`a8w8_blockscale_tuned_fmoe_glm5.csv` ships `(6144, 256, 257, topk 9)` — a **TP** table. +Our EP shape `(6144, 2048, 32)` has **zero** tuned rows in *any* shipped table. All +`model_dim=6144` coverage that exists at all: + + a8w8_blockscale_tuned_fmoe_glm5.csv inter 256 expert 257 topk 9 + minimax_m3_fp4 / mxfp8 inter 384/768 expert 129 topk 5 + tuned_fmoe.csv inter 4096 expert 8 topk 2 + +On a miss the kernel shape genuinely changes: `block_m = 64 if token > 32 else 16`, +whereas the tuned table specifies `block_m=32` at tokens 32/64. Two env handles for +zero-patch differential tests: `AITER_BYPASS_TUNE_CONFIG=1` (force default path) and +`AITER_ONLINE_TUNE=1` (gate the online tuner; only this appends to `untuned_fmoe.csv`). + +### 8.3 Upstream bug: `topk -= int(is_ep)` makes every EP row unreachable + +`aiter/fused_moe.py:1261` decrements `topk` before the table lookup, but every shipped +table stores the EP row with the fake expert slot in **both** `expert` and `topk` +(kimik2 384/8 + 385/9; qwen3.5 512/10 + 513/11; dsv3 257/9). So the runtime probes +e.g. `(257, 8)` — a row that exists for no model. Proven by AITER's own log line: +`topk=9` -> HIT (`kernelName1=_ZN5aiter49fmoe_bf16_blockscaleFp8_g1u1_novs_silu_1tg_32x256E`); +`topk=8` and `topk=7` -> `"using 1stage default"` (MISS). + +**Real bug, but NOT our bottleneck** — our key misses on `expert` and `inter_dim` too, +and per §8.1 MoE is only ~25% of the step. Report it upstream; don't chase it for TPOT. + +### 8.4 CUDA graph mode is a first-class TPOT knob (top current candidate) + +`PIECEWISE` cuts the graph at every op in `splitting_ops`. For DSA that list contains +**three** hot ops — `vllm::unified_mla_attention_with_output`, `vllm::sparse_attn_indexer`, +`vllm::rocm_aiter_sparse_attn_indexer` — so a 78-layer decode step pays on the order of +**~234 launch boundaries**. That is a fixed per-step cost independent of batch size, +which is exactly the §8.1 signature. + +Verified preconditions for full-decode graphs on this stack: + +| check | result | where | +|---|---|---| +| `ROCMAiterMLASparseBackend` CG support | `AttentionCGSupport.UNIFORM_BATCH` | `rocm_aiter_mla_sparse.py:364` | +| DSA indexer backend CG support | `AttentionCGSupport.UNIFORM_BATCH` | `mla/indexer.py:462-467` | +| mori-specific CG guard? | **none** (only `deepep_high_throughput` is force-disabled) | `config/compilation.py` | +| host syncs in mori hot path? | **none** — the `.item()`/`.cpu()`/`synchronize()` calls are all in `get_dispatch_src_token_pos`, a debug helper, not `dispatch`/`combine` | `mori/ops/dispatch_combine.py:1380` | + +`UNIFORM_BATCH` is precisely the level that permits `FULL_DECODE_ONLY` / +`FULL_AND_PIECEWISE` (vLLM's own v1 default). We were overriding *down* to `PIECEWISE`. +The knob is already plumbed per-role end-to-end: +`models.yaml DECODE_CUDAGRAPH_MODE` -> `run_xPyD_models.slurm:725` -> `connectors/moriio.sh:307`. +No code change needed. + +**Simulate the resolution before spending a boot.** Call the real resolver in-image +instead of relaunching and hoping — `set_splitting_ops_for_v1(all2all_backend, dp)` then +`resolve_cudagraph_mode_and_sizes(min_cg_support=..., min_cg_attn_backend=..., max_num_reqs=...)`: + + FULL_DECODE_ONLY mori_low_latency -> FULL_DECODE_ONLY decode=FULL mixed=NONE + FULL_AND_PIECEWISE mori_low_latency -> FULL_AND_PIECEWISE decode=FULL mixed=PIECEWISE + PIECEWISE mori_low_latency -> PIECEWISE decode=PIECEWISE mixed=PIECEWISE + +Both full modes survive unchanged under mori. **Prefer `FULL_AND_PIECEWISE`**: identical +`decode=FULL`, but `FULL_DECODE_ONLY` sets `mixed=NONE` (mixed prefill-decode batches drop +to eager) whereas `FULL_AND_PIECEWISE` keeps the piecewise graphs those batches already +have — a strict superset of current behaviour. + +Two constraints a careless cudagraph change breaks: +- `use_inductor_graph_partition` must stay **on** — it keeps the STABLE-ABI + `concat_and_cache_mla` out of the compiled graph. Without it the first *real* MLA decode + dies with `RuntimeError: unknown parameter type` (boot and warmup pass via the fake path, + so it looks healthy until the first request). +- Never express "no graphs" as bare `--enforce-eager` — it drops `+quant_fp8` and hits an + AITER `dynamic_per_token_scaled_quant` signature mismatch at engine init. Use + `cudagraph_mode: NONE` **with** `+quant_fp8`. + +**Verification hook:** grep the decode log for `Capturing CUDA graphs (decode` — the +`"decode"` label at `gpu_model_runner.py:7057`. Under PIECEWISE only the +`(mixed prefill-decode, PIECEWISE)` line ever appears. + +### 8.5 Block size: what the source actually supports + +| backend | `get_supported_kernel_block_sizes()` | file:line | +|---|---|---| +| `ROCMAiterMLASparseBackend` (**loaded**) | **`[1, 64]`** | `rocm_aiter_mla_sparse.py:282` | +| `DeepseekV32IndexerBackend` (ROCm) | **`[1, 64]`** | `indexer.py:136` | +| `DeepseekV4IndexerBackend` | `[256]` | `indexer.py:176` | +| `AiterMLABackend` | `[MultipleOf(1)]` | `rocm_aiter_mla.py:140` | + +**16 and 32 are NOT supported; 64 IS.** We run `block_size=1`, which makes the block +table enormous: `max_len/block_size` int32 entries per request row, so +`1048576 / 1 = 1,048,576` entries/row — vs 16,384 at `bs=64`, and 1,024 at `bs=64` + +64k context. Two knobs (`block_size` 1->64, `max_model_len` 1M->64k) compound to a +**1024x** smaller block table. + +**Unresolved contradiction, resolve before changing:** `models.yaml:204` asserts DSA +*requires* block-size 1, while `tests/argv_assert.sh:10` asserts `--block-size 16` for +the prefill role, and the source above allows 64. Do not change this blind. + +### 8.6 Perf-tuning method rules (learned the hard way) + +1. **Budget first, tune second.** §8.1 saved a whole restart-and-tune cycle chasing a + <=7 ms prize against a 44 ms deficit. +2. **Batch-independent TPOT means fixed per-step cost.** If 4x concurrency doesn't move + TPOT, stop looking at kernels and look at launches/collectives. +3. **Re-baseline inside the same engine boot.** Job 5642 produced 104.94 ms TPOT at + con=64 and then **146.30 ms** for the identical config ~25 min later (13,133 vs + 10,177 tok/s total). Runs were sequential (log names are **UTC**, mtimes **CDT** — + convert before concluding overlap), watchers held, zero throttle/preempt/recompute + markers, GPUs idle at 45-50 C junction on `auto`. **Cause still unknown.** Never A/B + against a number from a previous boot. +4. **Grep the engine's own decisions** rather than reading source and inferring — AITER, + vLLM compilation config, and cudagraph capture all log what they actually chose. + +### 8.7 Head-to-head results (ISL/OSL 28k/1k, 1P/1D EP8, GLM-5.2-FP8) + +| run | c16 tok/s | c32 tok/s | c64 tok/s | c64 TPOT | +|---|---|---|---|---| +| 171621 (run 1) | 3,978 | 7,423 | 13,133 | 104.94 ms | +| 174152 (run 2) | 3,001 | 5,739 | 10,177 | **146.30 ms** | + +NIAH accuracy **PASSED** in both runs. Ticket SLO is TPOT **50 ms** avg / TTFT < 7 s. + +### 8.8 Ranked lever list (current) + +1. **Decode CUDA graph mode** `PIECEWISE` -> `FULL_AND_PIECEWISE` (§8.4) — holds the ~64 ms. +2. **`block_size` 1 -> 64** (§8.5) — legal per source; resolve the `models.yaml:204` claim first. +3. **`max_model_len` 1M -> 64k** — 16x smaller block-table row; the existing "do not cap" + note is an *accuracy* argument only, not a perf one. +4. ~~AITER MoE tuning~~ — **deprioritised by measurement** (<=7 ms realistic). +5. ~~decode `max_num_batched_tokens` below 2048~~ — **exhausted**, curve is flat. +6. ~~`MORI_IB_GID_INDEX` 1 -> 3~~ — **closed, not a lever**. On AAC index 3 does not + exist (§5.7); 1 is correct. And it could not have mattered: in 1P/1D the decode EP8 is + 8 GPUs on ONE node over Infinity Fabric, so the NICs carry only the prefill->decode KV + handoff — GID choice moves **TTFT, not TPOT**. +7. AITER / MoRI / vLLM commit deltas vs the blog pins (`AITER e03fa6040`, `MoRI 42e895472b08`). + +### 8.9 Suspects eliminated — do NOT re-chase + +MoRI-EP kernel mode (correct: high_throughput prefill / low_latency decode) · absent +`FLYDSL_GPU_ARCH=gfx000` · DP5 empty-CUDA-graph asymmetry (on all 8 workers) · untuned +a8w8 GEMM as a *differential* vs the blog · CX7-vs-Ionic transport (not in the decode +loop) · graph capture failure (PIECEWISE=9 on all 8 workers, 3.69 GiB) · mnbt below +2048 · the glm5 `topk` off-by-one as *our* miss cause · thermal throttling and +double-launch as the run-2 regression cause. From 8212dfac8581638fbad90ee6651f7e645fc81ea0 Mon Sep 17 00:00:00 2001 From: raviguptaamd Date: Sun, 16 Aug 2026 17:34:06 -0700 Subject: [PATCH 17/41] vllm_disagg: add customer-SLO benchmark (agentic 80K/200K, Poisson, goodput-gated) benchmark_xPyD.sh is a tuning sweep and cannot answer "do we meet the SLO". This adds a harness that can, and exits non-zero when we do not. Four differences from the sweep, each of which was producing a wrong answer: 1. --request-rate inf is a saturation test. It fires everything at once, so TTFT is just queue drain time -- at con=64/isl=28672 the measured TTFT is exactly 458,752 tok / 39,340 tok/s = 11.66 s, an identity rather than a property of the model. The customer's "<7 s avg TTFT" is about a *served* request, so this runs at a finite Poisson rate derived from Little's Law: rate = concurrency / SLO latency. That asks "can it serve the offered load AT the SLO", which is the actual question. 2. The sweep reports mean/max; the customer asked for p50/p95/p99. Now passes --percentile-metrics ttft,tpot,itl,e2el --metric-percentiles 50,95,99. 3. Pass/fail was a human reading a log. The SLO is now handed to vLLM directly via --goodput ttft:7000 tpot:50, so the fraction of traffic actually served acceptably is a number the server computes. A cell can have a passing mean TTFT while a third of its requests miss; goodput catches that and the mean does not. 4. --random-prefix-len 0 makes every request unique. The customer's use case is agentic and they flagged interest in long prefix caching -- an agent loop re-sends a large static context every turn. vLLM's RandomDataset builds the prefix once and shares it across requests, so PREFIX_FRAC (default 0.5) models that. 0.5 is a declared assumption, not a measurement: the customer gave no reuse ratio, so sweep it. Scenarios are the customer's own: 80K/1K (256K context) and 200K/1K (1M context), concurrency to 256/DP. The 200K concurrency list stops at 64 on purpose. GLM-5.2 MLA KV is 43.88 KiB/token (kv_lora_rank 512 + qk_rope 64, FP8, 78 layers), so 200K = 8.37 GiB/request against a ~930 GiB single-node pool -- concurrency 256 would need ~2,142 GiB and cannot fit on one decode node. Defaulting to a list that fits means a failure is a real failure and not a predictable OOM. The arithmetic is in the header so the ceiling can be recomputed rather than rediscovered. slo_report.py divides throughput by --dp-ranks because the customer's targets are per-rank and vLLM reports aggregate; not doing so overstates by the DP degree. It labels prefill tok/s as an end-to-end lower bound (total_input_tokens/duration, which includes decode time for the same requests) rather than passing it off as an isolated prefill-engine number. --- scripts/vllm_dissag/benchmark_customer_slo.sh | 213 ++++++++++ scripts/vllm_dissag/gen_workload.py | 387 ++++++++++++++++++ scripts/vllm_dissag/slo_report.py | 183 +++++++++ 3 files changed, 783 insertions(+) create mode 100755 scripts/vllm_dissag/benchmark_customer_slo.sh create mode 100644 scripts/vllm_dissag/gen_workload.py create mode 100755 scripts/vllm_dissag/slo_report.py diff --git a/scripts/vllm_dissag/benchmark_customer_slo.sh b/scripts/vllm_dissag/benchmark_customer_slo.sh new file mode 100755 index 00000000..b90db5bc --- /dev/null +++ b/scripts/vllm_dissag/benchmark_customer_slo.sh @@ -0,0 +1,213 @@ +#!/bin/bash +# Customer-facing SLO benchmark for GLM-5.2 / HY4 on MI355X. +# +# WHY THIS EXISTS, AND WHY IT IS NOT benchmark_xPyD.sh +# ---------------------------------------------------- +# benchmark_xPyD.sh is a SWEEP: it walks a grid of isl/osl x concurrency at +# --request-rate inf and reports the max throughput per cell. That is the right tool +# for tuning -- it answers "did my change help". +# +# It is the WRONG tool for answering "do we meet the customer's SLO", for four +# reasons, each of which this script fixes: +# +# 1. --request-rate inf is a saturation test, not a service test. It fires all +# requests at once and measures how fast the queue drains. Under that regime +# TTFT is dominated by queueing: at con=64/isl=28672 the measured TTFT is just +# 458,752 tok / 39,340 tok/s = 11.66 s -- an identity, not a property of the +# model. The customer's "<7 s avg TTFT" is a statement about a SERVED request, +# so it must be measured at a finite arrival rate. +# 2. The reported stat is the mean/max. The customer specified p50/p95/p99. Those +# are different questions and vLLM will only print percentiles if asked +# (--percentile-metrics / --metric-percentiles). +# 3. The pass/fail is left to a human reading a log. Here the SLO is expressed to +# vLLM directly via --goodput, so "how many requests actually met the SLO" is a +# number the server computes, and this script exits non-zero when it is missed. +# 4. --random-prefix-len 0 means every request is unique. The customer's use case +# is AGENTIC and they explicitly flagged interest in long prefix caching -- an +# agent loop re-sends a large, mostly-static context every turn. Measuring that +# with zero shared prefix understates the real system by whatever the prefix +# cache would have saved. See PREFIX_FRAC below. +# +# WORKLOAD, AS SPECIFIED BY THE CUSTOMER +# --------------------------------------- +# use case agentic workflow -> shared prefix, multi-turn +# model GLM-5.2 (HY4 family) -> FP8 and MXFP4 +# ISL/OSL 80K/1K (256K context) +# 200K/1K (1M context) +# concurrency max 256 per DP rank +# TTFT < 7 s avg +# TPOT < 50 ms avg (equivalently "> 20 tok/s" per the sheet) +# prefill 34,000 tok/s per rank (MI355X) +# decode 670 tok/s per rank (MI355X) +# framework vLLM only +# +# Note the sheet gives TTFT/TPOT as *averages* but asks for p50/p95/p99 in the +# latency-SLO row. We report all of them and gate on the average, which is the +# stated acceptance criterion; the percentiles are reported so the tail is visible +# rather than hidden inside a mean. +# +# CAPACITY REALITY CHECK -- READ BEFORE SETTING CONCURRENCY +# --------------------------------------------------------- +# GLM-5.2 MLA KV is kv_lora_rank(512) + qk_rope_head_dim(64) = 576 elem/token/layer, +# FP8, 78 layers = 43.88 KiB/token. Per request that is: +# 28,672 tok -> 1.20 GiB 80,000 tok -> 3.35 GiB 200,000 tok -> 8.37 GiB +# On one 8x MI355X node (288 GB/GPU) the usable KV pool is roughly 930 GiB (FP8, +# util 0.80) after weights (88 GiB/GPU) and the 3.5 GiB MLA chunked-prefill +# workspace. So the KV-bound concurrency ceiling on ONE decode node is about: +# ISL 28,672 -> ~777 ISL 80,000 -> ~279 ISL 200,000 -> ~111 +# Concurrency 256 at 80K needs ~857 GiB and fits. **Concurrency 256 at 200K needs +# ~2,142 GiB and does NOT fit on a single decode node** -- it needs 3 (or MXFP4, +# which has ~1,440 GiB of pool and still only reaches ~172). This script does not +# silently truncate: it runs what you ask and reports what happened, but the 200K +# row is defaulted to a concurrency list that fits so that a failed run means a real +# failure and not an OOM you could have predicted with arithmetic. +# +# USAGE +# MODEL_PATH=/models/GLM-5.2-FP8 BENCHMARK_PORT=8000 ./benchmark_customer_slo.sh +# +# Exit code 0 only if every scenario met both the TTFT and the TPOT SLO. + +set -uo pipefail + +MODEL_PATH="${MODEL_PATH:?set MODEL_PATH}" +BENCHMARK_PORT="${BENCHMARK_PORT:-8000}" +HOST="${BENCHMARK_HOST:-127.0.0.1}" +LOG="${LOG:-/run_logs/${SLURM_JOB_ID:-local}/customer_slo}" +mkdir -p "$(dirname "$LOG")" 2>/dev/null || true + +# --- SLO, from the customer sheet. Milliseconds, because --goodput wants ms. --- +SLO_TTFT_MS="${SLO_TTFT_MS:-7000}" +SLO_TPOT_MS="${SLO_TPOT_MS:-50}" +# Per-rank throughput targets (MI355X row of the sheet). +TGT_PREFILL_TOK_S="${TGT_PREFILL_TOK_S:-34000}" +TGT_DECODE_TOK_S="${TGT_DECODE_TOK_S:-670}" + +# --- Number of DP ranks, to convert aggregate throughput into per-rank. --- +# The sheet's targets are PER RANK and vLLM reports an AGGREGATE, so getting this +# wrong scales every throughput verdict by 8x in whichever direction hurts. +DP_RANKS="${DP_RANKS:-8}" + +# --- Agentic shared prefix ----------------------------------------------------- +# Fraction of ISL that is a shared, cacheable system/tool preamble. vLLM's +# RandomDataset builds the prefix ONCE and prepends the same tokens to every +# request (datasets.py: "Generate prefix once"), so this is a genuine shared +# prefix that the prefix cache can hit -- not per-request random padding. +# 0.0 reproduces the old zero-sharing behaviour. 0.5 is a deliberate, declared +# assumption: the customer said "agentic" and "interested in long prefix caching" +# but did not give a reuse ratio. Sweep it rather than trusting one value. +PREFIX_FRAC="${PREFIX_FRAC:-0.5}" + +# --- Arrival model ------------------------------------------------------------- +# Poisson (burstiness 1.0) at a finite rate. NOT "inf". See reason 1 above. +BURSTINESS="${BURSTINESS:-1.0}" + +# Scenarios: "label:isl:osl:concurrency-list". Concurrency lists are chosen to fit +# the KV pool (see the capacity note); override with SCENARIOS=... to push past it. +DEFAULT_SCENARIOS="\ +256k-ctx:80000:1024:16 32 64 128 256|\ +1m-ctx:200000:1024:8 16 32 64" +IFS='|' read -ra SCENARIOS <<< "${SCENARIOS:-$DEFAULT_SCENARIOS}" + +ITERS="${SLO_ITERS:-1}" +RESULT_DIR="${RESULT_DIR:-$(dirname "$LOG")/slo_json}" +mkdir -p "$RESULT_DIR" + +echo "==============================================================" +echo " Customer SLO benchmark -- GLM-5.2 / MI355X" +echo " model : $MODEL_PATH" +echo " SLO : TTFT <= ${SLO_TTFT_MS} ms avg, TPOT <= ${SLO_TPOT_MS} ms avg" +echo " per-rank tgt : prefill ${TGT_PREFILL_TOK_S} tok/s, decode ${TGT_DECODE_TOK_S} tok/s (DP=${DP_RANKS})" +echo " prefix : ${PREFIX_FRAC} of ISL shared (agentic reuse)" +echo " arrival : Poisson, burstiness ${BURSTINESS}" +echo "==============================================================" + +# --------------------------------------------------------------------------- +# Shape warmup. Without this the first measured cell of a shape absorbs residual +# JIT and reports a wildly inflated TPOT (observed 302 ms vs ~89 ms steady state). +# Deliberately at low concurrency and few prompts so it is cheap. +# --------------------------------------------------------------------------- +warmup_shape() { + local isl=$1 osl=$2 pfx=$3 + echo "[WARMUP] isl=$isl osl=$osl prefix=$pfx" + timeout "${WARMUP_TIMEOUT:-3600}" vllm bench serve \ + --model "$MODEL_PATH" --backend vllm --host "$HOST" --port "$BENCHMARK_PORT" \ + --dataset-name random \ + --random-input-len "$isl" --random-output-len "$osl" --random-prefix-len "$pfx" \ + --num-prompts 4 --max-concurrency 2 --request-rate inf --ignore-eos \ + >>"${LOG}_warmup.log" 2>&1 +} + +# --------------------------------------------------------------------------- +# One measured cell. +# +# Request rate is derived, not guessed: to sustain `con` in flight with a +# per-request latency of about (TTFT + OSL*TPOT), Little's Law gives +# rate = con / latency +# Feeding a rate materially above that just rebuilds the infinite-rate queue and +# we are back to measuring drain time. We target the SLO latency, i.e. we ask +# "can the system serve the offered load AT the SLO", which is the actual question. +# --------------------------------------------------------------------------- +run_cell() { + local label=$1 isl=$2 osl=$3 con=$4 pfx=$5 iter=$6 + + local slo_lat_s + slo_lat_s=$(python3 -c "print(($SLO_TTFT_MS + $osl*$SLO_TPOT_MS)/1000.0)") + local rate + rate=$(python3 -c "print(round($con/$slo_lat_s, 4))") + + # Enough prompts that the measurement is not dominated by ramp-in/ramp-out. + # 4x concurrency, floored at 32, capped so a 200K cell stays affordable. + local prompts=$(( con * 4 )); [ "$prompts" -lt 32 ] && prompts=32 + [ "$prompts" -gt "${MAX_PROMPTS:-512}" ] && prompts="${MAX_PROMPTS:-512}" + + # Timeout scales with the work: prompts * (isl+osl) tokens, assuming a + # pessimistic floor rate, plus slack. + local tmo + tmo=$(python3 -c "print(int(900 + $prompts*($isl+$osl)/3000.0))") + + local jf="$RESULT_DIR/${label}_con${con}_iter${iter}.json" + echo "[RUNNING] $label isl=$isl osl=$osl con=$con prefix=$pfx rate=${rate}/s prompts=$prompts (timeout ${tmo}s)" + + timeout "$tmo" vllm bench serve \ + --model "$MODEL_PATH" --backend vllm --host "$HOST" --port "$BENCHMARK_PORT" \ + --dataset-name random \ + --random-input-len "$isl" --random-output-len "$osl" --random-prefix-len "$pfx" \ + --num-prompts "$prompts" \ + --max-concurrency "$con" \ + --request-rate "$rate" \ + --burstiness "$BURSTINESS" \ + --ignore-eos \ + --percentile-metrics ttft,tpot,itl,e2el \ + --metric-percentiles 50,95,99 \ + --goodput "ttft:${SLO_TTFT_MS}" "tpot:${SLO_TPOT_MS}" \ + --save-result --result-dir "$RESULT_DIR" \ + --result-filename "$(basename "$jf")" \ + 2>&1 | tee -a "${LOG}.log" + local rc=${PIPESTATUS[0]} + if [ "$rc" -eq 124 ]; then + echo "[STALL] $label con=$con timed out after ${tmo}s" | tee -a "${LOG}.log" "${LOG}_stalls.log" + fi + echo "$jf" +} + +for scenario in "${SCENARIOS[@]}"; do + IFS=':' read -r label isl osl conlist <<< "$scenario" + pfx=$(python3 -c "print(int($isl*$PREFIX_FRAC))") + echo; echo "########## scenario $label (isl=$isl osl=$osl, shared prefix=$pfx) ##########" + warmup_shape "$isl" "$osl" "$pfx" + for iter in $(seq 1 "$ITERS"); do + for con in $conlist; do + run_cell "$label" "$isl" "$osl" "$con" "$pfx" "$iter" >/dev/null + sleep 10 + done + done +done + +echo; echo "########## VERDICT ##########" +python3 "$(dirname "$0")/slo_report.py" \ + --result-dir "$RESULT_DIR" \ + --slo-ttft-ms "$SLO_TTFT_MS" --slo-tpot-ms "$SLO_TPOT_MS" \ + --target-prefill "$TGT_PREFILL_TOK_S" --target-decode "$TGT_DECODE_TOK_S" \ + --dp-ranks "$DP_RANKS" \ + --csv "${LOG}_slo.csv" +exit $? diff --git a/scripts/vllm_dissag/gen_workload.py b/scripts/vllm_dissag/gen_workload.py new file mode 100644 index 00000000..2fa7ec8b --- /dev/null +++ b/scripts/vllm_dissag/gen_workload.py @@ -0,0 +1,387 @@ +#!/usr/bin/env python3 +""" +Generate an explicit ISL/OSL distribution as a JSONL for `vllm bench serve +--dataset-name custom`. + +WHY THIS EXISTS +--------------- +The customer specified the workload as a DISTRIBUTION inside a CONTEXT WINDOW: + + ISL/OSL (p50 / p95 / p99) avg: 80K/1K (256K context) + avg: 200K/1K (1M context) + +benchmark_customer_slo.sh originally sent a single fixed --random-input-len 80000. +That is the *average* and nothing else: no spread, and the 256K/1M window is never +touched, so "it works at 256K context" would be assumed rather than measured. + +The obvious fix -- `--random-range-ratio` -- provably cannot express this. In +vllm/benchmarks/datasets/utils.py::get_sampling_params the draw is UNIFORM and +SYMMETRIC about the mean: + + input_low = floor(mean * (1 - r)) + input_high = ceil (mean * (1 + r)) + if not (0.0 <= r < 1.0): raise ValueError + +so the widest expressible support is [0, 2*mean]. Reaching a 256K tail from an 80K +mean needs r = 2.28; reaching 1M from 200K needs r = 4.24. Both are rejected by that +validator. It is also the wrong SHAPE -- uniform means the mean sits at the middle of +the range, whereas real agentic traffic is right-skewed (many short turns, a few very +long accumulated contexts). + +So we generate the lengths ourselves and feed them through CustomDataset, which reads +`{"prompt": ..., "output_tokens": N}` per line and therefore accepts an arbitrary +per-request length distribution. + +WHAT WE ASSUME, AND WHY IT IS DECLARED RATHER THAN HIDDEN +--------------------------------------------------------- +The sheet's header says "p50 / p95 / p99" but the cell only gives "avg: 80K" and a +window. Two of the three numbers we would need are simply not in the sheet. Rather +than invent three percentiles and present them as the customer's, we fix the two +things they DID state -- the mean and the window -- and pick the least-committal +shape that honours both: + + lognormal, solved so that mean = + p99 = TAIL_FRAC * + +TAIL_FRAC defaults to 1.0, i.e. the p99 request is a full-window request. That is the +aggressive reading and it is the one worth measuring, because it is the reading under +which the deployment is claimed to support a 256K/1M context at all. + +The achieved percentiles are RE-MEASURED from the generated sample and printed. They +are not the requested ones: clamping at the window and integer rounding both move +them. Report what came out, not what was asked for. + +SAMPLING ERROR IS THE DOMINANT DEVIATION, AND IT IS REPORTED, NOT HIDDEN +------------------------------------------------------------------------ +A lognormal with these parameters has CV = 0.64 (256K row) and 1.08 (1M row), so the +standard error of the realised mean is CV/sqrt(n): + + 256K row, n=256 -> 4.0% of the mean + 1M row, n=128 -> 9.5% of the mean + +Measured across 12 seeds the realised mean spanned 72,884-84,222 (14.2%) and +166,447-237,295 (35.4%). This is inherent to drawing few samples from a heavy tail -- +not a solver bug. Reaching a 2% standard error would need n=1,014 and n=2,903, which is +not affordable at 80K-200K tokens per request. + +Two consequences, both deliberate: + + * We do NOT stratify. Inverse-CDF placement would pin the mean to within 1-2% but it + caps the sample at the (n-0.5)/n quantile, so at n=128 the p99 reaches only 224,884 + instead of 262,144 -- it would buy mean fidelity by giving up the very thing the + window number is meant to test. The tail is the customer's stated requirement; the + mean deviation is a measurable property we can simply state. + * We therefore REPORT the deviation. Every run writes .meta.json carrying the + achieved distribution, the deviation from target as a percentage, the seed, and the + lognormal parameters -- so what goes to the customer is "our trace averaged 78,412 + tokens, 2.0% below the 80,000 target, p99 262,144" rather than an unqualified claim + to have hit 80K. Averaging several runs at DISTINCT seeds shrinks the error as + 1/sqrt(n_total): 10 seeds takes 4.0% -> 1.3% and 9.5% -> 3.0%. That is what + benchmark_avg_80K_ten.sh / benchmark_avg_200K_ten.sh do. + + Solving the lognormal: with X ~ LogN(mu, s), + ln(mean) = mu + s^2/2 ; ln(p99) = mu + z99*s (z99 = 2.32635) + Subtracting gives s^2/2 - z99*s + ln(p99/mean) = 0, so + s = z99 - sqrt(z99^2 - 2*ln(p99/mean)) (smaller root -> less skew) + which is real only while p99/mean <= exp(z99^2/2) = 14.97. Both customer rows are + well inside that (3.2x and 5.0x), but the guard is here because a future row may + not be. + +SHARED PREFIX +------------- +Agentic traffic re-sends a large, mostly-static preamble every turn. We emit the SAME +prefix text at the head of every prompt, so the server's prefix cache can actually hit +it -- this is the thing the customer meant by "interested in long prefix caching". +--prefix-frac is a fraction of the *median*, not of each request, so the shared block +is a fixed size and only the per-request remainder varies (which is how a real agent +behaves: constant system+tools, growing scratchpad). + +KV CAPACITY +----------- +A right-skewed ISL is not free: KV is charged on the ACTUAL length, so the tail costs +real HBM. GLM-5.2 MLA KV is (kv_lora_rank 512 + qk_rope_head_dim 64) = 576 +elem/tok/layer, FP8, 78 layers = 43.88 KiB/token. A single 262,144-token request is +therefore 10.97 GiB of KV on its own. This script prints the mean and p99 KV cost per +request so the concurrency you choose is a decision, not an OOM you discover later. + +USAGE + python3 gen_workload.py --mean-isl 80000 --context-window 262144 \\ + --osl 1024 --num-prompts 256 --tokenizer /models/GLM-5.2-FP8 \\ + --out /run_logs/wl_256k.jsonl + + # shape-only, no tokenizer needed -- prints the distribution and exits + python3 gen_workload.py --mean-isl 80000 --context-window 262144 --dry-run + +Then: + vllm bench serve --dataset-name custom --dataset-path /run_logs/wl_256k.jsonl \\ + --skip-chat-template --custom-output-len -1 ... + +--skip-chat-template matters: CustomDataset applies the chat template by default, +which prepends role tokens and would silently shift every length we just spent this +much effort placing. +""" + +import argparse +import json +import math +import os +import sys + +Z99 = 2.3263478740408408 # scipy.stats.norm.ppf(0.99) + +# GLM-5.2 MLA KV, bytes/token. See module docstring. +KV_BYTES_PER_TOKEN = (512 + 64) * 1 * 78 # 576 elem * 1 byte (fp8) * 78 layers + + +def solve_lognormal(mean, p99): + """Return (mu, sigma) of a lognormal with the given mean and 99th percentile.""" + ratio = p99 / float(mean) + if ratio <= 1.0: + raise ValueError( + "p99 target (%d) must exceed the mean (%d); a distribution cannot have " + "its 99th percentile at or below its mean." % (p99, mean) + ) + disc = Z99 * Z99 - 2.0 * math.log(ratio) + if disc < 0: + raise ValueError( + "p99/mean = %.2f exceeds the maximum %.2f attainable by a lognormal. " + "Lower --tail-frac or raise --mean-isl." % (ratio, math.exp(Z99 * Z99 / 2)) + ) + sigma = Z99 - math.sqrt(disc) + mu = math.log(mean) - sigma * sigma / 2.0 + return mu, sigma + + +def percentile(sorted_vals, q): + """Nearest-rank percentile. No numpy dependency -- this runs on the head node.""" + if not sorted_vals: + return 0 + k = max(0, min(len(sorted_vals) - 1, int(math.ceil(q / 100.0 * len(sorted_vals))) - 1)) + return sorted_vals[k] + + +def sample_lengths(n, mean, window, tail_frac, floor_tokens, seed): + """Draw n integer input lengths, clamped to [floor_tokens, window].""" + import random + + rng = random.Random(seed) + mu, sigma = solve_lognormal(mean, tail_frac * window) + lens = [] + for _ in range(n): + v = int(round(math.exp(rng.gauss(mu, sigma)))) + lens.append(max(floor_tokens, min(window, v))) + return lens, mu, sigma + + +def describe(lens, window, label="", target_mean=None, target_p99=None): + """Summarise a sample. Deviation vs target is printed as a PERCENTAGE because that + is the form the number has to travel in: "our trace averaged 78,412 tokens, 2.0% + below the 80,000 target" is reportable, "78412" alone invites the reader to assume + it was meant to be 80,000 exactly.""" + s = sorted(lens) + n = len(s) + mean = sum(s) / float(n) + p99 = percentile(s, 99) + stats = { + "n": n, + "min": s[0], + "p50": percentile(s, 50), + "p95": percentile(s, 95), + "p99": p99, + "max": s[-1], + "mean": mean, + "at_window": sum(1 for v in s if v >= window), + } + if target_mean: + stats["mean_dev_pct"] = 100.0 * (mean - target_mean) / float(target_mean) + if target_p99: + stats["p99_dev_pct"] = 100.0 * (p99 - target_p99) / float(target_p99) + + dev = "" + if target_mean: + dev = " [mean %+.1f%% vs target %d]" % (stats["mean_dev_pct"], target_mean) + print(" %-9s n=%d mean=%.0f p50=%d p95=%d p99=%d max=%d (%d clamped at window)%s" + % (label, stats["n"], mean, stats["p50"], stats["p95"], + stats["p99"], stats["max"], stats["at_window"], dev)) + return stats + + +def expected_se_pct(sigma, n): + """Standard error of the sample mean, as a percentage of the mean. + + For X ~ LogN(mu, sigma) the coefficient of variation is sqrt(exp(sigma^2) - 1), + independent of mu, and SE(mean)/mean = CV/sqrt(n). Printed alongside the achieved + mean so a reader can tell a normal draw from a genuinely anomalous one instead of + guessing -- a realised mean 4% off target at n=256 is expected, not a defect. + """ + cv = math.sqrt(math.exp(sigma * sigma) - 1.0) + return cv, 100.0 * cv / math.sqrt(n) + + +def kv_report(stats, concurrency_list): + gib = 1024.0 ** 3 + mean_gib = stats["mean"] * KV_BYTES_PER_TOKEN / gib + p99_gib = stats["p99"] * KV_BYTES_PER_TOKEN / gib + max_gib = stats["max"] * KV_BYTES_PER_TOKEN / gib + print(" KV per request: mean %.2f GiB, p99 %.2f GiB, max %.2f GiB" + % (mean_gib, p99_gib, max_gib)) + if concurrency_list: + print(" KV at concurrency (mean-length steady state / all-p99 worst case):") + for c in concurrency_list: + print(" con=%-4d %8.0f GiB / %8.0f GiB" % (c, c * mean_gib, c * p99_gib)) + print(" (one 8x MI355X node has ~930 GiB of KV pool at FP8/util 0.80)") + + +def build_prompts(lens, tokenizer_path, prefix_tokens, out_path, osl, trust_remote_code): + """Materialise prompts whose re-tokenised length matches the drawn length.""" + from transformers import AutoTokenizer + + tok = AutoTokenizer.from_pretrained(tokenizer_path, trust_remote_code=trust_remote_code) + vocab = tok.vocab_size + special = set(tok.all_special_ids or []) + allowed = [t for t in range(1000, min(vocab, 60000)) if t not in special] + if not allowed: + raise RuntimeError("tokenizer produced no usable token ids") + + import random + rng = random.Random(1234) + + # One shared prefix, identical text on every request -> real prefix-cache hits. + prefix_ids = [allowed[rng.randrange(len(allowed))] for _ in range(prefix_tokens)] + prefix_text = tok.decode(prefix_ids) if prefix_tokens else "" + prefix_real = len(tok(prefix_text).input_ids) if prefix_tokens else 0 + print(" shared prefix: requested %d tok, materialised %d tok" + % (prefix_tokens, prefix_real)) + + achieved = [] + with open(out_path, "w") as f: + for i, target in enumerate(lens): + body = max(1, target - prefix_real) + ids = [allowed[(rng.randrange(len(allowed)) + i) % len(allowed)] + for _ in range(body)] + text = prefix_text + tok.decode(ids) + got = len(tok(text).input_ids) + # decode -> re-encode is NOT length preserving: the decoded text can + # re-tokenise into a different number of tokens because adjacent pieces + # merge. Converge by adjusting the body length by the observed error. + # Bounded at 4 passes -- this is a benchmark input, not a proof; we + # report the achieved distribution rather than pretending it is exact. + for _ in range(4): + if got == target: + break + body = max(1, body + (target - got)) + ids = [allowed[(rng.randrange(len(allowed)) + i) % len(allowed)] + for _ in range(body)] + text = prefix_text + tok.decode(ids) + got = len(tok(text).input_ids) + achieved.append(got) + # "input_tokens" is NOT read by CustomDataset (it only looks at "prompt" + # and "output_tokens"; extra keys ride along unused). It is written so + # downstream tooling can size timeouts from the real token count instead + # of estimating from character length. + f.write(json.dumps({"prompt": text, "output_tokens": int(osl), + "input_tokens": int(got)}) + "\n") + if (i + 1) % 32 == 0: + print(" ... %d/%d" % (i + 1, len(lens)), file=sys.stderr) + return achieved + + +def main(): + ap = argparse.ArgumentParser(description=__doc__, + formatter_class=argparse.RawDescriptionHelpFormatter) + ap.add_argument("--mean-isl", type=int, required=True, + help="the 'avg' from the customer sheet, e.g. 80000") + ap.add_argument("--context-window", type=int, required=True, + help="the window from the sheet, e.g. 262144 or 1048576") + ap.add_argument("--tail-frac", type=float, default=1.0, + help="p99 target as a fraction of the window (default 1.0 = the " + "p99 request fills the window)") + ap.add_argument("--osl", type=int, default=1024) + ap.add_argument("--num-prompts", type=int, default=256) + ap.add_argument("--prefix-frac", type=float, default=0.5, + help="shared cacheable prefix as a fraction of the MEDIAN isl") + ap.add_argument("--min-isl", type=int, default=1024, + help="floor, so the left tail stays a plausible agent turn") + ap.add_argument("--seed", type=int, default=0) + ap.add_argument("--tokenizer", default="", + help="model path; required unless --dry-run") + ap.add_argument("--trust-remote-code", action="store_true") + ap.add_argument("--out", default="") + ap.add_argument("--dry-run", action="store_true", + help="print the distribution and KV cost; write nothing") + ap.add_argument("--concurrency", default="16 32 64 128 256", + help="space-separated list, for the KV table only") + a = ap.parse_args() + + print("workload: mean ISL %d, window %d, p99 target %d (tail-frac %.2f), OSL %d" + % (a.mean_isl, a.context_window, int(a.tail_frac * a.context_window), + a.tail_frac, a.osl)) + + try: + lens, mu, sigma = sample_lengths(a.num_prompts, a.mean_isl, a.context_window, + a.tail_frac, a.min_isl, a.seed) + except ValueError as e: + print("ERROR: %s" % e) + return 2 + p99_target = int(a.tail_frac * a.context_window) + cv, se_pct = expected_se_pct(sigma, a.num_prompts) + print(" lognormal mu=%.4f sigma=%.4f (analytic p50=%.0f, mean=%.0f)" + % (mu, sigma, math.exp(mu), math.exp(mu + sigma * sigma / 2))) + print(" CV=%.3f -> expected SE of the realised mean at n=%d is %.1f%%. A sample mean" + % (cv, a.num_prompts, se_pct)) + print(" within ~2 SE (%.1f%%) of target is a normal draw, NOT a defect; average" + % (2 * se_pct)) + print(" several runs at DISTINCT --seed to shrink it as 1/sqrt(n).") + + stats = describe(lens, a.context_window, "requested", a.mean_isl, p99_target) + cons = [int(x) for x in a.concurrency.split()] if a.concurrency else [] + kv_report(stats, cons) + + if a.dry_run: + print("dry run -- no file written") + return 0 + + if not a.tokenizer or not a.out: + print("ERROR: --tokenizer and --out are required unless --dry-run") + return 2 + + prefix_tokens = int(stats["p50"] * a.prefix_frac) + os.makedirs(os.path.dirname(os.path.abspath(a.out)) or ".", exist_ok=True) + achieved = build_prompts(lens, a.tokenizer, prefix_tokens, a.out, a.osl, + a.trust_remote_code) + ach = describe(achieved, a.context_window, "achieved", a.mean_isl, p99_target) + + # Sidecar. The ACHIEVED distribution is what we actually served, so it is what gets + # quoted upstream -- see the module docstring. Written next to the JSONL rather than + # only to stdout because stdout is a 30k-line benchmark log by the time anyone reads + # it, and this number has to survive the trip to a customer slide intact. + meta = { + "target": { + "mean_isl": a.mean_isl, + "context_window": a.context_window, + "tail_frac": a.tail_frac, + "p99_isl": p99_target, + "osl": a.osl, + }, + "lognormal": {"mu": mu, "sigma": sigma, "cv": cv, + "expected_se_pct_of_mean": se_pct}, + "seed": a.seed, + "prefix_tokens_requested": prefix_tokens, + "achieved": ach, + "kv_bytes_per_token": KV_BYTES_PER_TOKEN, + "note": ("Random draw, not stratified: the p99 must actually reach the context " + "window, which inverse-CDF placement at small n cannot do. The mean " + "therefore carries sampling error of about " + "%.1f%% at this n -- report 'achieved', not 'target'." % se_pct), + } + meta_path = a.out + ".meta.json" + with open(meta_path, "w") as f: + json.dump(meta, f, indent=2, sort_keys=True) + + print("wrote %s (%d requests)" % (a.out, len(achieved))) + print("wrote %s (achieved distribution -- quote THIS upstream)" % meta_path) + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/vllm_dissag/slo_report.py b/scripts/vllm_dissag/slo_report.py new file mode 100755 index 00000000..f920d5e6 --- /dev/null +++ b/scripts/vllm_dissag/slo_report.py @@ -0,0 +1,183 @@ +#!/usr/bin/env python3 +""" +Turn vLLM bench-serve result JSONs into a customer-facing SLO verdict. + +Reads the --save-result JSONs written by benchmark_customer_slo.sh and answers the +only two questions the customer actually asked: + + does it meet the SLO, and if not, by how much + +Design notes, because each of these was a real trap: + + * The customer's throughput targets are PER RANK ("prefill: 34000tokens/s + (per-rank)"). vLLM reports AGGREGATE. Dividing by --dp-ranks is therefore not + cosmetic -- reporting the aggregate against a per-rank target overstates the + result by exactly the DP degree. + + * "Prefill throughput" is not a field vLLM emits. We derive it as + total_input_tokens / duration, which is prompt tokens processed per wall second. + That is an END-TO-END rate that includes decode time for the same requests, so + it is a *lower bound* on the prefill engine's rate, not a kernel measurement. + Labelled as such below rather than quietly presented as the engine number. + + * Goodput is the honest headline. mean_ttft can pass while a third of requests + miss the SLO; request_goodput counts only the requests that met BOTH ttft and + tpot thresholds, so goodput/throughput is the fraction of traffic actually + served acceptably. + + * The gate is on the MEAN, because the acceptance row of the customer sheet says + "avg ttft: <7s" and "avg tpot: 50ms/token(avg)". Percentiles are reported so + the tail is visible, but they are not gated -- the customer did not set them. +""" + +import argparse +import csv +import glob +import json +import os +import sys + + +def _fmt(v, spec="%.1f"): + return "n/a" if v is None else spec % v + + +def load(result_dir): + rows = [] + for path in sorted(glob.glob(os.path.join(result_dir, "*.json"))): + try: + with open(path) as f: + d = json.load(f) + except (OSError, ValueError) as e: + print(" warning: could not read %s (%s) -- skipped" % (path, e)) + continue + if "mean_ttft_ms" not in d: + # Not a bench-serve result (or a run that died before reporting). + print(" warning: %s has no metrics -- run likely failed; skipped" + % os.path.basename(path)) + continue + # Recover the label/concurrency from the filename we chose in the runner: + #