Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
41 commits
Select commit Hold shift + click to select a range
56f83c7
vllm_dissag: GLM-5.1-FP8 (MLA+DSA) MoRI-EP WideEP disaggregated enabl…
raviguptaamd Jul 8, 2026
7408984
vllm_dissag: make NIAH harness thinking-model-aware (GLM-5.1)
raviguptaamd Jul 8, 2026
9cfe987
vllm_dissag: NIAH multi-seed support (NIAH_SEEDS) for variance-aware …
raviguptaamd Jul 8, 2026
74c69b7
vllm_dissag: NIAH gate robust to cold-start JIT (warmup + readiness p…
raviguptaamd Jul 10, 2026
8d0ecab
[GLM-5.1] v0.27 level-set: 3.4x decode speedup, 200K context, RDMA + …
raviguptaamd Aug 16, 2026
0724197
[GLM-5.1] Default WITH_NIXL=0 (lean MoRI-EP-only image)
raviguptaamd Aug 16, 2026
09d84ce
[GLM-5.1] Add long-context NIAH harness + vllm-disagg operational pla…
raviguptaamd Aug 16, 2026
550603f
[GLM-5.1] Remove dead-end MoRI EP env plumbing and its misleading com…
raviguptaamd Aug 16, 2026
69eb062
vllm_disagg: remove stray NFS silly-rename artifact
raviguptaamd Aug 16, 2026
0f660fe
vllm_disagg: build the GLM image for gfx950 (MI355X) + ionic, and fix…
raviguptaamd Aug 16, 2026
500353f
vllm_disagg/moriio: pass RDMA tuning via extra_config, widen GLM gate…
raviguptaamd Aug 16, 2026
d40c5e3
vllm_disagg/models.yaml: add GLM-5.2-FP8 and GLM-5.2-MXFP4 recipes fo…
raviguptaamd Aug 16, 2026
24b5c9c
vllm_disagg: make the launcher portable across sites, and add the AAC…
raviguptaamd Aug 16, 2026
adc163c
vllm_disagg/niah: make the accuracy check a real pass/fail gate, jitt…
raviguptaamd Aug 16, 2026
801e536
vllm_disagg/parse_to_csv: stop dropping the first benchmark cell; fix…
raviguptaamd Aug 16, 2026
bef784f
vllm_disagg: document the GLM-5.2 MI355X recipe
raviguptaamd Aug 16, 2026
8212dfa
vllm_disagg: add customer-SLO benchmark (agentic 80K/200K, Poisson, g…
raviguptaamd Aug 17, 2026
907cceb
vllm_disagg: pool ten seeds for the avg-ISL rows, and unhardcode the …
raviguptaamd Aug 17, 2026
52ba0c4
vllm_disagg: NIAH ladder in tokens to 950K, with per-needle depth rep…
raviguptaamd Aug 17, 2026
bc56811
vllm_disagg: register the new benchmarks in the launcher, and documen…
raviguptaamd Aug 17, 2026
02ba118
vllm_disagg: fix a stale assertion that failed on correct code
raviguptaamd Aug 17, 2026
e200a14
moriio_nic_sweep: two-node RDMA block-size sweep, results, and test i…
raviguptaamd Aug 17, 2026
bea677c
vllm_disagg: let MODELS_YAML reach the container, and document co-ten…
raviguptaamd Aug 18, 2026
cf413fc
vllm_disagg: record reproduced GLM-5.2-FP8 1P/1D sweep on a second no…
raviguptaamd Aug 18, 2026
4727924
vllm_disagg: add NIaH-256K ladder wrapper and wire into launcher
raviguptaamd Aug 23, 2026
65d9c54
vllm_disagg: avg-workload SLO harness with cell-warmup and min-iter
raviguptaamd Aug 23, 2026
343b346
vllm_disagg: GLM-5.2 recipe -- weights, measured sweep, NIaH and MoRI…
raviguptaamd Aug 23, 2026
eda3657
moriio_nic_sweep: support attaching a second pre-existing job for two…
raviguptaamd Aug 23, 2026
48d62b5
vllm_disagg: opt-in MoRI PR#558 (MORI_EP_OVER_RDMA) for cross-node EP…
raviguptaamd Aug 27, 2026
99d6629
vllm_disagg: EP16 cross-node recipe + Crusoe/ionic TP8/EP8/EP16 results
raviguptaamd Aug 27, 2026
a8f6049
vllm_disagg: GLM-5.2-MXFP4 boots on gfx950 -- Code-209 resolved on cu…
raviguptaamd Aug 27, 2026
4554161
vllm_disagg: config-selection guide -- when TP8 vs EP8 vs EP16 wins
raviguptaamd Aug 28, 2026
d78789b
vllm_disagg: long-context scaling-sweep orchestrator (50K-750K, 3 con…
raviguptaamd Aug 28, 2026
4b6a57b
GLM-5.2 MTP disagg: full portable patch stack + apply-all + README
raviguptaamd Aug 30, 2026
58d1905
GLM-5.2 EP16+MTP: keep decode cudagraphs (FULL_AND_PIECEWISE), don't …
raviguptaamd Aug 30, 2026
3971f59
vllm_disagg: fix longctx sweep var-name bug ($osl->$OSL) + NIAH v3
raviguptaamd Aug 30, 2026
0b0aedc
vllm_disagg: pin GLM-5.2 MTP/EP16 image to patched vLLM+AITER forks
raviguptaamd Aug 30, 2026
6079aa1
docker: fix MoRI wheel build — pin SETUPTOOLS_SCM_PRETEND_VERSION
raviguptaamd Aug 30, 2026
1678059
docker: pin libibverbs to 1.14/rdmav34 before MoRI build (ionic ABI)
raviguptaamd Aug 30, 2026
91b74e7
docker: re-pin libibverbs 1.14/rdmav34 as FINAL layer (runtime ABI)
raviguptaamd Aug 30, 2026
e194829
vllm_disagg: correct EP16-MTP status in docs — known limitation, not …
raviguptaamd Sep 1, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
45 changes: 45 additions & 0 deletions docker/mori_pr558_ionic.patch
Original file line number Diff line number Diff line change
@@ -0,0 +1,45 @@
diff --git a/src/application/transport/rdma/proxy/proxy_thread.cpp b/src/application/transport/rdma/proxy/proxy_thread.cpp
index e3728c9..c41d5b0 100644
--- a/src/application/transport/rdma/proxy/proxy_thread.cpp
+++ b/src/application/transport/rdma/proxy/proxy_thread.cpp
@@ -163,10 +163,18 @@ void ProxyThread::MainLoop() {
uint32_t wr_qp[kMaxBatch];
int batch_count = 0;

+ const bool hpdbg = getenv("MORI_PROXY_DEBUG") != nullptr;
+ unsigned long long spins = 0;
+ uint32_t last_head_dbg = 0xFFFFFFFFu;
while (!ring_->shutdown) {
batch_count = 0;

uint32_t head = ring_->gpu_head;
+ if (hpdbg && (head != last_head_dbg || (++spins % 200000000ULL) == 0)) {
+ fprintf(stderr, "[proxy-dbg] gpu_head=%u next_slot=%u qps=%zu ring=%p\n", head,
+ next_slot_, qps_.size(), (void*)ring_);
+ last_head_dbg = head;
+ }
while (next_slot_ < head && batch_count < kMaxBatch) {
uint32_t slot = next_slot_ & PROXY_RING_MASK;
volatile ProxyCmd* cmd = &ring_->cmds[slot];
diff --git a/src/application/transport/rdma/rdma.cpp b/src/application/transport/rdma/rdma.cpp
index db3b4c5..1d230f1 100644
--- a/src/application/transport/rdma/rdma.cpp
+++ b/src/application/transport/rdma/rdma.cpp
@@ -339,6 +339,17 @@ bool ReadIbEnableRelaxedOrderingEnv() {
}

int MaybeAddRelaxedOrderingFlag(int accessFlag) {
+ // Pensando ionic (AINIC) does not support REMOTE_ATOMIC-capable MRs; ibv_reg_mr
+ // returns EINVAL (errno 22) for any MR registered with IBV_ACCESS_REMOTE_ATOMIC,
+ // regardless of size. Set MORI_IO_DISABLE_ATOMIC_MR=1 to strip the atomic bit so
+ // MoRI-IO KV registration succeeds on ionic (RDMA write/read still work; the KV
+ // transfer path does not rely on NIC remote-atomic).
+ {
+ const char* v = getenv("MORI_IO_DISABLE_ATOMIC_MR");
+ if (v && v[0] == '1') {
+ accessFlag &= ~IBV_ACCESS_REMOTE_ATOMIC;
+ }
+ }
#ifdef IBV_ACCESS_RELAXED_ORDERING
if (ReadIbEnableRelaxedOrderingEnv()) {
return accessFlag | IBV_ACCESS_RELAXED_ORDERING;
448 changes: 448 additions & 0 deletions docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile

Large diffs are not rendered by default.

32 changes: 32 additions & 0 deletions models.json
Original file line number Diff line number Diff line change
Expand Up @@ -3420,6 +3420,38 @@
},
"args": "-N 2 -n 2"
},
{
"name": "pyt_vllm_disagg_mori_glm-5.1-fp8",
"url": "",
"dockerfile": "docker/vllm_disagg_inference.glmv5.1",
"scripts": "scripts/vllm_dissag/run_xPyD_models.slurm",
"data": "huggingface",
"n_gpus": "-1",
"owner": "mad.support@amd.com",
"training_precision": "",
"tags": [
"pyt",
"vllm",
"vllm_disagg",
"mori_ep",
"inference"
],
"timeout": -1,
"distributed": {
"launcher": "slurm_multi"
},
"env_vars": {
"DOCKER_IMAGE_NAME": "<supply-your-image>",
"MODEL_NAME": "GLM-5.1-FP8",
"xP": "1",
"yD": "1",
"RUN_MORI": "1",
"RUN_DEEPEP": "0",
"GLM_SKIP_PATCHERS": "1",
"BENCHMARK_COMBINATIONS": "1024/1024"
},
"args": "-N 2 -n 2"
},
{
"name": "pyt_vllm_disagg_mori_deepseek-v3-5layer",
"url": "",
Expand Down
100 changes: 100 additions & 0 deletions scripts/moriio_nic_sweep/Dockerfile
Original file line number Diff line number Diff line change
@@ -0,0 +1,100 @@
# MoRI-IO two-node RDMA block-size sweep -- test image
#
# WHY THIS DOCKERFILE EXISTS, AND WHAT IT IS *NOT*
# ------------------------------------------------
# It does NOT build a MoRI stack. It layers the missing *harness* onto the
# workload image that PR#205 already builds, so that the number this benchmark
# produces characterises the stack we actually ship for GLM-5.2 disagg -- not a
# freshly-cloned mori that nothing else in this repo runs.
#
# The recorded 378.1 GB/s result was produced against the image's own pinned
# libmori (MORI_REF=42e895472b08). Verified in-container:
# mori pkg : /usr/local/lib/python3.12/dist-packages/mori/__init__.py
# That is the number to trust precisely because the engine is the shipped one.
#
# The one thing the workload image lacks is tests/python/io/benchmark.py: the
# mori wheel installs the `mori` package but not the repo's test tree. The
# sweep scripts therefore need a benchmark.py from somewhere. Two ways:
#
# (a) bind-mount a mori checkout at /opt/mori (what moriio_sweep.slurm does
# by default; requires a FULL, non-sparse checkout on the host), or
# (b) build this image, which bakes a benchmark.py pinned to the SAME commit
# as the image's libmori, so harness and engine cannot drift.
#
# (b) is strictly better and is why this file exists. The recorded run used
# (a) with an older 6ad812c checkout, whose benchmark.py predates
# --mem-type/--max-chunks/--chunk-bytes. That cost nothing for the recorded
# sweep -- those flags only re-state defaults the planner already uses -- but it
# did mean CHUNK_SWEEP=1 could not run. Build this image and it can.
#
# BUILD
# # BASE must be the image PR#205 built (or a rebuild of
# # docker/vllm_disagg_inference.glmv5.1.ubuntu.amd.Dockerfile).
# docker build --network=host \
# --build-arg BASE_IMAGE=rocmshared/vllm-disagg:glm52-gfx950-ionic-v1 \
# -f Dockerfile -t moriio-nic-sweep:latest .
#
# RUN -- then point the sweep at it and skip the host mount entirely:
# export DOCKER_IMAGE_NAME=moriio-nic-sweep:latest
# export MORI_SRC=SKIP # harness is baked in at /opt/mori
# sbatch moriio_sweep.slurm
#
# The container still needs the host RDMA stack bind-mounted at run time (24
# mounts, 4 groups) -- see README "Exposing the host RDMA stack". That is a
# runtime concern; nothing here can bake it in.

ARG BASE_IMAGE=rocmshared/vllm-disagg:glm52-gfx950-ionic-v1
FROM ${BASE_IMAGE}

# Must match the MORI_REF baked into BASE_IMAGE. If you bump the base, bump
# this too, or the harness and the engine describe different code. The check
# below turns that mistake into a build failure instead of a silent mismatch.
ARG MORI_REF=42e895472b08

# Where the sweep scripts expect to find the mori repo (MORI_REPO / MORI_MNT).
ENV MORI_REPO=/opt/mori
ENV MORI_HARNESS_BAKED=1

# numactl: absent from the workload image, and its absence is visible in the
# recorded logs as "[warn] numactl absent -- NIC/CPU affinity left to the
# scheduler." MatchCpuNics() then orders rails unpinned. It cost nothing at the
# large-block end (0.5% rail spread) but small-block numbers are the ones that
# measure software overhead, so pin them properly here.
RUN apt-get update && apt-get install -y --no-install-recommends \
numactl git ca-certificates \
&& rm -rf /var/lib/apt/lists/*

# Fetch ONLY the harness, at the engine's commit. --filter=blob:none keeps this
# cheap; we then materialise the full tree (NOT a sparse checkout -- a sparse
# checkout tracks tests/ without materialising it, so benchmark.py can be "in
# git" and still absent from disk, which is a preflight failure the sweep
# scripts explicitly check for).
RUN git clone --filter=blob:none https://github.com/ROCm/mori.git ${MORI_REPO} \
&& git -C ${MORI_REPO} checkout ${MORI_REF} \
&& test -f ${MORI_REPO}/tests/python/io/benchmark.py \
|| (echo "FATAL: benchmark.py absent at MORI_REF=${MORI_REF}" && exit 1)

# Assert harness == engine. The image's libmori is authoritative; if the pin
# above disagrees with what is installed, fail here rather than publish a
# number attributed to the wrong commit.
RUN set -eux; \
installed="$(grep -oE 'MORI_REF=[^@]+' /app/versions.txt | head -1 | cut -d= -f2 || echo UNKNOWN)"; \
echo "harness MORI_REF=${MORI_REF} image MORI_REF=${installed}"; \
if [ "${installed}" != "UNKNOWN" ] && [ "${installed}" != "${MORI_REF}" ]; then \
echo "FATAL: harness pin (${MORI_REF}) != image libmori (${installed})."; \
echo " Rebuild with --build-arg MORI_REF=${installed}."; exit 1; \
fi; \
echo "MORIIO_SWEEP_HARNESS_REF=${MORI_REF}" >> /app/versions.txt

# Do NOT put ${MORI_REPO} on PYTHONPATH. It would be harmless in practice --
# the repo's python package lives at python/mori, not the tree root, so it
# cannot shadow the installed wheel -- but relying on that is fragile. The
# sweep runs the harness as `python3 -m tests.python.io.benchmark` from
# MORI_REPO, which needs cwd, not PYTHONPATH. `import mori` must keep
# resolving to the installed wheel; that is the whole point.

COPY run_moriio_sweep.sh /opt/moriio_nic_sweep/run_moriio_sweep.sh
COPY aggregate_sweep.py /opt/moriio_nic_sweep/aggregate_sweep.py
RUN chmod +x /opt/moriio_nic_sweep/run_moriio_sweep.sh /opt/moriio_nic_sweep/aggregate_sweep.py

WORKDIR ${MORI_REPO}
Loading