Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
94 commits
Select commit Hold shift + click to select a range
3db7de5
feat(vllm_dissag): TP-aware wideEP topology + multi-node pools + niah…
Cemberk Aug 13, 2026
a5d7eb6
feat(vllm_dissag): honor models.yaml dp: flags on moriio wideEP; add …
Cemberk Aug 13, 2026
01c56bb
feat(kimi-k3): MI300X image + colocated multi-node vLLM harness
Cemberk Aug 13, 2026
13c2005
feat(kimi-k3): register MI300X models + emit perf CSV from NIAH runs
Cemberk Aug 13, 2026
893d4d0
docs(kimi-k3): MI300X multi-node recipes under benchmark/kimi_k3/mi300x/
Cemberk Aug 13, 2026
a7dd951
fix(kimi-k3): correct node sizing, NIAH model tag, and colocated run …
Cemberk Aug 17, 2026
8cf7ca9
fix(vllm_dissag): fail fast when the allocation is smaller than xP+yD
Cemberk Aug 17, 2026
da621b7
refactor(kimi-k3): report NIAH results on the narrow madengine CSV co…
Cemberk Aug 18, 2026
09d3f99
fix(kimi-k3): size the SLURM allocation from slurm.nodes, not distrib…
Cemberk Aug 25, 2026
66144b0
fix(kimi-k3): make the MI300X recipes actually runnable on gfx942
Cemberk Aug 27, 2026
c5e4b1e
docs(kimi-k3): record the MI300X failure modes and site prerequisites
Cemberk Aug 27, 2026
06ddd99
feat(vllm_dissag): integrate Kimi-K3-MXFP4 2P/2D as first-class model
MIR-AMD Aug 28, 2026
a30a1e3
fix(kimi-k3): unstick the colocated shutdown and make NIAH results fa…
Cemberk Aug 28, 2026
bad212c
update default isnt being passed through
Cemberk Aug 28, 2026
db276a6
Survive cross-node startup skew in the colocated multinode launcher
Cemberk Sep 3, 2026
889fc0d
Default the checkpoint pre-warm off and bound it with a budget
Cemberk Sep 3, 2026
d4ccac7
Let NCCL_DEBUG through the container so the sglang logs can be quieted
Cemberk Sep 7, 2026
1f1a4a6
Run KV-cache transfer over the CX7 rail NICs, and verify placement first
Cemberk Sep 8, 2026
92788a0
Capture only the batch sizes the benchmark actually requests
Cemberk Sep 8, 2026
b2e6094
Install the launcher's Python deps in the image, not on every node
Cemberk Sep 8, 2026
cdff5d5
Merge origin/develop into the kimi-k3 branch
Cemberk Sep 8, 2026
f074955
Probe local NVMe for weights first, and stop the probe faking a faile…
Cemberk Sep 8, 2026
ee3c5d0
Make the migrated registry entries relative to their own scripts dir
Cemberk Sep 8, 2026
364f0c0
Forward the full NIAH knob set from the colocated launcher, and asser…
Cemberk Sep 8, 2026
84fd837
Make BENCHMARK_ITR actually iterate, and fix the parser it was masking
Cemberk Sep 8, 2026
6688d67
Drop run-specific identifiers from code comments
Cemberk Sep 8, 2026
d9a22b7
Give COLOCATED_FORWARD_ENV a default so the launcher can start
Cemberk Sep 10, 2026
b87f50f
Load colocated weights from local NVMe when every node has them
Cemberk Sep 10, 2026
cfa4f5e
Put the site's facts in one sourced file, for both CI paths
Cemberk Sep 10, 2026
a7830f7
Capture only the decode batch sizes the benchmark asks for
Cemberk Sep 11, 2026
d9c93e2
Say which kind of miss the weight probe hit
Cemberk Sep 11, 2026
de61b04
Say what CI-created directories are, so someone else can clean them
Cemberk Sep 11, 2026
1e42e8f
Keep apostrophes out of the srun body, where one ends the script
Cemberk Sep 11, 2026
b85ac2b
Mount the RDMA sonames, not only the versioned files
Cemberk Sep 14, 2026
c772cdf
Refuse a weight path that differs per node, and fill the CSV provenance
Cemberk Sep 14, 2026
5334f51
Require local weights for Kimi-K3 instead of falling back to NFS
Cemberk Sep 14, 2026
26fcdb3
Resolve weights by MODEL_WEIGHTS_NAME when a card is an alias
Cemberk Sep 15, 2026
1de0c6f
Detect the fabric archetype instead of hardcoding one cluster's
Cemberk Sep 15, 2026
52338b0
Add a way-4 twin of the proven Kimi-K3 card, for parity testing
Cemberk Sep 16, 2026
a84416a
kimik3 2p2d: KV write-fence fix, PR241 env parity, MI300X+CX7 fabric
Sep 16, 2026
6436bde
kimik3 2p2d: fix GPU_MEMORY_UTILIZATION 0.85->0.80 for MI300X baseline
Sep 17, 2026
16e2a04
Keep comments out of the docker run continuation, where one ends the …
Cemberk Sep 17, 2026
f93c019
Kimi-K3-MXFP4: align disagg recipe to PR#241-validated config
Sep 17, 2026
9213331
Get my own prose out of the srun body, which it was truncating
Cemberk Sep 17, 2026
45f11b6
Check that both execution paths still generate a runnable invocation
Cemberk Sep 17, 2026
cde355a
Kimi-K3-MXFP4: enable 8th RDMA rail (mlx5_9) for disagg
MIR-AMD Sep 18, 2026
6e4ab58
Kimi-K3-MXFP4: genericize TP-within-EP via EP_TP_SIZE knob
MIR-AMD Sep 18, 2026
f26b991
Kimi-K3-MXFP4: finish genericization (P1/P3) + EP_TP_SIZE divisibilit…
MIR-AMD Sep 19, 2026
b4006e3
Clean up the container when the run ends early, not only when it succ…
Cemberk Sep 21, 2026
2f0060b
Forward AITER_SITUV2_A8W4 into the container
Cemberk Sep 21, 2026
a3df9e7
Drop the site layer from the way-4 twin
Cemberk Sep 21, 2026
f5518e1
vllm_dissag: docs+tests for in-tree Kimi-K3 Dockerfile & generic topo…
MIR-AMD Sep 22, 2026
c1c1f67
docker: add in-tree Kimi-K3-MXFP4 disagg Dockerfile
MIR-AMD Sep 22, 2026
499611e
moriio: keep /tmp/vllm_cache default (revert unneeded /opt flip)
MIR-AMD Sep 22, 2026
2d4a577
rixl: revert eval->_model_config_to_array (out of scope for K3 wideEP)
MIR-AMD Sep 22, 2026
2f86ba6
Set WIDE_EP=1 on the cards whose models require it
Cemberk Sep 24, 2026
b7e477c
Merge PR 237 (Kimi-K3-MXFP4) onto one TP-within-EP knob
Cemberk Sep 25, 2026
4a2ec0c
Refuse the wrong GPU in the launcher, where both CI paths meet
Cemberk Sep 25, 2026
c8fff40
Advertise pod hosts only when the router addresses the whole pool
Cemberk Sep 25, 2026
5ea0f65
Add a native-MXFP4 MI355X Kimi-K3 disagg recipe, and follow the GPU i…
Cemberk Sep 25, 2026
04332cb
Fail a disagg job fast, on every node, with the reason in the log
Cemberk Sep 28, 2026
983ce6f
Fail an sglang disagg job fast, on every node
Cemberk Sep 28, 2026
511a70c
Let NCCL_DEBUG through the vllm_dissag container too
Cemberk Sep 28, 2026
ed3163d
Pin rocSHMEM in the ROCm 7.2 disagg images
Cemberk Sep 28, 2026
b6bba68
Print a dead server's first error lines, not only its tail
Cemberk Sep 28, 2026
cdbfc16
Stop a failing node's servers before it exits
Cemberk Sep 28, 2026
5fa145d
Show who holds the node's GPU memory when a server fails to start
Cemberk Sep 28, 2026
4a28f2c
Pass the resolved GPU_MEMORY_UTILIZATION to vllm serve on rixl too
Cemberk Sep 28, 2026
2eee2d0
Publish agentic runs' metrics as the perf.csv madengine collects
Cemberk Sep 28, 2026
f1ce8d7
Kill wedged servers with SIGKILL, and fail on a worker's RCCL error
Cemberk Sep 29, 2026
c4d26ea
Turn off RCCL's MSCCL path for the llama-3.3-70B TP recipe
Cemberk Sep 29, 2026
52ab5e1
Give rixl/deepep the recipe's KV block size, dtype and cache bytes
Cemberk Sep 29, 2026
834b8d6
Stop a node's container when its SLURM task is cancelled
Cemberk Sep 29, 2026
3ea8dc5
Print GPU memory-access faults among a dead server's first errors
Cemberk Sep 29, 2026
54e77dc
Keep one noisy warning from hiding a dead server's real errors
Cemberk Sep 29, 2026
5130bda
Give each SLURM cluster an allocation profile both CI paths read
Cemberk Sep 29, 2026
5c727ab
Document running the multinode workloads through madengine or sbatch
Cemberk Sep 29, 2026
f0a5daa
Describe past failures in comments without internal job IDs or hostnames
Cemberk Sep 29, 2026
0b7070a
Run the GLM-5.1 1P/1D cards TP8 within each pool again
Cemberk Sep 29, 2026
07376af
Check that the GLM-5.1 1P/1D cards keep each pool TP8
Cemberk Sep 29, 2026
b19a665
Give rixl/deepep the recipe's AITER and decode cudagraph settings
Cemberk Sep 29, 2026
8370e44
Size the MI300X Kimi-K3 disagg KV cache to fit beside its weights
Cemberk Sep 29, 2026
ec39445
File each sweep result under its own concurrency, and fail the cells …
Cemberk Sep 30, 2026
b3c5702
Capture only Kimi-K3 decode's real batch sizes, and fail fast on a wo…
Cemberk Sep 30, 2026
4c2dd5e
Restore the Kimi-K3 recipe's KV cache and decode capture sizes
Cemberk Sep 30, 2026
761d477
Request the model under the name the recipe serves it as
Cemberk Sep 30, 2026
b210fe4
Build every Kimi-K3 card from one Dockerfile, with AITER from source
Cemberk Sep 30, 2026
08c45d3
Name the m2m cluster's GPU for image builds
Cemberk Sep 30, 2026
7d60f60
Drop the cluster profile; the allocation defaults are madengine's pre…
Cemberk Sep 30, 2026
c652f31
Fail an SGLang disagg job whose nodes failed or that produced no results
Cemberk Sep 30, 2026
6a7c24b
Read SGLang sweep results cell by cell, and fail the cells that failed
Cemberk Sep 30, 2026
1688cf0
Fail vLLM result cells that printed nothing, NIAH lengths that errore…
Cemberk Sep 30, 2026
6f863d6
Add a docs/ entry point that takes a reader from zero to running and …
Cemberk Sep 30, 2026
e9c4e1e
Fetch the Kimi-K3 AITER commit by SHA; a clone of ROCm/aiter does not…
Cemberk Sep 30, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,9 @@

MAD is a platform that consists of curated list of AI models that allow us to run on various GPU architectures seamlessly while tracking performance and generating dashboards for insights.

**Documentation:** start at [docs/README.md](docs/README.md). It takes you from running your first
model to running, configuring and extending the multinode inference workloads.

## Blueprints

This repository provides state-of-the-art deep learning recipes for training, inference and easy deployment on AMD Instinct GPUs.
Expand Down
5 changes: 5 additions & 0 deletions benchmark/kimi_k3/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,11 @@ MAD supports Kimi-K3 day-0 inference across **three** serving frameworks on AMD
| **SGLang** | `pyt_sglang_kimi-k3` | `lmsysorg/sglang-rocm:rocm720-mi35x-k3-20260727` | [SGLang cookbook](https://docs.sglang.io/cookbook/autoregressive/Moonshotai/Kimi-K3) |
| **ATOM** | `pyt_atom_kimi-k3` | `rocm/atom-dev:rocm7.2.4_ubuntu24.04_py3.12_pytorch2.10.0_20260727_kimi_k3` | ROCm ATOM |

> **On MI300X (gfx942)?** All four models above carry `skip_gpu_arch: gfx942` —
> the ~1.5 TB checkpoint does not fit a single 8-GPU MI300X node, so K3 needs
> multi-node sharding there. See [`mi300x/`](mi300x/README.md) for the 2-node
> colocated and 4-node prefill/decode-disaggregated recipes.

## Hardware requirements

- **8x MI350X or MI355X** (TP8)
Expand Down
5 changes: 5 additions & 0 deletions benchmark/kimi_k3/mi300x/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
# Kimi-K3 inference on AMD Instinct MI300X (gfx942)

This guide moved to [docs/kimi-k3.md](../../../docs/kimi-k3.md), which covers Kimi-K3 on MI300X and
MI355X, colocated and disaggregated: the cards, the image, site settings, known failure modes and
results.
4 changes: 4 additions & 0 deletions docker/sglang_disagg_inference.ubuntu.amd.Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,10 @@ WORKDIR /sgl-workspace

RUN pip install --upgrade sglang-router

# Runtime deps of the disagg launcher and proxy. Baked in so the launcher scripts
# do not have to install them on every node of every run.
RUN pip install py-spy flask pyyaml

WORKDIR /sgl-workspace/mori

ARG MORI_COMMIT="158c7e8335a0b19b3f1f422ff134d7869252135e"
Expand Down
9 changes: 8 additions & 1 deletion docker/vllm_disagg_inference.ubuntu.amd.Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -225,6 +225,13 @@ ENV _ROCM_DIR=/opt/rocm \
_RIXL_BRANCH=f33a5599 \
_RIXL_INSTALL_DIR=/usr/local/RIXL/install \
_NIXLBENCH_INSTALL_DIR=/usr/local/RIXL
# rocSHMEM is pinned like every other source here. It tracked rocm-systems `develop`
# until rocm-systems 16dc5673 (2026-09-25, "replace hip atomic builtins with scoped
# counterparts") made projects/rocshmem/src/atomic.hpp use scoped atomics on float and
# double, which this image's ROCm 7.2 compiler rejects ("address argument to atomic
# operation must be a pointer to integer or pointer"); every build of this image failed
# from then on. This is that commit's parent, the rocSHMEM the image last built with.
ARG ROCSHMEM_REF=9c43b23229cf2f835b838f72c0d1af3d4876167c
RUN if [ "${WITH_NIXL}" != "1" ]; then \
echo "WITH_NIXL=${WITH_NIXL}: skipping UCX/RIXL/rocSHMEM/DeepEP (MoRI-EP + base DeepEP only)"; \
else set -e && \
Expand Down Expand Up @@ -257,7 +264,7 @@ RUN if [ "${WITH_NIXL}" != "1" ]; then \
--config-settings=setup-args="-Ddisable_gds_backend=true" . && \
# rocSHMEM (DeepEP dep)
cd /tmp && git clone --no-checkout --filter=blob:none https://github.com/ROCm/rocm-systems.git && \
cd rocm-systems && git sparse-checkout set --cone projects/rocshmem && git checkout develop && \
cd rocm-systems && git sparse-checkout set --cone projects/rocshmem && git checkout "${ROCSHMEM_REF}" && \
mkdir -p /tmp/rocshmem-build && cd /tmp/rocshmem-build && \
/tmp/rocm-systems/projects/rocshmem/scripts/build_configs/all_backends \
-DUSE_EXTERNAL_MPI=OFF -DGPU_TARGETS="${GFX_COMPILATION_ARCH}" && \
Expand Down
343 changes: 343 additions & 0 deletions docker/vllm_kimi_k3.ubuntu.amd.Dockerfile

Large diffs are not rendered by default.

123 changes: 123 additions & 0 deletions docs/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,123 @@
# MAD documentation

MAD (Model Automation and Dashboarding) is a curated list of AI models that run on various GPU
architectures while tracking performance and generating dashboards for insights. It provides deep
learning recipes for training, inference and deployment on AMD Instinct GPUs.

You do not run MAD directly. You run it with **madengine**, a command-line tool from
[ROCm/madengine](https://github.com/ROCm/madengine). madengine reads the model definitions in this
repository, builds a Docker image for each model, runs the model's script inside a container (or
submits it to a SLURM cluster), and writes the results to a CSV file.

This folder is a guide from first contact to writing your own workloads. Start at step 1 of the
learning path and stop when you have what you need.

## Who should read what

| You want to | Read |
|---|---|
| Run one model on one GPU host | [Getting started](getting-started.md) |
| Add a model, a Dockerfile or a script | [Adding a model](adding-a-model.md) |
| Understand multinode inference (disaggregated prefill/decode, colocated) | [Multinode overview](multinode-overview.md) |
| Run a multinode card on a SLURM cluster | [Running multinode workloads](multinode-running.md) |
| Change a setting and know which layer wins | [Configuration](configuration.md) |
| Read or compare benchmark results | [Benchmarks and results](benchmarks-and-results.md) |
| Look up every knob of one launcher or one model | [vLLM disaggregated](vllm-disagg.md), [SGLang disaggregated](sglang-disagg.md), [Kimi-K3](kimi-k3.md) |

## Learning path

1. **Run a single-node model.** Install madengine, run a model with one command, and learn what
happens during a run and where the results go.
[getting-started.md](getting-started.md)
2. **Add a model.** Write a model card, a Dockerfile and a run script, and report performance in
the format madengine expects. [adding-a-model.md](adding-a-model.md)
3. **Learn the multinode concepts.** Disaggregated prefill/decode versus colocated multinode,
launchers, KV connectors, EP backends and node topology.
[multinode-overview.md](multinode-overview.md)
4. **Run a multinode workload.** Prerequisites, running a card through madengine, `sbatch` or
`salloc`, reading logs and results, and fixing failures.
[multinode-running.md](multinode-running.md)
5. **Configure.** Every configuration layer, its precedence, and where to change a given setting.
[configuration.md](configuration.md)
6. **Benchmark and read results.** Throughput sweeps, needle-in-a-haystack (NIAH), agentic replay,
and the `perf.csv` schema. [benchmarks-and-results.md](benchmarks-and-results.md)
7. **Use the references.** Full per-launcher and per-model references:
[vllm-disagg.md](vllm-disagg.md), [sglang-disagg.md](sglang-disagg.md),
[kimi-k3.md](kimi-k3.md).

## Page index

| Page | What it covers |
|---|---|
| [README.md](README.md) | This page: what MAD is, the learning path, the glossary and the index. |
| [getting-started.md](getting-started.md) | Prerequisites, installation, running models, tags, timeouts, debugging, model discovery, build versus run, and where results go. |
| [adding-a-model.md](adding-a-model.md) | Every model-card field, Dockerfile resolution, the GPU architecture build argument, run scripts, the performance reporting contract, and multinode cards. |
| [multinode-overview.md](multinode-overview.md) | Concepts: disaggregated prefill/decode versus colocated multinode, launchers, KV connectors, EP backends, topology, architecture diagrams. |
| [multinode-running.md](multinode-running.md) | Running a multinode card through madengine, `sbatch` or `salloc`; logs, results, failure modes, troubleshooting and offline checks. |
| [configuration.md](configuration.md) | Every configuration layer and its precedence, `models.yaml` recipes, connector environment, `cluster.sh`, madengine presets and `--additional-context` keys, `mad-config.yaml`. |
| [benchmarks-and-results.md](benchmarks-and-results.md) | Throughput sweep, NIAH, agentic replay, the `perf.csv` schema and status semantics. |
| [vllm-disagg.md](vllm-disagg.md) | Full reference for [`scripts/vllm_dissag`](../scripts/vllm_dissag). |
| [sglang-disagg.md](sglang-disagg.md) | Full reference for [`scripts/sglang_disagg`](../scripts/sglang_disagg). |
| [kimi-k3.md](kimi-k3.md) | Kimi-K3 on MI300X and MI355X, colocated and disaggregated. |

## Blueprints

These are the supported model families, with the documentation that ships next to each one.

| Blueprint | Description | Models |
|-----------|-------------|--------|
| [Kimi-K3 inference (vLLM / SGLang / ATOM)](../benchmark/kimi_k3/README.md) | Kimi-K3 (2.8T) day-0 inference on MI350X/MI355X across three frameworks. See also [kimi-k3.md](kimi-k3.md). | [moonshotai/Kimi-K3](https://huggingface.co/moonshotai/Kimi-K3) |
| [xDiT diffusion inference](../benchmark/xdit/README.md) | Diffusion Transformer inference using xDiT | FLUX.1, FLUX.1 Kontext, FLUX.2, FLUX.2 Klein, HunyuanVideo, HunyuanVideo 1.5, LTX-2, Stable Diffusion 3.5, Wan 2.1, Wan 2.2, Z-Image Turbo |
| [JAX MaxText training](../benchmark/jax_maxtext/README.md) | Train LLMs on AMD Instinct GPUs using JAX MaxText | Llama 2 7B/70B, Llama 3/3.1 8B/70B, Llama 3.1 405B, Llama 3.3 70B, DeepSeek-V2-lite 16B, Mixtral-8x7B |
| [vLLM inference](../benchmark/vllm/README.md) | LLM inference with vLLM on AMD Instinct GPUs | DeepSeek-R1, gpt-oss-20b/120b, Kimi-K3, Llama-2-70b, Llama-3.1-8b/405b, Llama-3.3-70b, Llama-4-Scout/Maverick, Mixtral-8x7b/8x22b, Phi-4, Qwen3-8b/32b/30b-a3b/235b-a22b |
| [SGLang inference](../benchmark/sglang/README.md) | LLM inference with SGLang on AMD Instinct GPUs | DeepSeek-R1-Distill-Qwen-32B, Kimi-K3 |
| PyTorch training | Train LLMs on AMD Instinct GPUs using AMD's Primus. Primus notes are in [benchmark/primus/README.md](../benchmark/primus/README.md). | Llama 2/3/3.1/3.2/3.3/4, GPT-OSS 20B/120B, Qwen2/2.5/3, Flux, SDXL, DLRM, and others |
| [PyTorch inference](../benchmark/pytorch_inference/README.md) | Inference recipes for multimodal, video and vision transformer models | Mochi video, Chai-1, CLIP (ViT-B-32), Wan2.1, Janus-Pro-7B, HunyuanVideo |
| Megatron-LM training | Train LLMs on AMD Instinct GPUs using ROCm Megatron-LM | Llama 2 7B/70B, Llama 3/3.1 8B/70B, Llama 3.3 70B, DeepSeek-V2-lite, DeepSeek-V3, Mixtral 8x7B/8x22B, Qwen 2.5 7B/72B |
| MPT-30B training (llm-foundry) | LLM training for Mosaic Pretrained Transformer (MPT) models using llm-foundry | MPT-30B |
| PyTorch PEFT/FSDP fine-tuning | Fine-tuning a Hugging Face model with the LoRA approach and the FSDP strategy | Llama-2-70b-chat-hf |
| [Large EP microbenchmark](../scripts/large-ep-benchmark/README.md) | MoE large expert parallelism with MoRI-EP and DeepEP communication microbenchmarks | No specific models |
| [vLLM disaggregated P/D inference](../scripts/vllm_dissag/README.MD) | Distributed inference with prefill/decode disaggregation in vLLM (Default, MoRI EP, DeepEP). See also [vllm-disagg.md](vllm-disagg.md). | DeepSeek-R1, DeepSeek-V3, DeepSeek-V3-5layer, amd-Llama-3.3-70B-Instruct-FP8-KV, Llama-3.1-405B-Instruct-FP8-KV, gpt-oss-120b |
| [SGLang disaggregated P/D inference](../scripts/sglang_disagg/README.MD) | Distributed inference with prefill/decode disaggregation in SGLang (MoRI IO, Mooncake). See also [sglang-disagg.md](sglang-disagg.md). | Llama-3.1-8B, Qwen3-32B, Llama-3.3-70B-FP8, Llama-3.1-405B-FP8, Mixtral-8x7B, DeepSeek-V3, DeepSeek-R1 |
| [SGLang disaggregated P/D inference with WideEP/LargeEP](../scripts/sglang_disagg/README.MD) | Distributed inference with prefill/decode disaggregation in SGLang with WideEP/LargeEP | DeepSeek-V3, DeepSeek-R1 |
| [KVCache Transfer Bench](../scripts/kvcache_transfer_bench/README.md) | Inter-node transfer benchmark | No specific models |

The root [README](../README.md) links the PyTorch training, Megatron-LM, MPT-30B and PEFT/FSDP
blueprints to files that are not in this checkout, so those rows have no link here.

## Glossary

Terms are listed in the order you meet them.

| Term | Meaning |
|---|---|
| **madengine** | The CLI from [ROCm/madengine](https://github.com/ROCm/madengine) that discovers, builds, runs and reports MAD models. `pip install -r requirements.txt` in this repository installs it. Its main commands are `discover`, `build`, `run`, `report` and `database`. |
| **Model card** (or **card**) | One JSON object in a `models.json` file that describes one workload: its name, Dockerfile, script, GPU count, tags, arguments and, for multinode work, its launcher, node count and environment. Cards live in `scripts/<dir>/models.json`. A directory can also generate cards in Python with `get_models_json.py`. See [adding-a-model.md](adding-a-model.md). |
| **Tag** | A label in a card's `tags` list, such as `pyt`, `vllm` or `inference`. `madengine run --tags X` selects every card whose name or tags match `X`. See [getting-started.md](getting-started.md#selecting-models-with-tags). |
| **Recipe** | The per-model serving flags and environment for a multinode workload, kept apart from the card. For vLLM disaggregated serving it is the model's entry in [`scripts/vllm_dissag/models.yaml`](../scripts/vllm_dissag/models.yaml); SGLang has [`scripts/sglang_disagg/models.yaml`](../scripts/sglang_disagg/models.yaml). See [configuration.md](configuration.md). |
| **Launcher** | Two related meanings. (1) The batch script a multinode card runs, for example `run_xPyD_models.slurm` or `run_multinode.slurm`. (2) The value of a card's `distributed.launcher`, which tells madengine how to start the workload (`torchrun`, `vllm`, `sglang`, `slurm_multi` and others). |
| **slurm_multi** | The madengine launcher used by every MAD multinode inference card. madengine writes a wrapper SBATCH script that exports the card's `env_vars`, then runs the card's own `.slurm` script on the head node. That script starts the per-node Docker containers itself with `srun`. The hyphenated `slurm-multi` is accepted as an alias. |
| **Connector** | In disaggregated serving, the component that moves the KV cache from the prefill server to the decode server. vLLM uses `rixl` (NixlConnector) or `moriio` (MoRIIOConnector). SGLang uses MoRI IO or Mooncake as its transfer backend. See [multinode-overview.md](multinode-overview.md). |
| **EP backend** | The all-to-all communication library used for wide expert parallelism (wideEP) in mixture-of-experts models: `mori` (MoRI-EP) or `deepep` (DeepEP). In vLLM each connector pairs with its own backend: `moriio` with `mori`, `rixl` with `deepep`. |
| **xP/yD** | The shape of a disaggregated run: `xP` prefill nodes and `yD` decode nodes. The job needs `xP + yD` nodes. For example `1P/1D` is 2 nodes. |
| **`--additional-context`** | A JSON string passed to `madengine build` or `madengine run` that adds or overrides configuration: `gpu_vendor`, `guest_os`, `docker_env_vars`, `docker_build_arg`, `env_vars`, a `slurm` block, and more. `--additional-context-file` reads the same JSON from a file, and `--additional-context` is merged over it key by key. |
| **Build manifest** | `build_manifest.json`, written by `madengine build` (and by the build half of `madengine run`). It records each built image, the card it was built for, the registry image if one was pushed, and the context. `madengine run --manifest-file build_manifest.json` runs from it without building again. |
| **perf.csv** | The results table. `madengine run` appends one row per model (or one row per result for cards with `multiple_results`), including failed runs as `FAILURE` rows. Change the file name with `-o`. See [benchmarks-and-results.md](benchmarks-and-results.md). |
| **skip_gpu_arch** | A card field listing GPU architectures (comma-separated, for example `gfx942`) the card must not run on. madengine checks it before running, locally or on SLURM, and writes a `SKIPPED` row instead. `--disable-skip-gpu-arch` turns the check off. |
| **MAD_SYSTEM_GPU_ARCHITECTURE** | The host GPU architecture, for example `gfx942` (MI300X) or `gfx950` (MI355X). madengine sets it as an environment variable inside the container, and passes it as a Docker build argument so a Dockerfile can build for one architecture. See [adding-a-model.md](adding-a-model.md#the-mad_system_gpu_architecture-build-argument). |

## madengine documentation

madengine has its own documentation in the
[`docs/` folder of ROCm/madengine](https://github.com/ROCm/madengine/tree/main/docs). These pages
are the ones MAD users need most.

| Page | What it covers |
|---|---|
| [installation.md](https://github.com/ROCm/madengine/blob/main/docs/installation.md) | Installing madengine, setting up the MAD package, testing Docker GPU access for ROCm and CUDA, and fixing import, permission and ROCm-path problems. |
| [usage.md](https://github.com/ROCm/madengine/blob/main/docs/usage.md) | The five commands, the three model discovery methods, the build and run workflows, timeouts, debugging, profiling, reporting and MongoDB upload. |
| [cli-reference.md](https://github.com/ROCm/madengine/blob/main/docs/cli-reference.md) | Every option of `discover`, `build`, `run`, `report` and `database`, with defaults, examples and exit codes. |
| [configuration.md](https://github.com/ROCm/madengine/blob/main/docs/configuration.md) | Every `--additional-context` key: defaults, log error pattern scan, pinned image digests, Docker environment, build arguments, mounts, timeouts, Kubernetes and SLURM blocks, profiling, pre/post scripts, credentials and configuration priority. |
| [launchers.md](https://github.com/ROCm/madengine/blob/main/docs/launchers.md) | Each distributed launcher (`torchrun`, DeepSpeed, Megatron-LM, TorchTitan, Primus, vLLM, SGLang, SGLang disaggregated, `slurm_multi`), with a comparison matrix and troubleshooting. |
| [distributed-config.md](https://github.com/ROCm/madengine/blob/main/docs/distributed-config.md) | How distributed inference configuration splits by scope (the run, the site, the model, the measurement) across madengine `--config`, `cluster.sh`, `models.yaml`, `configs/*.yaml` and `mad-config.yaml`. |
| [deployment.md](https://github.com/ROCm/madengine/blob/main/docs/deployment.md) | Deploying workloads to Kubernetes and SLURM: workflow, configuration examples and troubleshooting. |
Loading