STRIDE (STRIp-walked Triangulated Residual Integer Decoder) is a per-meshlet, GPU-decodable mesh compression format. One fused CUDA kernel decodes upward of 1.8 G triangles per second on a consumer NVIDIA RTX 3090, producing vertex and index buffers in the layout expected by modern mesh-shader pipelines (DX12 / Vulkan / DirectX12 Ultimate).
Compared head-to-head against AMD's Dense Geometry Format (DGF) on the same hardware, STRIDE is strictly smaller on every test mesh (1.38× to 1.93× fewer bytes) while remaining within 6–17 % of DGF's decode throughput on the meshes of AMD's DGF benchmark suite.
The implementation, full benchmark harness, and reproducibility recipe live in
this repository. The design is documented in the paper at
docs/paper.md (LaTeX
submission bundle under docs/paper_tex/). The end-to-end
reproduction recipe is in ARTIFACT.md.
Eight-mesh corpus at uniform 12-bit per-axis (bbox-relative) quantization, NVIDIA RTX 3090:
| Mesh | Tris | STRIDE BPV | STRIDE GPU (M tris/s) | DGF BPV | DGF GPU (M tris/s) | DGF/STRIDE bytes |
|---|---|---|---|---|---|---|
| fandisk | 13 K | 45.71 | 73 | 71.96 | 438 | 1.57 |
| stanford-bunny | 69 K | 42.54 | 350 | 64.21 | 1,356 | 1.51 |
| horse | 97 K | 41.62 | 564 | 59.66 | 1,464 | 1.43 |
| Monkey | 1.01 M | 33.35 | 1,682 | 49.05 | 2,049 | 1.47 |
| Happy Buddha | 1.09 M | 43.24 | 1,554 | 51.26 | 2,002 | 1.19 |
| Crab | 2.14 M | 42.15 | 1,746 | 48.36 | 2,052 | 1.15 |
| tank | 3.51 M | 33.76 | 1,804 | 49.04 | 2,181 | 1.45 |
| xyz-dragon | 7.22 M | 33.62 | 1,904 | 46.39 | 2,233 | 1.38 |
GPU times are means over 300 individually event-bracketed launches (std in
bench_decode_variance.csv); DGF = our CUDA implementation of the DGF v1.2.0
block format, validated bit-exact against the SDK's DGFLib reference decoder
(bench_dgf_validation.csv). On the 7.2 M-triangle Stanford XYZ RGB Dragon:
STRIDE 15.2 MB / 3.79 ms decode vs. DGF 20.9 MB / 3.23 ms. Full multi-codec comparison (Draco, meshoptimizer,
Corto) is in paper Section 5.
git clone --recurse-submodules https://github.com/maletsden/meshpress.git
cd meshpress
pip install -r requirements.txt
python scripts/download_models.py --paper # auto-fetch 5 of 8 paper meshes
python scripts/bench_stride_decode_sweep.py # GPU decode timing per mesh
python scripts/bench_competitors.py # full multi-codec sweepEncode an OBJ into STRIDE bytes:
from reader import Reader
from encoder import STRIDEEncoder
model = Reader.read_from_file("assets/stanford-bunny.obj")
enc = STRIDEEncoder(max_verts=256, precision_error=0.0005)
out = enc.encode(model)
print(f"{len(out.data)} B, {out.bits_per_vertex:.2f} bpv")Decode on the GPU (CuPy CUDA):
from utils.paradelta_v5_cuda import ParaDeltaV5GpuDecoder
dec = ParaDeltaV5GpuDecoder(data)
v, t = dec.decode_to_host() # numpy float32 positions + uint32 indices| Path | Contents |
|---|---|
encoder/paradelta_v5.py |
Reference STRIDE encoder (bit-exact with paper §3.6). |
encoder/paradelta_codec.py |
prepare_paradelta_arrays — partition + strip-walk + plan. |
encoder/paradelta_v5_nb.py |
Numba-JIT hot kernels for the encoder. |
encoder/_irlp_fit.py |
Constrained integer-rational linear predictor (IRLP) per-mesh fit (paper §3.5): affine weights n0+n1+n2 = 2^K, per-axis canonical fallback. |
utils/paradelta_v5_cuda.py |
Fused CUDA decoder (paper §4). |
utils/meshlet_gen_joint*.py |
Joint-learned meshlet partitioner (paper §3.2). |
utils/meshlet_gen_mo_nb.py, utils/meshlet_gen_greedy_nb.py, utils/meshlet_tunneling_nb.py |
Other meshlet generators. gen_method = meshopt (numba port of meshopt_buildMeshlets, no DLL; smallest bitstream on most meshes), joint_learned (paper default), tunneling, greedy, portfolio[+members] (encode several, keep the smallest; encoder/partition_portfolio.py). |
utils/meshlet_plan_nb.py |
Strip-emit traversal + AMD GTS-style connectivity. |
reader/fast_obj.py |
pandas-backed OBJ reader + .cache.npz sidecar. |
scripts/bench_*.py |
Bench harnesses. Each writes a CSV. |
scripts/verify_*.py |
Round-trip + crack-free verifiers. |
scripts/viz/ |
Paper figure generators (matplotlib). |
bench_cpp/ |
C++/CUDA bench binaries for STRIDE, DGF, and meshopt. See bench_cpp/README.md. |
docs/paper.md |
Paper (markdown). |
docs/paper_tex/ |
LaTeX submission bundle. |
docs/bitstream_spec.md |
Standalone bitstream layout reference. |
models/meshlet_gen_weights.json |
Joint-learned partitioner weights (paper §3.2). |
scripts/reproduce_r1.py |
One-command reproduction of every manuscript number (stages in dependency order); scripts/check_paper_numbers.py is the consistency gate. |
ARTIFACT.md |
Reproducibility recipe (tables + figures). |
legacy/ |
Pre-STRIDE encoders (wavelet, LOD, ellipsoid). Not on the paper path. |
third_party/DGF-SDK |
AMD DGF reference (submodule). |
third_party/corto |
Corto codec (submodule). |
Required for the DGF GPU comparison and the C++ meshopt timings. See
bench_cpp/README.md. Builds against CUDA 12 and either
MSVC (Windows) or gcc / clang (Linux).
If you use STRIDE or MeshPress in your research, please cite the accompanying article (currently under review):
Maletskyi, D., Vyklyuk, Y., Li, F. STRIDE: STRIp-walked Triangulated Residual Integer Decoder for Per-Meshlet GPU Mesh Compression. Under review.
@article{maletskyi_stride,
title = {{STRIDE}: {STRIp}-walked Triangulated Residual Integer Decoder
for Per-Meshlet {GPU} Mesh Compression},
author = {Maletskyi, Denys and Vyklyuk, Yaroslav and Li, Fengping},
note = {Under review}
}The source code and the raw benchmark results of the run the paper reports
(results/full/final/) are archived on Zenodo (concept DOI, always the latest
release): 10.5281/zenodo.20942759.
The evaluation meshes are third-party and are fetched by
scripts/get_data.py instead.
@software{meshpress_stride,
title = {{MeshPress}: {STRIDE} GPU mesh compression},
author = {Maletskyi, Denys and Vyklyuk, Yaroslav and Li, Fengping},
publisher = {Zenodo},
doi = {10.5281/zenodo.20942759},
url = {https://doi.org/10.5281/zenodo.20942759},
version = {1.1.0}
}MIT. See LICENSE.