Any model → UAII IR → any hardware.
UAII is a modular C++ inference runtime: load a model into a common intermediate representation, optimize and schedule it, then execute on CPU or GPU backends through one session API. It is built as an execution platform—loaders, operators, backends, schedulers, and storage are pluggable—not a single-family engine hard-wired to one format.
Ready to recommend today: Hugging Face decoder-only folders (*ForCausalLM / GPT-2 / Mixtral, layout auto-detect) and GGUF blk.*, with a default laptop RAM dial (pin trunk, stream the rest — kimi-style). Vision / SSM / RWKV / HF MLA are still roadmap — see docs/huggingface_support.md and upgrade.md.
| Language | C++17 |
| Build | CMake 3.20+ |
| Stable ABI | C API 0.3.0 (uaii_capi) |
| License | MIT |
- Ingest HF decoder directories, GGUF, Safetensors, ONNX, MLX (weights + config), or PyTorch export sidecars into UAII IR.
- Validate & plan the graph (shapes, dtypes, fusion, memory reuse, optional disk plan cache).
- Execute via a
Sessionon a chosen backend, with competitive CPU GEMM, in-memory quantized MatMul, KV-cache generation for HF / GGUF transformers, and optional weight streaming. - Integrate through the
uaiiCLI (pull/generate/chat), the C ABI shared library, or the Python SDK.
GGUF / Safetensors / ONNX / MLX / PyTorch
│
▼
┌──────────┐
│ UAII IR │ validate · fuse · plan
└────┬─────┘
│
▼
┌──────────┐
│ Session │ KV cache · quant GEMM · streaming
└────┬─────┘
│
┌─────────────┼─────────────┐
▼ ▼ ▼
CPU CUDA* Metal* / Vulkan* / WebGPU* / ROCm*
(oneDNN / cuBLASLt
OpenBLAS / + kernels
tiled ref)
*GPU paths require the matching UAII_WITH_*=ON build and a present device. Capability strings and uaii doctor report host-fallback honestly—there is no silent “GPU name, CPU math.”
Absolute kernel microbenchmarks from in-tree uaii_bench (schema uaii_bench/v3). Full methodology: docs/benchmarks.md.
Published sample: local WSL2 · Intel Core i9-14900HX · 32 threads · median of 21 trials.
JSON: benchmarks/results/local_wsl.json · linked: ref, openblas
| Workload | Result |
|---|---|
| f32 GEMM openblas 256³ / 512³ / 1024³ | 141 / 284 / 425 GFLOP/s |
| f32 GEMM ref-tiled 256³ / 512³ / 1024³ | 7.7 / 11.8 / 15.1 GFLOP/s |
| STREAM triad (~256 MiB) | 16.6 GB/s |
| Attention B1 H8 S512 D64 (e2e, ref GEMM) | 59.7 ms median |
| Session 8×512 MatMul+ReLU stack | 2.77 ms median |
| Q4_0 weights (2048×4096) | 4.5 MiB vs 32 MiB f32 (7.11×) |
# WSL
sudo apt install -y libopenblas-dev libdnnl-dev # vendors
TRIALS=21 bash scripts/run_bench_wsl.sh
# Native
cmake -S . -B build -DUAII_WITH_OPENBLAS=ON -DUAII_WITH_ONEDNN=ON -DUAII_BUILD_BENCHMARKS=ON
cmake --build build --target uaii_bench --parallel
./build/benchmarks/uaii_bench --suite all --providers all --trials 21 --jsonCI benchmarks job builds with oneDNN/OpenBLAS when available, uploads JSON artifacts, and prints a job summary — see docs/benchmarks.md for how to cite and rerun.
New here? → TRY.md (5-minute path).
| Format | Role |
|---|---|
| Hugging Face directory | First-class: config.json + tokenizer* + *.safetensors (± shards); Llama/Mistral/Qwen/Gemma/Phi-class allowlist; RoPE, GQA, tied emb — see docs/huggingface_support.md |
| GGUF | Any architecture with llama.cpp-style blk.* decoder tensors (not a Llama-only allowlist); arch-prefixed metadata (qwen2.*, gemma.*, …); RMSNorm, QKV, RoPE, Attention + KV, dense SwiGLU/GELU or MoE (ffn_*_exps + router), tied embeddings |
| Safetensors | Single-file weight graphs / non-directory layouts → UAII IR |
| ONNX | Import to IR (companion .uaii.json or ONNX proto when enabled) |
| MLX | Directory with config.json + .safetensors (weights + config, not the Apple MLX runtime) |
| PyTorch | .pt / .pth via exported .onnx or .uaii.json sidecar (UAII_WITH_LIBTORCH reserved for TorchScript) |
Convert anything the loader registry accepts:
uaii convert model.gguf -o model.uaii.json
uaii convert model.onnx -o model.uaii.json- In-memory GGUF block quants without full f32 unpack:
Q4_0,Q4_1,Q5_0,Q5_1,Q8_0,Q2_K–Q6_K - Pack/unpack helpers: F16, BF16, INT8, INT4, NF4, MXFP4
- Session policy:
compute_dtype(F32/F16),keep_quantized_weights - Unsupported / IQ* GGUF types fail closed (no silent Ones weights)
See docs/gguf_support.md.
- Prefill + greedy decode:
Session::generate/uaii_session_generate - First-class KV cache (per-layer K/V, context limit from metadata, fail-closed overflow)
- Long-sequence awareness;
UAII_MAX_LAYERSoptional layer cap - Tokenizers: BPE, SentencePiece (
UAII_WITH_SENTENCEPIECE), GGUFtokenizer.ggml.*, plus SimpleTokenizer for demos
| Backend | Without vendor SDK | With UAII_WITH_*=ON + device |
|---|---|---|
| CPU | Full f32 kernels; tiled/ref GEMM | + oneDNN / OpenBLAS when linked |
| CUDA | Host-fallback executable | Device memory, cuBLASLt, elementwise/norm kernels; Attention may host-fallback (advertised) |
| Metal | Host-fallback | Shared buffers; MatMul / Add / RMSNorm via runtime MSL |
| Vulkan | Host-fallback | Device buffers; Add via compute when available; honest caps for remaining ops |
| WebGPU | Host-fallback | Buffer path when headers present; compute still limited |
| ROCm | Host-fallback | HIP memory; MatMul (rocBLAS); Add / RMSNorm HIP kernels |
CPU GEMM provider is selected at runtime (ref / onednn / openblas); uaii doctor prints the active provider. Details: docs/backend_support.md.
- Graph validator, JSON + binary IR (
.uaii.json/.uaii) - Planner: op fusion, memory reuse, storage plan, disk plan cache
- Weight streaming with double-buffer host staging; CUDA async H2D overlap when native
- Chrome-trace profiler, benchmark CLI, cache status/clear
- Plugin operator ABI (
Negexample) and plugin discovery - Fail-closed defaults (
weight_init=none)
- C++17 compiler (MSVC 2019+, Clang 10+, or GCC 9+)
- CMake 3.20+
- Ninja (recommended) or your platform generator
- Git
Optional: CUDA Toolkit, Vulkan SDK, ROCm, oneDNN, OpenBLAS, SentencePiece, Node.js (docs site), Python 3.10+ (SDK).
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release \
-DUAII_BUILD_TESTS=ON -DUAII_BUILD_PLUGINS=ON
cmake --build build --config Release --parallelOn Windows with Ninja + MinGW/MSVC, the same commands apply; multi-config generators place binaries under build/Release/.
Paths below assume a single-config build (build/uaii or build/libs/uaii-cli/uaii). Adjust for your generator.
uaii version
uaii doctor
uaii doctor --load-plugins
# IR
uaii validate examples/ir/toy_mlp.uaii.json
uaii inspect examples/ir/toy_mlp.uaii.json
uaii graph examples/ir/toy_mlp.uaii.json --format text
# Execute
uaii run --demo toy_mlp
uaii run --demo tiny_block
uaii run --demo gguf
uaii run --demo parity
uaii run examples/ir/toy_mlp.uaii.json \
--weight-init ones \
--input x=1,2,3,4 \
--output y_prob
# Tokenize
uaii tokenize encode hello world
uaii tokenize encode "hello" --bpe vocab.json --merges merges.txt
# Optimize / profile
uaii run --demo optimize
uaii profile --demo --output uaii_profile.json
uaii benchmark --demo
uaii cache statusSeparate from the public website/ docs site. This is an operator console over the uaii CLI (Node + browser) — not a signed “download and double-click” consumer app. Real GGUF chat with sampling controls, models, doctor, benches, OpenAI /v1:
# Build uaii first, then:
cd dashboard
npm run install:all
npm run build && npm start # → http://127.0.0.1:8787
# Self-host: UAII_DASH_BIND=0.0.0.0 UAII_DASH_TOKEN=secret npm startDocs: TRY.md · website /docs/dashboard/ · dashboard/README.md.
ctest --test-dir build -C Release --output-on-failure| Command | Description |
|---|---|
uaii doctor |
Environment, modules, GEMM provider, backends, plugins |
uaii validate <path> |
Validate a UAII IR graph |
uaii inspect <path> |
Tensors, nodes, metadata |
uaii graph <path> [--format text|dot|json|plan] |
Dump or visualize the graph |
uaii convert <model> -o <out.uaii.json> |
GGUF / Safetensors / ONNX / MLX / PyTorch sidecar → IR |
uaii tokenize encode|decode … |
Simple / BPE / SentencePiece / GGUF tokenizer |
uaii generate … |
HF / GGUF / --demo; --backend auto (GPU if present, else CPU); sampling; --json / --stream |
uaii chat --jsonl … |
Warm chat worker for the Operator UI |
uaii run … |
Run IR or built-in demos (--backend, --weight-init, …) |
uaii profile |
Chrome-trace JSON |
uaii benchmark |
Timing harness |
uaii cache |
Plan-cache status / clear |
uaii help / uaii version |
Help and version |
Global options: --config <toml>, --log-level <level>, --no-color, --load-plugins.
Built-in demos include toy_mlp, tiny_block, gguf, safetensors, moe, parity, optimize, streaming, profile, and quant.
| Option | Default | Meaning |
|---|---|---|
UAII_BUILD_TESTS |
ON |
Unit / smoke tests |
UAII_BUILD_PLUGINS |
ON |
Example plugins |
UAII_BUILD_PYTHON |
OFF |
pybind11 extension |
UAII_WARNINGS_AS_ERRORS |
OFF |
-Werror / /WX |
UAII_WITH_ONEDNN |
OFF |
Intel oneDNN CPU GEMM |
UAII_WITH_OPENBLAS |
OFF |
OpenBLAS CPU GEMM |
UAII_WITH_CUDA |
OFF |
CUDA device path (cuBLASLt + kernels) |
UAII_WITH_METAL |
OFF |
Metal device path |
UAII_WITH_VULKAN |
OFF |
Vulkan device path |
UAII_WITH_WEBGPU |
OFF |
WebGPU device path |
UAII_WITH_ROCM |
OFF |
ROCm / HIP device path |
UAII_WITH_SENTENCEPIECE |
OFF |
SentencePiece tokenizer |
UAII_WITH_LIBTORCH |
OFF |
LibTorch loader hook |
UAII_WITH_ONNX |
ON |
ONNX loader features |
Useful environment variables: UAII_LOG_LEVEL, UAII_NUM_THREADS, UAII_GEMM, UAII_MAX_LAYERS, UAII_CAPI_PATH, UAII_PLUGIN__DIRS.
Shared library target: uaii_capi. Header: include/uaii/c_api/uaii.h.
Current version: 0.3.0 (pre-1.0). Callers must set uaii_session_options.struct_size after uaii_session_options_init.
Highlights:
- Create / destroy sessions, set inputs, run, read outputs
uaii_session_generatefor token-in / token-out generation- Options: backend, fusion, streaming, profiler,
compute_dtype,keep_quantized_weights,max_context - Fail-closed weight init by default
Stability rules: docs/c_api_stability.md.
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release --parallel
python bindings/python/scripts/bundle_native.py --build-dir build
pip install -e bindings/python
python examples/python/load_run_profile.pyIf the native library is not found automatically:
# Windows PowerShell
$env:UAII_CAPI_PATH = "C:\path\to\uaii_capi.dll"Wheel builds are available via .github/workflows/wheels.yml (tag or manual dispatch). More detail: bindings/python/README.md.
Default file: configs/uaii.toml. Environment overlays override file keys (for example UAII_LOG_LEVEL → log.level).
Dynamic libraries export uaii_plugin_get_info, uaii_plugin_init, and uaii_plugin_shutdown. The host rejects mismatched UAII_PLUGIN_ABI_VERSION.
See include/uaii/c_api/plugin_abi.h and plugins/example_probe / plugins/example_op.
include/uaii/ Public C++ headers and C ABI
libs/uaii-core/ Errors, logging, config, plugin host
libs/uaii-ir/ Graph, registry, validator, serialize, plan
libs/uaii-runtime/ Session, schedulers, demos
libs/uaii-kernels/ CPU kernels, IGemm, quant GEMM
libs/uaii-backends/ CPU + CUDA / Metal / Vulkan / WebGPU / ROCm
libs/uaii-loaders/ GGUF, Safetensors, ONNX, MLX, PyTorch
libs/uaii-tokenizers/ Simple, BPE, SentencePiece, GGUF wiring
libs/uaii-quant/ Quant formats and GGUF dequant
libs/uaii-memory/ Arena / pool / budget allocator
libs/uaii-storage/ File provider, mmap, streaming weights
libs/uaii-planner/ Fusion, memory/storage plans, cache
libs/uaii-profiler/ Chrome-trace profiler
libs/uaii-capi/ Stable C ABI shared library
libs/uaii-cli/ `uaii` command-line tool
bindings/python/ Python SDK
website/ Next.js documentation site
examples/ IR samples, models notes, Python examples
schemas/ FlatBuffers IR contract
configs/ Default TOML
docs/ Architecture and support matrices
tests/ Unit and smoke tests
plugins/ Example plugins
| Document | Contents |
|---|---|
| docs/architecture.md | Module and IR design |
| docs/benchmarks.md | Microbench results + how to reproduce |
| docs/gguf_support.md | GGUF arches, quants, generate |
| docs/huggingface_support.md | HF directory layouts, allowlist, CLI |
| docs/backend_support.md | Backend capability matrix |
| docs/c_api_stability.md | C ABI versioning |
| docs/vision.md | Product mission |
| docs/feature.md | Feature catalog |
| website/ | Static docs site (npm ci && npm run build) |
MIT © 2026 Arjun Shukla and contributors