Skip to content

chore: fix benchmark script and bench both implementations - #51

Merged
AmitMY merged 5 commits into
mainfrom
chore/bench-both-impls
Jul 23, 2026
Merged

AmitMY merged 5 commits into
mainfrom
chore/bench-both-impls

Conversation

@AmitMY

@AmitMY AmitMY commented Jul 23, 2026 •

Copy link
Copy Markdown
Contributor

What

benchmarks/bench_tokenizers.py imported complex_tokenization.examples.*, which no longer exists — the script didn't run at all. Rewritten to:

  • Bench BPE, BNE n=4, BoundlessBPE, SuperBPE for the reference (Python) and the fast (Rust) implementation side by side (fast rows appear when the package is installed), plus a HuggingFace baseline row.
  • Each case runs in its own subprocess and reports OS-level peak RSS (ru_maxrss). tracemalloc only tracks Python allocations — Rust rows showed 0.04 MB — and in a shared process the first case absorbed one-time allocations (BPE looked heavier than BNE). Subprocesses fix both and cover Rust memory for real.
  • The corpus is loaded once and handed to subprocesses via a temp file (importing datasets + streaming rows cost seconds per child).
  • Output is a markdown table, printed at the end (HF's progress bars write straight to the OS fd mid-run), ready to paste into PRs/issues.
  • md5 digest of each merge list per row, so output divergence between implementations is visible in every run — the table below shows the fast/ output diverges from reference: n>2 merges never apply via try_merge + bytes tie-break #47 divergence that fix(fast): apply n-ary merges and match the reference tie-break exactly #52 fixes.
  • --samples / --merges for scale. Default calibrated to 2,000 rows / 100 merges → the whole run takes ~30s (measured 26.5s), with the slowest single case (Python Boundless) at ~5s. For reference: Python BPE/BNE handle the full split within 10s; connected mode caps at ~3,700 rows per 10s; --samples 0 --merges 500 is the full-scale run (~25 min, Python Boundless dominating).

Default run output (2,000 rows ≈ 626k chars, 100 merges, ~30s total)

Tokenizer Implementation Time Peak RSS Merges Digest
BPE HuggingFace 2.853s 795 MB 100 b4e28c5ada
BPE reference (Python) 0.517s 122 MB 100 20d930a435
BNE n=4 reference (Python) 1.000s 177 MB 100 6598b73bb3
Boundless BPE reference (Python) 4.939s 140 MB 100 9999e69912
Super BPE reference (Python) 4.203s 144 MB 100 9999e69912
BPE fast (Rust) 0.263s 100 MB 100 05af392274
BNE n=4 fast (Rust) 0.230s 150 MB 100 04c28a32d1
Boundless BPE fast (Rust) 4.213s 214 MB 100 787721b689
Super BPE fast (Rust) 4.330s 293 MB 100 787721b689

fast digests differ from reference pending #52 (tie-break + n>2 merge fix) — once it lands, matching digests here become the parity check on every run.

Full-scale numbers (--samples 0 --merges 500, ~2M words) from this script's previous run

Tokenizer Implementation Time
BPE HuggingFace 3.8s
BPE reference / fast 12.3s / 4.9s
BNE n=4 reference / fast 23.4s / 4.4s
Boundless BPE reference / fast 679s / 179s
Super BPE reference / fast 499s / 216s

🤖 Generated with Claude Code

AmitMY and others added 5 commits July 23, 2026 10:22
benchmarks/bench_tokenizers.py imported complex_tokenization.examples.*,
which no longer exists, so it did not run at all. It now benches BPE,
BNE n=4, BoundlessBPE, and SuperBPE for both the reference and the fast
implementation (when installed) plus HuggingFace, prints a digest of each
merge list so output divergence between implementations is visible (see
issue #47), and takes --samples/--merges so the same script scales from a
smoke run to the full wikitext-2 train split.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
tracemalloc only tracks Python allocations, so fast (Rust) rows showed
~0MB. Each case now runs in its own subprocess and reports ru_maxrss,
which covers Rust memory too and isolates cases from each other.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rges)

Calibrated on wikitext-2: BPE and BNE handle the full split within 10s,
Boundless (connected) caps at ~3700 rows, so 3500 keeps a default run's
slowest case near 10s while still exercising ~10% of the dataset.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Each subprocess was paying seconds to import datasets and stream rows; the
parent now writes the corpus to a temp file the children read. Default
samples calibrated to 2000 so the whole default run (9 cases, 100 merges)
stays around 30 seconds.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AmitMY
AmitMY force-pushed the chore/bench-both-impls branch from 8515541 to f67e5da Compare July 23, 2026 08:22
@AmitMY
AmitMY merged commit 0c9472e into main Jul 23, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant