Repository navigation
perf: incremental candidate counting for disconnected training - #42
Conversation
c92c053 to
6627325
Compare
6627325 to
25b3773
Compare
|
Added a gigatoken-inspired commit (marcelroed/gigatoken keys all training work on unique words, never occurrences): the incremental update loop now merges and recounts each distinct subgraph object once instead of once per occurrence, sharing the result across duplicates. Zipf corpus (20k words, 2k-word vocab, 300 merges), byte-identical digests:
138 tests pass, ruff clean. |
fb7b73f to
615fff5
Compare
|
Reworked per the review suggestion: instead of deduping inside
Tie-break is preserved exactly (uniques kept in first-occurrence order ⇒ candidate insertion order unchanged); digests byte-identical on all cases, 138 tests pass, ruff clean. This supersedes #45 for the training path — and unlike #45 it also speeds up the all-unique-words case rather than taxing it. |
615fff5 to
25b3773
Compare
The training loop rebuilt Counter(graph.get_merges()) from scratch every step — recounting the whole forest, even the subgraphs the chosen merge never touched. For a disconnected forest (BPE/BNE, where each merge hits only the few words containing the pair) almost all of it is wasted. _train_incremental keeps each subgraph's candidate counts plus a running global total, and after a merge updates only the subgraphs that contained it (located via an index). total is summed in subgraph order, so picking the first max-score candidate reproduces max(Counter(...), key=score) exactly, including tie-breaks — output is byte-identical. Gated to the disconnected case (Tokenizer passes incremental=not connected); connected/draw/verbose keep the plain recount loop, so Boundless/SuperBPE are unchanged. The two loops share _steps() for the progress scaffolding; their cores differ (Counter-rebuild vs total+index update). Trade-off: the persistent count state raises peak memory. Stacked on the bytes-Node change: BPE ~6.6x, BNE ~7.4x over the original baseline. 138 tests pass; digests identical. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
25b3773 to
02b2be8
Compare
The training loop rebuilt
Counter(graph.get_merges())from scratch every step — re-counting and re-hashing every candidate in the whole forest, even the subgraphs the chosen merge never touched. Profiling showed this recount is ~70% of training time. For a disconnected forest (BPE/BNE — each merge hits only the few words containing the pair) almost all of it is wasted._train_incrementalkeeps each subgraph's candidate counts plus a running globaltotal, and after a merge updates only the subgraphs that contained it (located via anindex):Correctness:
totalis summed in subgraph order, so picking the first max-score candidate reproducesmax(Counter(graph.get_merges()), key=score)exactly, including tie-breaks. Output is byte-identical — merge digests unchanged, 138 tests pass.Gated to the disconnected case (
Tokenizerpassesincremental=not connected); connected / draw / verbose keep the plain recount loop, so Boundless and SuperBPE are unchanged (incremental doesn't help when a merge touches most of the few large components).Trade-off: the persistent per-subgraph count state (
comp_counts+index+total) raises peak memory ~3× for the disconnected tokenizers — a deliberate speed-for-memory choice for the dominant BPE/BNE path.🤖 Generated with Claude Code