Milestones
List view
Q4 quality measured against upstream, and a GGUF reader. Kept inside November because the user's instruction puts every usable-level feature there. Not on Gemstone's November critical path (Gemstone defers Q4 and GGUF), so it runs in parallel and never takes a gate slot from TN-M2.
Due by November 26, 2026•0/3 issues closedMoved out of November, with reasons. Paged-attention kernel (#14): a speed kernel; Gemstone's November path uses transformers' sdpa_paged / eager_paged, which run on plain ops. NPU on-device proof (#21): needs devices or user-run steps, and adding agents cannot parallelise it. S7.1 and S7.2 (#18, #19): not on the serving path. S6.7 federation extras (#20): not on the serving path. Performance optimisation: correctness first; optimising before the path agrees would optimise the wrong thing.
Due by March 31, 2027•0/9 issues closedGemstone's November serving path, usable on cpu and mps and agreeing with upstream: transformers continuous batching (ContinuousBatchingManager / generate_batch) with attn_implementation sdpa_paged and eager_paged, which are plain torch ops. Acceptance per #15: two concurrent requests equal sequential (greedy), and the same seed gives the same output with or without batching (sampling). The blocking-op list ships by 11-06. Three days before Gemstone's 2026-11-16. No new paged-attention kernel is needed; that moved to TN-M3.
Due by November 13, 2026•3/10 issues closedStreaming generation on Llama 3.2 and SmolLM2 with q8_0 agreeing with upstream; the gate-blocking defect (#8); S6.5 and S6.1; the milestone record; the AGENTS.md migration. Checkpoint four weeks before the usable release.
Due by October 24, 2026•11/17 issues closed