-
Notifications
You must be signed in to change notification settings - Fork 1.2k
Pull requests: FlashML-org/FreeToken
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
fix(engine): reject fa attention backend below sm_90 at config time
#468
opened Sep 14, 2026 by
SanjayS45
Loading…
feat(gemma4): serve image input on the Gemma-4 releases
feature
New feature or request
multimodal
#467
opened Sep 14, 2026 by
jason-fxz
Collaborator
Loading…
fix(utils): tolerate coalesced msgpack frames in zmq pull queues
#466
opened Sep 14, 2026 by
Cyber-Marty
Loading…
fix(scheduler): match stop strings with incremental decoding
#464
opened Sep 14, 2026 by
taking-lying-flat
Contributor
Loading…
fix(server): loopback-bind the TP rendezvous store, add --dist-port
#461
opened Sep 13, 2026 by
shikhakath
Loading…
feat(models): support deepseek v4.1 with native fp8-fp4 kv storage
#460
opened Sep 13, 2026 by
ArqAlice
Loading…
fix(engine): respect explicit --moe-cache-size with the cpu moe strategy
#457
opened Sep 13, 2026 by
BenMohStem
Loading…
docs(moe): report shared expert banks across replicas
#451
opened Sep 12, 2026 by
kuangrenaigc-stack
Loading…
feat(moe): owner-local expert parallelism and tensor parallelism
#447
opened Sep 11, 2026 by
leiyu1980
Loading…
fix(engine): resolve moe-strategy auto to fused on unified-memory GPUs (GB10)
#445
opened Sep 11, 2026 by
iamanishx
Loading…
fix(models): build o_proj row-parallel in the remaining column-parallel families
#440
opened Sep 11, 2026 by
gberasmus87
Loading…
fix(scheduler): read the prefill chunk cap from the pool, not a snapshot
#439
opened Sep 10, 2026 by
gberasmus87
Loading…
fix(server): resolve JSON Schema refs and union types in tool arguments
#435
opened Sep 10, 2026 by
Retloldin
Loading…
fix(glm4_moe): load compressed-tensors NVFP4 expert checkpoints
#430
opened Sep 10, 2026 by
gberasmus87
Loading…
qwen4_exp: build o_proj row-parallel, as every other family does
#429
opened Sep 10, 2026 by
gberasmus87
Loading…
fix(kernel): order the batch-memcpy probe against the current stream
#414
opened Sep 8, 2026 by
tspeaks
Loading…
fix(models): support llm-compressor NVFP4 MoE export variants
#413
opened Sep 8, 2026 by
Romeriz
Loading…
fix(bench): show how to use custom bandwidth profiles
#412
opened Sep 7, 2026 by
earlvanze
Contributor
Loading…
docs(models): qualify full GLM-5.3 on H200 Serverless
#406
opened Sep 7, 2026 by
earlvanze
Contributor
Loading…
fix(models): serve qwen4_exp FTW checkpoints with PLE streamed from source
#405
opened Sep 7, 2026 by
Eng-Ahmd
Loading…
fix(cpu-moe): honor padded fp8 scale strides, add avx512f tier and float64 parity (builds on #36)
#399
opened Sep 6, 2026 by
ChenyuHeee
Loading…
Previous Next
ProTip!
Type g i on any issue or pull request to go back to the issue listing page.