-
Notifications
You must be signed in to change notification settings - Fork 80
Pull requests: meta-pytorch/MSLK
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
SMEM stage prune for FP8 rowwise H100 autotune (#517)
cla signed
meta-exported
#517
opened Sep 21, 2026 by
njriasan
Contributor
Loading…
Serve the stacked BF16 grouped GEMM from FlyDSL on ROCm
cla signed
module: rocm
#512
opened Sep 8, 2026 by
aryaman-gupta
Contributor
Loading…
Stop the CUDA conda-build CI job from OOMing the worker
cla signed
meta-exported
#511
opened Sep 4, 2026 by
sryap
Contributor
Loading…
Autotune async copy for rowwise FP8 GEMM (#494)
cla signed
meta-exported
#494
opened Aug 21, 2026 by
warrendeng
Loading…
Fix flash_attn_bench crash on the flash_attn.cute (FA4) backend: pass window_size=(-1, -1) instead of None
#397
opened Jun 19, 2026 by
Maurits-de-Groot
Loading…
[CUDA] [PERFORMANCE] Increase speed of bf16bf16bf16_grouped_wgrad via indicating that ElementC is void / nullptr
cla signed
#329
opened Apr 19, 2026 by
benediktjohannes
Loading…
ProTip!
Updated in the last three days: updated:>2026-09-18.