-
Notifications
You must be signed in to change notification settings - Fork 289
Pull requests: lightseekorg/tokenspeed
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
docs: clarify tokenspeed-scheduler release workflow
#1653
opened Sep 20, 2026 by
lightseek-bot
Contributor
Loading…
feat: support tp shard on MoE shared expert
#1652
opened Sep 20, 2026 by
byshiue
Collaborator
Loading…
perf(moe): fuse FlashInfer NVFP4 routing-map padding initialization
#1644
opened Sep 19, 2026 by
byshiue
Collaborator
Loading…
perf(amd): give gfx1250 MXFP4 decode's spare warp bit to N
#1641
opened Sep 18, 2026 by
jerryyin
Member
Loading…
refactor(kernel): unify PyTorch references under numerics/reference
#1626
opened Sep 17, 2026 by
antiagainst
Member
•
Draft
[Don't Merge][Draft] feat(kimi-k3) DEP + TP o projection on attention
#1610
opened Sep 17, 2026 by
byshiue
Collaborator
Loading…
fix(runtime): aggregate cache flush results across data parallel schedulers
#1608
opened Sep 17, 2026 by
yechank-nvidia
Collaborator
Loading…
perf(sampling): skip unused fallback-index reduction in target-only v…
#1607
opened Sep 17, 2026 by
yechank-nvidia
Collaborator
Loading…
perf(moe): pass unpacked routes to FlashInfer
#1606
opened Sep 17, 2026 by
yechank-nvidia
Collaborator
Loading…
fix(runtime): preserve pending drains on empty memory resume
#1605
opened Sep 17, 2026 by
yechank-nvidia
Collaborator
Loading…
[Speculative Decoding] add synthetic acceptance length
#1601
opened Sep 16, 2026 by
edwingao28
Loading…
fix(cache): allow uneven ordinary cache groups
#1589
opened Sep 16, 2026 by
yechank-nvidia
Collaborator
Loading…
perf(kpool): overlap prefill pooled-cache writes with indexer query preparation
#1588
opened Sep 16, 2026 by
Dovis01
Contributor
Loading…
feat(scheduler): keep ordinary state chunks request-local
#1585
opened Sep 16, 2026 by
dongjiyingdjy
Contributor
Loading…
Add Kimi K3 AgentX Slurm benchmark harness
#1564
opened Sep 15, 2026 by
Xiangyi1996
Collaborator
•
Draft
perf(mla): tokenspeed_mla GB200 decode optimizations (-14.6%)
#1525
opened Sep 12, 2026 by
Dogacel
Collaborator
Loading…
feat(memory): reserve the CUDA-graph pools in the KV cache budget
#1509
opened Sep 12, 2026 by
nperrin-fr
Collaborator
•
Draft
feat(sampling): add the sonic-sampler backend
#1506
opened Sep 11, 2026 by
nperrin-fr
Collaborator
•
Draft
perf(autotune): tune exact decode shapes and refresh GEMM routes
#1421
opened Sep 7, 2026 by
chenht2022
Contributor
Loading…
perf(prefill-graph): share bucket outputs and size the ceiling to a chunk
#1350
opened Sep 1, 2026 by
nperrin-fr
Collaborator
•
Draft
ProTip!
Add no:assignee to see everything that’s not assigned.