SGLang
Fast serving framework for large language and vision models
Only patch tags reach this feed, and every one of them is frontier-model firefighting
◆Recent moves
- 17d ago
Patch fixes GLM 5.2 under disaggregation and FP4 MoE NaNs
Repairs GLM 5.2 IndexShare under both prefill/decode disaggregation and context parallelism, plus NaN outputs from FlashInfer FP4 MoE kernels on long inputs. Model-specific stabilisation across the serving paths that are hardest to get right.
View source ↗ - 2mo ago
Patch cherry-picks twelve DeepSeek V4 stability fixes
Twelve fixes backported onto the release branch, mostly DeepSeek V4: garbled single-token decode on B200/B300 from a scale-packing path, and a sliding-window allocator assertion that crashed EAGLE/MTP disaggregated decode around 2000 requests. Failures that appear only under sustained load.
View source ↗ - 3mo ago
Patch bumps FlashInfer to fix its JIT cubin downloader
A single dependency bump resolving a FlashInfer JIT cubin download failure. The smallest possible release, and another instance of the kernel layer driving patch traffic.
View source ↗