I build GPU kernels and the CAKE kernel agent. Feel free to reach out!
Tracking all the fresh slices of CAKE served in FlashInfer — grab a quick inference taste, or go forward and take a slice back for training! 🍰
I build GPU kernels and the CAKE kernel agent. Feel free to reach out!
Tracking all the fresh slices of CAKE served in FlashInfer — grab a quick inference taste, or go forward and take a slice back for training! 🍰
FlashInfer: Kernel Library for LLM Serving
SGLang is a high-performance serving framework for large language models and multimodal models.
Building the Virtuous Cycle for AI-driven LLM Systems
Train speculative decoding models effortlessly and port them smoothly to SGLang serving.
Mirage Persistent Kernel: Compiling LLMs into a MegaKernel