FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems
Published in MLSys 2026, 2026
FlashInfer-Bench establishes a reproducible pathway for deploying AI-generated GPU kernels into live LLM serving systems. It standardizes GPU workload descriptions through FlashInfer Trace, provides realistic benchmarks derived from production LLM deployments, and integrates kernel generation, benchmarking, and deployment through a unified schema with production systems such as SGLang and vLLM.
Recommended citation: Shanli Xing, Yiyan Zhai, Alexander Jiang, Yixin Dong, Yong Wu, Zihao Ye, Charlie Ruan, Yingyi Huang, Yineng Zhang, Liangsheng Yin, Aksara Bayyapu, Luis Ceze, Tianqi Chen. (2026). "FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems." MLSys 2026.
Download Paper
