Sitemap

A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.

Pages

Posts

FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems Permalink

15 minute read

Published:

Introducing FlashInfer-Bench, a benchmark and infrastructure that lets AI systems optimize themselves: standardized GPU workload descriptions via FlashInfer Trace, realistic benchmarks from production LLM deployments, and a seamless path for deploying AI-generated kernels into live serving systems.

portfolio

publications

FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems

Published in MLSys 2026, 2026

A framework that closes the loop between AI-generated GPU kernels and production LLM serving: a unified schema for kernel generation and benchmarking, with deployment paths into SGLang and vLLM.

Recommended citation: Shanli Xing, Yiyan Zhai, Alexander Jiang, Yixin Dong, Yong Wu, Zihao Ye, Charlie Ruan, Yingyi Huang, Yineng Zhang, Liangsheng Yin, Aksara Bayyapu, Luis Ceze, Tianqi Chen. (2026). "FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems." MLSys 2026.
Download Paper

talks

teaching