FlashInfer
ReleasesPostsDocumentation Slack Paper GitHub

Posts

  • Sep 22, 2026

    Accelerate LLM Inference with Open Kernels and Smarter Autotuning with FlashInfer v0.7

  • Sep 22, 2026

    MegaMoE in FlashInfer: Fused Expert-Parallel MoE Kernels

  • Sep 22, 2026

    Move Fast. Don’t Break Things.

  • Sep 22, 2026

    FlashInfer Autotuner v2: Tune the Way You Serve

  • Oct 21, 2025

    FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems

  • Mar 10, 2025

    Sorting-Free GPU Kernels for LLM Sampling

  • Dec 16, 2024

    FlashInfer 0.2 - Efficient and Customizable Kernels for LLM Inference Serving

  • Feb 2, 2024

    Cascade Inference: Memory Bandwidth Efficient Shared Prefix Batch Decoding

  • Feb 2, 2024

    Accelerating Self-Attentions for LLM Serving with FlashInfer

subscribe via RSS

FlashInfer

Copyright © 2023-2026, FlashInfer team

  • flashinfer-ai

Introduce Techniques to accelerate Large Language Model Deployment