vllm-project/aibrix

Cost-efficient and pluggable Infrastructure components for GenAI inference

View on GitHub ↗Jump to charts ↓Open shareable report →

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 17 minutes ago
Added to GitGenius on February 24th, 2025
Created on June 10th, 2024
Open Issues & Pull Requests: 378 (-2)
GitHub issues: Enabled
Number of forks: 721
Total Stargazers: 5,128 (+0)
Total Subscribers: 52 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 6.2 hours
Mean response time: 8.8 days
90th percentile: 14.3 days
Tracked items: 1,043

Maintainer activity

50 people did triage or write work on this repository in the last 12 months.

Counts unlabeled, assigned, unassigned, milestoned, demilestoned, locked, unlocked over the last 12 months. These are issue and pull request events that require triage or write permission. Commits and code review are not counted. labeled and renamed are excluded because GitHub issue forms record the issue author as the actor. Figures from October 7, 2026. This count is not comparable across projects: each project's automation decides which of these events a person emits.

How this project is maintained

About 13% of issues opened in the past year have never received a reply. 37% of open issues come from outside the core team, a mix of external reports and the maintainers' own roadmap. 59% of tracked open issues have had no activity in three months. Only 55% of issues opened in the past year have been closed. Three people close 82% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 340
New in 7 days: 13
Closed in 7 days: 22
Avg open age: 209 days
Stale 30+ days: 301
Stale 90+ days: 257

Recent activity

Opened in 7 days: 10
Closed in 7 days: 20
Comments in 7 days: 11
Events in 7 days: 28

Top labels

  • area/gateway (215)
  • kind/bug (178)
  • kind/feature (160)
  • priority/critical-urgent (157)
  • priority/important-soon (105)
  • kind/enhancement (103)
  • help wanted (90)
  • area/autoscaling (76)

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

AIBrix is an open-source infrastructure framework written in Go that provides essential building blocks for constructing scalable GenAI inference systems. The project is specifically designed to address enterprise needs in deploying, managing, and scaling large language model inference workloads in cloud-native environments.

The framework delivers a comprehensive set of infrastructure components tailored for LLM deployment. Key features include high-density LoRA management for lightweight model adaptations, an LLM gateway and routing system for efficiently directing traffic across multiple models and replicas, and an LLM app-tailored autoscaler that dynamically adjusts inference resources based on real-time demand. The unified AI runtime operates as a versatile sidecar component enabling metric standardization, model downloading, and management across the infrastructure. AIBrix supports distributed inference architecture to handle large workloads across multiple nodes and implements distributed KV cache functionality for high-capacity, cross-engine KV reuse, which is critical for efficient LLM serving.

Cost efficiency is a central design principle, with the framework offering cost-efficient heterogeneous serving capabilities that enable mixed GPU inference while maintaining service level objective guarantees. The system includes proactive GPU hardware failure detection to prevent infrastructure degradation. The project has evolved through multiple releases, with versions ranging from v0.1.0 released in November 2024 through v0.7.0 released in June 2026, demonstrating consistent development and feature expansion.

AIBrix has gained visibility in the broader infrastructure community, with the team delivering keynotes and presentations at major conferences including KubeCon North America 2025, KubeCon China 2025, and KubeCon EU 2025. The project was also featured at the ASPLOS 2025 workshop, positioning it as a significant contribution to system research in LLM inference. The framework is licensed under Apache 2.0 and maintains comprehensive documentation, a blog for release announcements, and an active developer community through Slack channels.