vllm-project/production-stack

vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization

View on GitHub ↗Jump to charts ↓Open shareable report →

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 1 hour ago
Added to GitGenius on February 11th, 2025
Created on January 21st, 2025
Open Issues & Pull Requests: 230 (+1)
GitHub issues: Enabled
Number of forks: 516
Total Stargazers: 2,662 (+0)
Total Subscribers: 29 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 15.7 hours
Mean response time: 12.8 days
90th percentile: 20.8 days
Tracked items: 298

Maintainer activity

1 person did triage or write work on this repository in the last 12 months.

Counts unlabeled, assigned, unassigned, milestoned, demilestoned, locked, unlocked over the last 12 months. These are issue and pull request events that require triage or write permission. Commits and code review are not counted. labeled and renamed are excluded because GitHub issue forms record the issue author as the actor. Figures from October 7, 2026. This count is not comparable across projects: each project's automation decides which of these events a person emits.

How this project is maintained

About 13% of issues opened in the past year have never received a reply. 86% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. 68% of issues opened in the past year have been closed, leaving a working backlog. Three people close 70% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 103
New in 7 days: 1
Closed in 7 days: 4
Avg open age: 240 days
Stale 30+ days: 93
Stale 90+ days: 81

Recent activity

Opened in 7 days: 1
Closed in 7 days: 4
Comments in 7 days: 0
Events in 7 days: 2

Top labels

  • feature request (116)
  • bug (109)
  • help wanted (14)
  • good first issue (12)
  • question (9)
  • discussion (7)
  • documentation (5)

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

The vLLM Production Stack is a reference implementation for deploying vLLM inference clusters in Kubernetes environments with community-driven performance optimization. Written primarily in Python, the project provides a complete system for scaling large language model serving from single instances to distributed deployments without requiring application code changes. The repository officially launched on January 22, 2025, and has already established itself as an active project with official documentation and cloud deployment tutorials for major platforms including AWS EKS, Google GCP, Lambda Labs, and Azure.

The core architecture consists of three main components working together. The serving engine runs multiple vLLM instances that host different language models. A request router directs incoming requests to appropriate backends based on routing keys or session IDs to maximize KV cache reuse and improve performance. An observability stack built on Prometheus and Grafana monitors backend metrics through a web dashboard, providing real-time insights into system health and performance.

The project provides step-by-step tutorials covering the complete deployment lifecycle, from Kubernetes environment setup through minimal installation, configuration customization, model loading, launching multiple models, and enabling KV cache offloading with LMCache. Deployment is managed through Helm charts, allowing users to deploy the stack with standard Kubernetes tools. The deployed system exposes the same OpenAI API interface as vLLM, ensuring compatibility with existing applications.

The Grafana dashboard offers comprehensive monitoring capabilities including available vLLM instance counts, request latency distribution, time-to-first-token metrics, active and pending request tracking, GPU KV cache usage percentages, and GPU KV cache hit rates. The router supports multiple deployment patterns including routing to endpoints running different models, session-ID based routing, round-robin routing, and prefix-aware routing currently in development. It also provides automatic service discovery and fault tolerance through the Kubernetes API.

The project hosts bi-weekly community meetings every other Tuesday at 5:30 PM PT and maintains active Slack channels for both the production stack and the related LMCache project.

The 2026 roadmap includes planned features for autoscaling based on vLLM-specific metrics, support for disaggregated prefill processing, and router improvements including more performant implementations using non-Python languages, KV-cache-aware routing algorithms, and enhanced fault tolerance. The project is licensed under Apache License 2.0 and welcomes community contributions. GMI Cloud is listed as a sponsor supporting development and benchmarking efforts.