jundot/omlx

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

View on GitHub ↗Jump to charts ↓Open shareable report →

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 57 minutes ago
Added to GitGenius on August 17th, 2026
Created on February 13th, 2026
Open Issues & Pull Requests: 614 (-19)
GitHub issues: Enabled
Number of forks: 1,987
Total Stargazers: 22,649 (+1)
Total Subscribers: 117 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 6.2 hours
Mean response time: 8.8 days
90th percentile: 14.9 days
Tracked items: 1,849

Maintainer activity

2 people did triage or write work on this repository in the last 12 months.

Counts unlabeled, assigned, unassigned, milestoned, demilestoned, locked, unlocked over the last 12 months. These are issue and pull request events that require triage or write permission. Commits and code review are not counted. labeled and renamed are excluded because GitHub issue forms record the issue author as the actor. Figures from October 7, 2026. This count is not comparable across projects: each project's automation decides which of these events a person emits.

How this project is maintained

About 14% of issues opened in the past year have never received a reply. 100% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Almost all tracked open issues have seen activity in the last three months. 83% of issues opened in the past year have been closed, leaving a working backlog. Three people close 72% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 373
New in 7 days: 41
Closed in 7 days: 695
Avg open age: 16 days
Stale 30+ days: 191
Stale 90+ days: 1

Recent activity

Opened in 7 days: 32
Closed in 7 days: 683
Comments in 7 days: 642
Events in 7 days: 2,419

Top labels

  • outdated (598)
  • has-pr (43)
  • in progress (8)
  • planned (2)
  • needs time (1)
  • priority (1)
  • weekend (1)

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

oMLX is an LLM inference server optimized for Apple Silicon Macs that runs from the macOS menu bar.

The tool addresses the friction of running large language models locally on Apple Silicon by implementing continuous batching and tiered key-value caching across both in-memory and SSD storage. This two-tier caching strategy allows the server to keep frequently used models pinned in memory while automatically swapping larger models to disk on demand. Crucially, KV cache persists across context changes within conversations, meaning past context remains cached and reusable even when requests change, making it practical for extended coding sessions with tools that maintain state across multiple interactions.

oMLX suits developers who want to run local LLMs on their Mac without sacrificing either convenience or control. The menu bar interface provides straightforward management without requiring terminal expertise, while the CLI shim enables integration with terminal commands and Apple Shortcuts for automation. The project targets users who need context limits, model pinning, and intelligent memory management—particularly those working with coding assistants that benefit from persistent context across requests. The macOS app distribution with built-in auto-update lowers the barrier to staying current.

Development activity shows consistent engagement with the project. The maintainer actively responds to user feedback and maintains the codebase. The tool receives regular updates that address both performance improvements and user-reported issues. Documentation is maintained across multiple languages, indicating attention to accessibility for a broader user base. The project includes benchmarking infrastructure to track performance characteristics, suggesting a data-driven approach to optimization decisions.