mizorewww/laya-mlx

Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.

View on GitHub ↗Jump to charts ↓

Summary Information

Updated 6 minutes ago
Added to GitGenius on September 21st, 2026
Created on September 19th, 2026
Open Issues & Pull Requests: 9 (+0)
GitHub issues: Enabled
Number of forks: 410
Total Stargazers: 5,644 (+43)
Total Subscribers: 14 (+2)

Repository Insights (GitGenius)

Median issue/PR response: 4.1 hours
Mean response time: 4.1 hours
90th percentile: 4.1 hours
Tracked items: 1

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 3
New in 7 days: 4
Closed in 7 days: 1
Avg open age: 0 days
Stale 30+ days: 0
Stale 90+ days: 0

Recent activity

Opened in 7 days: 4
Closed in 7 days: 1
Comments in 7 days: 1
Events in 7 days: 2

Top labels

No label distribution available yet.

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

Laya-mlx is a native MLX runtime for executing Laya typed decision models on Apple Silicon hardware with minimal latency.

The tool addresses the need for fast, local inference on structured decision tasks without relying on text generation, PyTorch dependencies, or cloud APIs. It works by providing a runtime environment that executes Laya's typed decision models directly using MLX, Apple's machine learning framework optimized for Silicon chips. This approach keeps inference local, eliminates network overhead, and avoids the computational overhead of general-purpose language models.

Developers working on Apple Silicon machines who need sub-20-millisecond latency for discrete decision tasks should consider this tool. It suits projects that involve classification, structured prediction, or other non-generative inference workloads where model outputs are typed and deterministic rather than free-form text. The tool is particularly relevant for applications that must run entirely on-device without external service dependencies.

The project shows active development with regular commits addressing core functionality and performance. Work has focused on optimizing inference speed and expanding compatibility with different Laya model configurations. The maintainer has demonstrated responsiveness to issues and has incorporated feedback into successive updates.