kyuz0/amd-strix-halo-toolboxes

AMD Strix Halo Llama.cpp Toolboxes is a container system that enables running large language models on AMD Ryzen AI Max "Strix Halo" integrated GPUs.

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 55 minutes ago
Added to GitGenius on November 11th, 2025
Created on July 28th, 2025
Open Issues & Pull Requests: 61 (+0)
Number of forks: 188
Total Stargazers: 1,872 (+0)
Total Subscribers: 62 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 3.8 hours
Mean response time: 4.4 days
90th percentile: 9.2 days
Tracked items: 72

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 100% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. 70% of tracked open issues have had no activity in three months, so the open count overstates what is actively being worked. Only 5% of issues opened in the past year have been closed.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 47
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 65 days
Stale 30+ days: 42
Stale 90+ days: 33

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Detailed Description

AMD Strix Halo Llama.cpp Toolboxes is a container system that enables running large language models on AMD Ryzen AI Max "Strix Halo" integrated GPUs.

The project solves the problem of deploying LLMs efficiently on Strix Halo hardware by providing pre-built Toolbx containers with multiple backend options. Toolbx is Fedora's standard developer container system, now compatible with Ubuntu, openSUSE, and Arch. The toolboxes bundle llama.cpp with different GPU backends—Vulkan via Mesa RADV or AMDVLK, and ROCm—allowing users to choose based on stability and performance needs. The project includes host configuration guidance, performance benchmarks, memory planning tools, and distributed inference support.

Developers working with Strix Halo hardware who want to run LLMs should consider this tool if they are using compatible Linux distributions. The Vulkan RADV backend is recommended as the most stable and compatible option for most users and models. The AMDVLK backend offers the fastest performance but has a single buffer allocation limit that prevents loading some larger models. ROCm containers provide access to AMD's supported stack, with versions available for both the latest and previous stable releases. The project includes specific guidance on stable kernel versions and firmware to avoid, indicating that hardware compatibility requires careful configuration.

The project maintains a substantial base of real-world adopters, as evidenced by most open issues being raised by outside users rather than the core team. Maintainers respond to new issues and pull requests within hours, indicating active engagement with the community.