containers/ramalama

RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in...

View on GitHub ↗Jump to charts ↓Open shareable report →

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 2 hours ago
Added to GitGenius on January 27th, 2025
Created on July 24th, 2024
Open Issues & Pull Requests: 123 (+0)
GitHub issues: Enabled
Number of forks: 381
Total Stargazers: 3,076 (+1)
Total Subscribers: 33 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 3.1 hours
Mean response time: 4.1 days
90th percentile: 7.1 days
Tracked items: 600

Maintainer activity

23 people did triage or write work on this repository in the last 12 months.

Counts unlabeled, assigned, unassigned, milestoned, demilestoned, locked, unlocked over the last 12 months. These are issue and pull request events that require triage or write permission. Commits and code review are not counted. labeled and renamed are excluded because GitHub issue forms record the issue author as the actor. Figures from October 7, 2026. This count is not comparable across projects: each project's automation decides which of these events a person emits.

How this project is maintained

About 6% of issues opened in the past year have never received a reply. 60% of open issues come from outside the core team, a mix of external reports and the maintainers' own roadmap. 58% of tracked open issues have had no activity in three months. 67% of issues opened in the past year have been closed, leaving a working backlog. Three people close 79% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 78
New in 7 days: 1
Closed in 7 days: 0
Avg open age: 166 days
Stale 30+ days: 60
Stale 90+ days: 50

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 2
Events in 7 days: 3

Top labels

  • bug (185)
  • enhancement (89)
  • stale-issue (85)
  • good first issue (58)
  • hacktoberfest (3)
  • help wanted (2)
  • question (2)
  • rag (1)

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

RamaLama is an open-source developer tool written in Python that simplifies the local serving and inference of AI models through containerized workflows. The project brings familiar container-centric development patterns to AI use cases by treating models similarly to how Podman and Docker treat container images, allowing engineers to work with AI models using common container commands and OCI container registries.

The core functionality of RamaLama eliminates the complexity of configuring host systems for AI workloads. Rather than requiring users to manually set up dependencies and hardware optimizations, RamaLama automatically detects GPUs present on the host system and pulls an appropriate accelerated container image tailored to that specific hardware. The tool supports multiple accelerators including NVIDIA CUDA, AMD ROCm, Intel Arc GPUs, Apple Silicon, Ascend NPUs, and Moore Threads GPUs, with fallback to CPU-based inference when no accelerators are detected. For macOS users with Apple Silicon, RamaLama also supports the MLX runtime for optimized inference without containerization.

RamaLama provides multiple installation paths across different operating systems. On macOS, users can download a self-contained installer package that includes Python and all dependencies. Fedora users can install via DNF, while Windows users can run RamaLama through Docker Desktop or Podman Desktop with WSL2. The tool is also available via PyPI for Python-based installation and supports installation on immutable systems like Fedora Silverblue through Toolbox containers.

Security is a primary design consideration in RamaLama. Models run in rootless containers by default, isolating them from the underlying host system. Models are mounted as read-only volumes within containers, and the tool defaults to no network access with automatic cleanup of temporary data on application exit. Users can interact with models through either a REST API or a chatbot interface.

RamaLama is classified across multiple domains including OCI standards, container runtimes, Kubernetes integration, CI/CD frameworks, and DevOps tooling. The project supports Hacktoberfest contributions and maintains community channels on Discord and Matrix. The tool's approach of using containers as the abstraction layer for AI model serving represents a significant shift in how developers can approach local AI inference, combining the reproducibility and portability benefits of containerization with the practical needs of AI model deployment and testing.