tracel-ai/burn

Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability.

View on GitHub ↗Jump to charts ↓Open shareable report →

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 1 hour ago
Added to GitGenius on July 23rd, 2025
Created on July 18th, 2022
Open Issues & Pull Requests: 184 (+2)
GitHub issues: Enabled
Number of forks: 1,081
Total Stargazers: 16,066 (+1)
Total Subscribers: 101 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 25.0 hours
Mean response time: 68.1 days
90th percentile: 243.8 days
Tracked items: 1,140

Maintainer activity

11 people did triage or write work on this repository in the last 12 months.

Counts unlabeled, assigned, unassigned, milestoned, demilestoned, locked, unlocked over the last 12 months. These are issue and pull request events that require triage or write permission. Commits and code review are not counted. labeled and renamed are excluded because GitHub issue forms record the issue author as the actor. Figures from October 7, 2026. This count is not comparable across projects: each project's automation decides which of these events a person emits.

How this project is maintained

About 6% of issues opened in the past year have never received a reply. 65% of open issues come from outside the core team, a mix of external reports and the maintainers' own roadmap. Work labelled "store" is answered fastest, typically in about 20 hours, while "feature" waits about 9 days. Almost all tracked open issues have seen activity in the last three months. 80% of issues opened in the past year have been closed, leaving a working backlog.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 153
New in 7 days: 16
Closed in 7 days: 37
Avg open age: 373 days
Stale 30+ days: 2
Stale 90+ days: 2

Recent activity

Opened in 7 days: 12
Closed in 7 days: 33
Comments in 7 days: 2
Events in 7 days: 7

Top labels

  • bug (375)
  • enhancement (180)
  • feature (138)
  • onnx (79)
  • performance (62)
  • wgpu (61)
  • documentation (47)
  • question (42)

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

Burn is a modular and flexible PyTorch library focused on accelerating and simplifying the deployment of machine learning models, particularly large language models (LLMs), across a diverse range of hardware backends. It aims to bridge the gap between research and production by providing a unified interface for model compilation, optimization, and execution, while supporting various precision levels (FP32, FP16, BF16, INT8, INT4) and quantization techniques. The core philosophy revolves around composability – allowing users to mix and match different components to tailor the deployment pipeline to their specific needs and hardware constraints.

At its heart, Burn utilizes a graph-intermediate representation (IR) called `GraphTensor`. This IR decouples the model definition from the underlying hardware, enabling optimizations to be applied independently of the original framework. Models are first loaded from popular frameworks like PyTorch and then converted into this `GraphTensor` representation. This conversion process is handled by "backends" which are responsible for translating the PyTorch operations into the Burn IR. Currently, Burn supports a growing list of backends including CUDA, CPU, Metal (Apple Silicon), and WebGPU, with ongoing work to expand this list. This backend-centric design allows for easy addition of new hardware support without modifying the core Burn library.

A key feature of Burn is its support for various quantization techniques. It offers both post-training quantization (PTQ) and quantization-aware training (QAT). PTQ allows for quick and easy model compression with minimal retraining, while QAT provides higher accuracy at lower precision by incorporating quantization into the training loop. Burn’s quantization capabilities are designed to be highly configurable, allowing users to fine-tune the quantization parameters to achieve the best trade-off between accuracy and performance. Specifically, it supports techniques like GPTQ, AWQ, and GGML/GGUF formats, making it compatible with a wide range of quantized models.

Burn’s modularity extends to its compilation and optimization stages. Users can choose from different compilation strategies, including just-in-time (JIT) compilation and ahead-of-time (AOT) compilation. AOT compilation can significantly improve performance by pre-compiling the model for a specific hardware target. Furthermore, Burn incorporates various optimization passes, such as operator fusion, memory layout optimization, and kernel selection, to maximize performance on the target hardware. These optimizations are applied to the `GraphTensor` representation before execution.

The repository includes examples demonstrating how to load, quantize, and deploy models on different backends. It also provides tools for benchmarking and profiling, allowing users to evaluate the performance of their models and identify bottlenecks. Burn is actively developed and maintained, with a strong focus on community contributions and expanding its capabilities. Its goal is to become a leading solution for deploying LLMs and other machine learning models efficiently and effectively across a wide range of hardware platforms, making powerful AI accessible to more users and applications.