p-e-w/heretic

Fully automatic censorship removal for language models

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 49 minutes ago
Added to GitGenius on February 18th, 2026
Created on September 21st, 2025
Open Issues & Pull Requests: 76 (+0)
Number of forks: 3,039
Total Stargazers: 28,064 (+0)
Total Subscribers: 128 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 7.8 hours
Mean response time: 5.2 days
90th percentile: 6.9 days
Tracked items: 195

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 88% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Work labelled "enhancement" is answered fastest, typically in about 3 hours, while "bug" waits about 3 days. 65% of tracked open issues have had no activity in three months, so the open count overstates what is actively being worked. Only 10% of issues opened in the past year have been closed.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 48
New in 7 days: 2
Closed in 7 days: 0
Avg open age: 62 days
Stale 30+ days: 44
Stale 90+ days: 31

Recent activity

Opened in 7 days: 2
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 1

Top labels

  • bug (19)
  • enhancement (13)
  • question (2)
  • wontfix (2)

Detailed Description

Heretic is a tool that removes safety alignment restrictions from transformer-based language models without requiring expensive retraining.

The tool addresses the problem of censorship in language models by applying directional ablation, an automated technique that identifies and modifies specific model parameters responsible for refusal behavior. Rather than requiring manual expertise in transformer internals, Heretic uses a parameter optimizer based on tree-structured Parzen estimation to automatically find abliteration settings that minimize refusals while preserving the model's original capabilities, measured by KL divergence on harmless prompts. The process runs completely unsupervised with default configuration, making it accessible to users without deep technical knowledge of how language models work internally.

Heretic suits developers and researchers who want to remove safety restrictions from language models and have access to sufficient computational resources. The tool supports most dense transformer architectures including multimodal models, multiple mixture-of-experts designs, and hybrid architectures, though pure state-space models and certain research architectures are not yet supported. According to the README, Heretic's automatically generated decensored models achieve comparable refusal suppression to manually created abliterations while introducing less degradation to the original model's performance.

The project maintains a substantial user base, with nearly all open issues coming from external adopters rather than the core team. Maintainers typically respond to new issues and pull requests within a day. Work in the issue tracker centers on bug reports, enhancement requests, and user questions.