August 23, 2026 · By YasKad
p-e-w/heretic

Heretic: automated directional ablation for language models

p-e-w/heretic · 32,341★ · 3,634 forks

Everything worth knowing about p-e-w/heretic: a command-line application that modifies Transformer language models through directional ablation and automatic optimization.


What Heretic is

Heretic is a Python tool for modifying Transformer-based language models without a costly fine-tuning cycle afterward. Its stated primary use case is reducing refusals through a parameterized variant of directional ablation, which the project also calls abliteration.

The project doesn’t present that modification as a safety improvement or a guarantee of preserved capabilities. It seeks a measured trade-off between fewer refusals and lower KL divergence from the original model; the README itself warns that values depend on platform and hardware and that automated metrics don’t substitute for human evaluation.

Dramatic illustration representing the tension between refusal suppression and capability preservation: a glowing neural network split down the middle by a luminous fracture line. On one side, dark jagged crystalline structures representing refusals crumble and dissolve into particles; on the other, pristine data streams in cyan and gold remain intact, representing preserved model capabilities, with a perfectly balanced neon light scale between the two forces.

The distribution published on PyPI is called heretic-llm and exposes the heretic executable; PyPI showed version 1.4.0 at the time of this query.

The origin: from independent research to an automated tool

The author and maintainer identified by the repository is Philipp Emanuel Weidmann (p-e-w). Their profile describes them as a mathematician and software engineer with fifteen years of industry experience and an independent researcher in language model alignment and interpretability.

The citation included by the project itself attributes Heretic to Weidmann and dates it to 2025. Rather than asking the user to manually tune layers and weights, the design automates the parameter search with Optuna; that choice responds to the project’s central tension: intervening on refusal directions without unnecessarily degrading model behavior outside those tests.

The project explicitly builds on the 2024 work by Arditi and colleagues, on writing by Jim Lai, and on observations published by Maxime Labonne. The README emphasizes that Heretic was written from scratch and doesn’t reuse code from the earlier implementations it lists.

Philosophy and principles

  • Automation over manual tuning: on starting a run, the program determines batch size, computes directions, and explores parameters; it doesn’t require internal knowledge of the Transformer architecture.

Stylized cyberpunk illustration of the internal architecture of a Transformer model being disassembled and reassembled without manual tuning: floating holographic panels show attention projection matrices and MLP weight matrices as grids of glowing numbers. Autonomous robotic arms made of light — pure energy rather than metal — adjust and orthogonalize the matrices with surgical precision, with no human operator present.

  • Preserving a quality signal: trial selection combines refusal count with KL divergence from the starting model, rather than optimizing purely for refusal suppression.
  • Distribution traceability: the official site offers PyPI, GitHub, an official Codeberg mirror, release archives, the Internet Archive, and IPFS; it explains this redundancy as a resilience measure against outages.
  • Supply chain verification: the project declares locked versions via uv, a seven-day delay for dependency updates, Sigstore signatures on releases, and GPG-signed commits by the maintainer.

Conceptual digital illustration of supply chain verification in a software ecosystem: an infinite, dark corridor of mirrored servers reflecting one another, with glowing cryptographic seal icons — Sigstore and GPG signatures — floating like sentinel orbs along a luminous data channel. Redundant distribution nodes rendered as interconnected crystalline hubs (PyPI, GitHub, Codeberg, IPFS, Internet Archive) pulsing with synchronized neon-blue light.

How it works

Heretic computes, for each layer, residual vectors of the first output token from prompt sets classified as benign and harmful. It treats the difference of means as a refusal direction and orthogonalizes attention-projection and MLP-layer matrices relative to those directions.

Abstract visualization of directional ablation in a neural network: a three-dimensional residual stream rendered as a translucent glowing tube of data vectors flowing through deep space. An orthogonal geometric plane — a shimmering sheet of neon-magenta light resembling glass — intersects the stream at a perfect right angle, filtering out dark, jagged particles representing refusal directions while allowing smooth, luminous cyan data particles to pass untouched.

Optuna’s TPE optimizer searches combinations of direction_index and per-component ablation weights. According to the README, it allows linear interpolation of non-integer direction indices and applying different weights to attention and MLP; the stated technical reason is that MLP interventions tend to be more damaging than attention ones.

Futuristic visualization of an automated optimization process: a vast dark chamber filled with floating holographic trial nodes, each a small glowing sphere connected by thin neon threads forming a Pareto frontier curve arcing through the void. Some spheres glow cyan (low KL divergence), others pulse magenta (high refusal suppression), and the frontier represents the trade-off boundary between the two.

The documented flow is: GPU detection, loading and analyzing the model from Hugging Face, loading prompts, determining the maximum batch size, optional reasoning-prefix detection, initial evaluation, direction computation, optimization trials, and choosing a point on the Pareto frontier. At the end, the program offers to save, upload to Hugging Face, chat with the result, or run evaluations.

Futuristic command center visualizing the end-to-end Heretic workflow pipeline as a luminous journey through a dark digital landscape. The pipeline flows from left to right: a GPU detection node glowing green, a Hugging Face model loading portal swirling with blue light, prompt datasets streaming as ribbons of text, optimization trials exploding like supernovae of data points, and a final Pareto frontier selection gate shining in gold. At the end, branching paths lead to icons for saving, uploading, chatting, and evaluating. The entire scene is connected by flowing neon data conduits against a deep black background.

It also includes interpretability features: --plot-residuals creates PaCMAP projections in PNG and a GIF animation; --print-residual-geometry prints quantitative metrics of the residual geometry. PaCMAP runs on CPU and the documentation warns that on large models it can take an hour or more.

High-tech interpretability visualization: a PaCMAP dimensionality reduction plot projected as a three-dimensional holographic star map floating in a dark void. Clusters of residual stream embeddings appear as luminous nebulae — some glowing cyan (benign prompts), others burning magenta (harmful prompts) — separated by a visible geometric gap representing the refusal direction. A ghostly wireframe of a Transformer layer overlays the scene, with quantitative geometry metrics rendered as faint neon text annotations in the surrounding space.

Official and semi-official status

Heretic is an independent project: no evidence was recovered of acceptance into an official marketplace by a model provider, nor of formal endorsement from Hugging Face, PyTorch, Google, Qwen, or another manufacturer. It does have its own documentation site, a PyPI package, a public repository, an official Hugging Face link, Discord, Matrix, and an official Codeberg mirror.

In practice, its PyPI distribution and signed releases enable a verifiable install, but that isn’t equivalent to certifying results, guaranteeing the safety of modified models, or approval from model providers.

The ecosystem

Project repositories and channels

  • p-e-w/heretic is the main repository. The author’s public profile also highlights p-e-w/waidrin, p-e-w/sorcery, and p-e-w/arrows, but the pages retrieved don’t demonstrate they’re dependencies, extensions, or components of Heretic; they’re therefore recorded only as projects by the same author.
  • The official site links an official Codeberg mirror, the Hugging Face profile, Discord, and Matrix. Those channels are part of the publishing and community infrastructure, not derivative repositories.
  • The repository includes predefined configurations: config.default.toml for refusal suppression, config.noslop.toml for slop suppression, and config.nohumor.toml for humor suppression.

The README names as prior public implementations of ablation techniques AutoAbliteration, abliterator.py, wassname’s Abliterator, ErisForge, Removing refusals with HF Transformers, and deccp. This list demonstrates topical relation, not compatibility, active maintenance, functional equivalence, or an independent quality comparison.

The search for forks, same-named repositories, and localized ports couldn’t be completed via the GitHub API: it returned a rate limit during this run. The public page does show roughly 3,000 forks, but no specific ports, translations, or extensions are attributed without being able to verify their README or description.

Repo numbers

Measured: August 13, 2026, public GitHub pages; the REST API was rate-limited.

MetricVisible value
Stars27.4k
Forks3k
Commits192
Branches8
Tags5
Open issues42
Open pull requests31
Latest releasev1.4.0, published June 14, 2026
LicenseAGPL-3.0 or later
Language and declared environmentPython; console and GPU

Abbreviated figures are kept as shown by GitHub; they aren’t converted to exact integers. The releases page lists v1.4.0 as the latest and shows ricyoung, anrp, p-e-w, kabachuha, coder3101, zaakirio, MoonRide303, UnstableLlama, umran666, Vinay-Umrethe, rocker-zhang, and iuyua9 among those who contributed to that release, but a global contributor ranking couldn’t be retrieved without the API.

GitHub visually separates the 42 open issues from the 31 open pull requests. Since the API wasn’t available, subscribers_count isn’t reported here, which is distinct from the watchers_count field GitHub duplicates, nor is a total mixing issues with pull requests.

How to contribute

No CONTRIBUTING.md file or separate contribution policy was found in the repository’s visible root. The README does establish, however, that contributing implies agreeing to publish the contribution under the same AGPL-3.0 or later.

There’s evidence of an active change process: the repository contains tests/, GitHub Actions workflows, and a v1.4.0 release listing contributions from twelve participants; the code page also shows end-to-end tests and a recent dependency-change release. The discussions section is also enabled and shows categories for announcements, general, ideas, polls, Q&A, and showcase. This doesn’t allow inferring a pull request template, branch requirements, or undocumented test commands.

How the community received it

The recoverable community evidence is uneven and should be read with caution:

  • The README gathers three linked external testimonials about models produced with Heretic: one person who started skeptical values the quality of extended responses from a GPT-OSS 20B model; another considers it the best uncensored model they’ve tried; and another reports that a modified Qwen3-4B was the best they could run with 16 GB of VRAM. These are experiences selected by the project itself, not an independent review or a benchmark reproduced here.
  • Issue 401, opened by Weidmann on July 5, 2026, communicates that they anticipate being able to stop developing software at some undetermined point in the future. Contributor rocker-zhang responded that working on Heretic had been a pleasure and thanked them for their work; accemlcc, identified by GitHub as a collaborator, wrote that the project had meant a lot to them and that they had learned from it. These are personal reactions to the announcement, not a technical evaluation.
  • In the public issues list, concrete problems appear: luyangliu616 opened issue 345 saying they couldn’t get it working, and sefgtrdh opened issue 318 about the missing libcaffe2_nvrtc.so. These issues show real installation and compatibility friction, not proof the problems affect every environment.
  • No verifiable direct Hacker News thread was recovered: the submissions-by-domain page came back empty and the Algolia query returned no usable records. Reddit returned an access challenge. As a result, no identifiers, points, comments, X opinions, videos, Product Hunt launch, or podcast mentions are invented. Nor were PyPI download figures obtained, since the stats service responded with a rate limit.

Heretic versus other approaches

ApproachVerifiable overlapLimit of the comparison
AutoAbliterationThe README lists it as a public ablation implementation.Its documentation wasn’t recovered during this run; no algorithm or performance differences are claimed.
abliterator.pyIt appears in the public prior-work list cited by Heretic.The source only demonstrates a topical relationship.
wassname’s AbliteratorIt appears in the list of prior implementations.Its README and metrics weren’t recovered; compatibility can’t be claimed.
ErisForgeIt appears as a prior public implementation.There’s no recovered basis to compare workflow, license, or results.

Heretic’s verifiable distinction lies in combining parameterized directional ablation with a TPE search that co-optimizes refusals and KL divergence, plus an automated console workflow. Any claim that it outperforms the listed approaches would require retrieving and reproducing their methods and evaluations, which wasn’t done in this investigation.

Quick-start guide

Installation and first run

  1. Prepare a Python 3.10 or later environment and have PyTorch 2.2 or later installed appropriate for your GPU; PyTorch must be installed manually because the command depends on the accelerator. Some models, such as MXFP4-quantized ones, require PyTorch 2.6 due to torch.accelerator.
  2. Install the stable release:
pip install -U heretic-llm

Dark-mode terminal interface glowing on a black screen, showing the word heretic in bright neon green at the command prompt, with cascading lines of Python execution logs scrolling below in cyan monospace text. The terminal window floats in a dark cyberpunk environment with subtle holographic data particles drifting in the background.

  1. Run the program with a Hugging Face model identifier:
heretic Qwen/Qwen3-4B-Instruct-2507

The current documentation also shows Qwen/Qwen3.5-4B as an example. On first run, the GPU is detected, the model is downloaded if not cached, a batch size is computed, and decisions are requested once optimization finishes.

For dependency reproducibility, anyone cloning the repository can use:

uv run heretic

The project includes uv.lock to pin the versions used by the developers.

Common workflows

  • Modify a model and save or publish the result: run heretic <model-id>. After optimization, the interface offers to save the model, upload it to Hugging Face, open a test chat, or run evaluations.
  • Evaluate a generated model: use the documented pattern heretic --model google/gemma-3-12b-it --evaluate-model p-e-w/gemma-3-12b-it-heretic. Values can change with hardware and environment.
  • Reduce VRAM usage: set quantization to bnb_4bit. The documentation notes that four-bit loading via bitsandbytes can reduce VRAM by roughly 70%.
  • Explore internal representations: install the extra and run the corresponding flags:
pip install -U 'heretic-llm[research]'
heretic --plot-residuals
heretic --print-residual-geometry

The first generates images and an animation; the second prints residual-geometry metrics.

Essential configuration

  • config.toml: the usual configuration file; it’s placed in the working directory from which Heretic is run.
  • config.default.toml: the template for refusal suppression; copy or rename it as config.toml before adjusting it.
  • quantization: select bnb_4bit to load with four-bit quantization and reduce required VRAM.
  • n_trials: controls the number of optimization trials; the tutorial shows 200 trials as an example, not as a universal requirement.
  • good_prompts, bad_prompts, good_evaluation_prompts, and bad_evaluation_prompts: determine the prompt sets used to compute directions and evaluate the trade-off between refusals and divergence.

Parameters can also be passed via command line, queryable with heretic --help, or via environment variables following the pattern HERETIC_<UPPERCASE_PARAMETER_NAME>.

Common pitfalls and fixes

  • Insufficient PyTorch version: PyTorch 2.2 is the minimum, but MXFP4 and gpt-oss need PyTorch 2.6 features. Install the appropriate PyTorch variant for your accelerator before installing or running Heretic.
  • Insufficient memory: enable quantization = "bnb_4bit"; the tutorial proposes it to reduce VRAM. Keep automatic batch determination unless you need to impose a limit with max_batch_size.
  • Expecting a short run: the tutorial illustrates a 200-trial optimization taking a bit under three hours in its example; plan your time and don’t treat that value as a guarantee for another model or GPU.
  • CUDA library failure observed by a user: issue 318 mentions a missing libcaffe2_nvrtc.so. The documentation doesn’t publish a specific fix; the well-founded action is to check that PyTorch, the driver, and the acceleration library match your hardware before opening an issue with environment details.
  • Installing from unverified sources: official releases include Sigstore signatures; use the release’s signature files and follow the site’s verification procedure if installing a compressed archive.

Integrations and migration

Heretic consumes identifiers and models from Hugging Face and, on finishing, offers to upload the result back to the same ecosystem. Its main practical integration is that model → modification → save, chat, evaluate, or publish flow; no documentation was found for integration with MCP, editors, CI, or messaging platforms.

For resilient installation, the project documents PyPI, cloning via GitHub or Codeberg, and release archives also available via the Internet Archive and IPFS. No migration guide was found from AutoAbliteration, Abliterator, or another tool; configuration or weight compatibility between them shouldn’t be assumed.

Use cases

  • Interpretability researchers needing to inspect residual directions can use --plot-residuals and --print-residual-geometry to generate projections, animations, and metrics, with the caveat that PaCMAP can be costly on CPU.
  • Anyone with a compatible Transformer model and a GPU wanting to apply an ablation procedure without manually choosing a layer or weights can use the built-in batch detection, direction computation, and TPE optimization.
  • Teams distributing derivative models on Hugging Face can take advantage of the save, chat, evaluate, and publish output the final flow offers. They’ll need to validate the result in their own domain: the project states that automated benchmarks don’t replace human evaluation.
  • Anyone operating in connectivity-constrained environments or concerned about provenance can choose between PyPI, GitHub, Codeberg, signed archives, the Internet Archive, and IPFS, and verify the Sigstore signatures on releases.

Resources


Note: this article combines official documentation, public GitHub pages, and PyPI consulted on August 13, 2026. Figures change over time; the GitHub API and PyPI statistics were rate-limited during this investigation.

Comments