Also, as I said in a top level comment, what this project wants to achieve has been done for a while and it's called Heretic: https://github.com/p-e-w/heretic
(Not vibecode by a twitter influgrifter)
And yeah, doing stuff like deleting layers or nulling out whole expert heads has a certain ice pick through the eye socket quality.
That said, some kind of automated model brain surgery will likely be viable one day.
It also seems the influgrifter has a lot of bots (or perhaps cultists) working this thread...
I use Berkley Sterling from 2024 because I can trick it. No abliteration needed.
Strategy What it does Use case
.......................................................
layer_removal Zero out entire transformer layers
head_pruning Zero out individual attention heads
ffn_ablation Zero out feed-forward blocks
embedding_ablation Zero out embedding dimension ranges
https://github.com/elder-plinius/OBLITERATUS?tab=readme-ov-f...It's interesting that people are writing tools that go inside the weights and do things. We're getting past the black box era of LLMs.
That may or may not be a good thing.
However, after a few rounds of conversation, it gets into loops and just repeats things over and over again. The main JOSIE models worked the best of all and was still useful even after abliteration.