Every command+data channel in existence has been and will continue to be exploited one way or another, because the solution space is for all intents and purposes unbounded. Sure, highly defensive escaping reduces attack surface dramatically, but e.g. prepared statements eliminate the whole class of bugs.
As far as I understand, current LLMs are architecturally incapable of this separation. Given the inherently recursive nature of GenAI, the model itself is part of the input space, making validation essentially impossible.
You see, you can (usually) easily tell what a particular escaping transformation does. That does not tell you neither how it will be interpreted down the line, nor what should be done.
Arguably the most common problem is double escaping. This typically manifests as various double escaping bugs.
If your hand rolled implementation just chains `.replaceAll("<sepcial>", input)` and `.replaceAll("<escape>", input)` the escaping result depends on evaluation order, at least on already pre-escaped inputs if your instruction sequences are single-element. Even if you get it right and don't reduce escaped sequences to double escape + unescaped, you are still dealing with stray escape sequences: "John o\'Doe".
You have to meticulously (all the way through your call tree and even through persistent storage cycle!) track if a value has already been escaped and whether it needs escaping. As paradoxical as it may sound, meticulous and defensive escaping produces tons of bugs. If the project decides to escape raw inputs right when passed in, escaping them once more before passing out (to storage, another process) is a bug that you cannot easily statically test against.
On top of that, various modules that you interact with (storage, libraries, modules pulled from another team) will have different escaping semantics: some will apply escaping on their own, some will expect input to be "sanitized" and treat it raw. The semantics can even be different on different paths: write to storage module accepts input as is, but retrieval method "helpfully" runs escaping.
Furthermore, in different contexts the escaping rules are going to be different. In a web world, what's safe to write directly to html, pass to js `alert`, and pass to sql query are entirely different things. Input safe to dump into html is not necessarily safe to pass to string-interpolating SQL DTO layer, and vice versa. Then the DBA changes config to allow variables in queries and your escaping _semantics_ are now entirely different.
It's a minefield with essentially unbounded surface. You will trip up. People have tried to solve the problem for decades. Very smart people have tried. They all have failed.
In the case of writing to the DOM, that means you use .innerText rather than .innerHTML. In the case of writing to a database, that means you use the driver and insert variables rather than directly into the string. Both of these are technology-specific escaping APIs. It's just that the browser is much better at making sure HTML is passed safely than you are.
This is even enforceable with Trusted Types.
That's not true. Most have failed, but those who used the right tool for the job - a rich, static type system in a functional language - did succeed. It's just that such type systems are rare, and even if nominally a type system is good enough, the required boilerplate might be uneconomical to maintain. Scala and F# are probably the only two languages that are mainstream-adjacent, at least, and have type systems expressive enough that using them to track escaping is not an absolute hassle. And they're both tiny in terms of the number of users.
In the end, we did settle on APIs that hide the escaping process (prepared statements, innerText vs. innerHTML, etc.) just because it's a) good enough; and b) possible to implement more or less uniformly across the TIOBE Top 20. That's good, but if you happen to use a language that's powerful enough, you probably also should track the escaped/unescaped status in the type system - it can be a cheap, additional safety net.
Those are interfaces where escaping is no longer needed precisely because they deinterlace instructions from data. Escaping problem inherent to instruction+data channels. Yes, deinterlacing is not solution, not futile attempts to escape.
Specifically, I serve the model with this shell script on my M2 Max: https://github.com/shawwn/scrap/blob/master/llama-serve
It's pretty good. I used it to do some pesticide research. (Normal models all refuse due to guardrails about bioweapons.)
As someone who has never once had any need whatsoever to research pesticides I… don’t think it’s bad at all? I don’t want anyone to have the capability to invent a human-targeted pesticide who isn’t verified not crazy?
I don't think you can 100% ensure your queries have absolutely no biology and cyber keywords inside. No matter how harmless, they always trigger. People complain Fable aborts even when they try to make a login page for showing "username" and "password".
This makes models like Fable 5 impossible to use in any serious agentic task, because you can't even guarantee the model, which is a basic thing you need to build on.
Obviously. Almost everything is a precursor to something dangerous, to the extent that if some model isn't aware of the risk it will wander into it blindly, e.g. suggesting leaving raw garlic and olive oil alone for a week without awareness this will likely breed botulism bacteria.
> This makes models like Fable 5 impossible to use in any serious agentic task, because you can't even guarantee the model, which is a basic thing you need to build on.
This is binary thinking: "100% ensure", "impossible to use", "can't even guarantee the model".
Outside computers, most work is not binary, it's probability, e.g. "this skyscraper will probably survive being hit by an aircraft; oh we didn't mean a 747 we meant a small Cessna, but what's the chances of a 747 crashing into it soon after takeoff?".
Fable being too cautious for its own good (especially since the other models were not) is a fair criticism, but this isn't a binary question.
I’m pretty sure that I encountered this the other day. I gave it a copy of a paper by biologist Michael Levin and mentioned off hand that it should be much easier to replicate that his other work (because most of his work is biological lab work and this paper was about sorting algorithms) and it immediately told me that I couldn’t use Fable for this.
This just isn’t feasible. These jackasses spent the last few years telling the world that their products are going to destroy the world to make them seem edgy and to justify regulations that benefit the entrenched players and now they’re going to be the ones to decide what we do with this technology?
History is going to look back at this time and how we let such foolishly inconsistent people make such grand choices for everyone poorly.
Assuming we have a future history. We've already got "history slop" with AI rewriting the past by their incompetence.
Given they're "such foolishly inconsistent people", would you rather they err on the side of caution like this? Or the side of boldness, like Musk has been doing with FSD/Autopilot or Grok porn, all of which he's getting in legal trouble over?
I distrust Musk and Zuckerberg (to put it mildly), so it's fair if you say you don't believe anyone's public statements; but I also hang out with some of the researchers on this, and a fear of e.g. ending up with something as criminally unhinged in cyber-work as Grok was with porn is the least of their worries. Plenty of them also fear a corporation centralising power with such tools (such power is Musk's entire sales pitch for why line go up in future).
Who could you get to work on this inherently bullsh*t tech, but inherent bullsh*tters?
So pretty much like all of history before this point.
Oh? How, exactly?
Has signed presidential orders against them preventing them from doing business as usual.
They probably should change that, and it is corproate speak, but I read it as:
"Our (forced by presidential order or otherwise we couldn't offer you this model at all) broad safeguards (now allow us) to deliver more capabilities.
There's lots to complain about with some of these companies. But let's pile on where it's deserved.
Though at this point the safeguards are making the product useless.
In that context, I'd say probably the current admin is indeed the cause of that. They certainly claim to be. And they used the full power of the federal government along with interventionist policies and, it would seem, a gangster like mentality to punish those they disagree with.
So would it be as obnoxious otherwise? I don't think so. And there would have to be guardrails regardless, or apparently anything the tool is used for is considered what you support and want. So can you imagine a platform with no guardrails?
But doesn't the timeline of events indicate that these guardrails are internally-driven, not externally-driven? I'm trying to attribute the guardrails and their severity appropriately.
The veracity of that? I have no idea, really. I simply know that the current admin is staggering around like an abusive, angry, drunk of a parent, lashing out at everything and everyone, making the world we all inhabit a crappier place.
So there was an EO, then "things were fixed" to abusive man's standards, so I presume that means crappier in some way.
It's the best I've got.
I don't want anyone to have the capability to rape women.
Any physics teacher should know the theory for constructing a nuclear bomb. Should we be controlling that knowledge too?
What about flight simulators? Don't want a load of people knowing how to fly.
This isn't computer science, the hard bit is getting the materials and equipment, not the knowledge.
Maybe not the best example, since that knowledge is some of the most highly controlled in the world.
But to mirror the point I made in a different post, the difficult part of making a nuclear bomb is not finding the theory behind like Little Boy. It's making an entire industry to generate HEU, etc.
That bozo baq should actually go ahead and write out the costs and then he will quickly realise he has no bloody idea what hes talking about.
Another deluded bozo.
Uhh.. we ( for a value of we ) are. Sure, it is not overt, but if you have not seen funnels, social stigma associated with some otherwise benign activities, you are not paying attention.
Also, if you live in America, it is much easier and more effective to create a mass casualty event with, say, a few cases of fireworks and a pressure cooker or an AR-15.
He is either pushing AI for whatever reason or he is in psychosis. Completely disconnected from reality.
These models are, ultimately, tools. I would never trust some random corporation (particularly one with a profit motive and hypocritical stance, which includes both OpenAI and Anthropic, to be clear) to decide what isn't and is considered "crazy" and who and who isn't "verified" not to be "crazy". Especially when these companies have time and time again demonstrated (1) that they cry wolf way too much which leads to nobody taking their claims about how "dangerous" their models are seriously and (2) incidents like this where OpenAI makes a claim ("Look at how dangerous our models are!") and then doesn't be smart and just... Slow the fuck down (and when testing these things, actually sandbox them properly, which obviously wasn't done here or this attack wouldn't have been even possible).
Correction: There's tens of thousands of them. They're easy to create, which is why everyone publishes their own.
Just put "uncensored", "abliterated", or "heretic" into search on huggingface/ollama/etc and pick any them. Fair warning: most aren't very good, essentially lobotomized, and totally broken if you enable thinking.
That is surely the point, most of the "uncensored" weights released for free on HuggingFace aren't being very successful at this. There is a stark difference in output quality between the official weights and all these "uncensored" variants that appears days afterwards.
Even as I write this the ‘abliterated’ word is denoted a typo. Does it not at your end?
https://en.wikipedia.org/wiki/Ablation
"Ablation (Latin: ablatio – removal) is the removal or destruction of something from an object by vaporization, chipping, erosive processes, or by other means."
which is what one does to the model during... well, ablation.
https://en.wikipedia.org/wiki/Ablation_(artificial_intellige...
Have you read some (any) actual papers from the word this all and LLMs are?
Well here some:
https://arxiv.org/abs/1901.08644
Tis ablation ablation and ablation everywhere save for some kid’s fav. model names and their X pals.
Note that they're all talking about the same technique, namely applying ablation specifically to the refusal vector as described in https://arxiv.org/abs/2406.11717.
I really recommend letting this one go.
> The term abliteration has been coined for the process of using ablation to uncensor large language models by modifying internal functions to completely eliminate refusal behaviors while preserving the remaining functions of the model. The word is a portmanteau that combines the words ablation and obliteration.
So are you dropping this now?
curprev 22:29, 29 May 2025 ZeevoX talk contribs 2,320 bytes +2,320 Create page for "abliterate" Tag: added link
not very credible.
Except for the initial comment, they were aggressive and sometimes disrespectful claims that "the word is spelled ablation, here are some papers that use it" in response to people saying "yes, we know what ablation means, but GP is intentionally using a separate word 'abliteration' that is the accepted word for the kind of ablation he's talking about, see links to respectable sources using or defining it". I downvoted them because I feel the discussion would be more valuable and feel nicer to read without them, and you could just read any of the offered links before responding and save the trouble.
The wiki "ablation" article you yourself linked dedicates a section to explain "abliteration".
Googling abliteration arxiv yields at least one page of papers that mention it in the abstract (all but one in the title too). I counted 9 unique papers.
If it was shown that all of these were posted after this discussion started, or by people related to the person you tried to correct, I would be convinced that abliteration is not a real word. But evidence keeps pointing otherwise, nnd you keep arguing with evidence that proves "ablate" is a word (to which we all agree), not evedence that proves "abliterate" isn't.
You did show evidence that it's a pretty new word (of course it is! it's a pretty new technique in a field that didn't exist before the first open source RLHF'd models were released in '23!), and indeed, this (different) wiktionary page contains its origin story from '24: https://en.wiktionary.org/wiki/abliterate#English
To your question, I am not a bot, my LinkedIn is in my profile if you want to know who I am.
I'm persisting because, I guess, I'm really confused by your own insistence, and feel curious to get to the bottom of the weird misunderstanding we must be having (maybe you're trying to argue something different than "user chmod775 intended to refer to 'ablation' and was mistaken to write 'Just put "uncensored", "abliterated", or "heretic" into search on huggingface/ollama/etc and pick any them' instead of 'Just put "uncensored", "ablated", or "heretic" into search on huggingface/ollama/etc and pick any them'?). And I have a lot of free time on a climbing vacation where my brain is too mushy to do anything more productive than talk to people on the interblags.
Are you for real? I mean - how is "legit" THE word for legitimate? Tis not, you see.
Abliteration is at best a marketing/community term coined in some forum. And all you claim it is THE terminology, well sorry, this sounds ridiculous.
So, we can keep having this argument, though the more I dig, the more this goes in the direction that all you folks are defending a made-up word for something, coined by someone who is not in a position to do so. Unlike mr. Karpahty for example, who by accident coined 'vibe coding' and is still struggling to replace it with "agentic dev" and fails doing so.
And yes, dig papers - it says "ablation studies" in most, if not all, that I have read. This abliteration thing is some abomination thing, and gets underscored as typo even as I write this comment.
p.s. see your linkedin for incoming contact request by s.o. starting with Gh in one of the names.
I am unsure if this is terminology I am unfamiliar with, a typo of Devs, or a 2001 reference.
Devs tell computers what to do. Computers tell Daves "I can't do that."
what a time to be alive.