Humanising LLM Outputs Is Dumb
kuber.studio
kuber.studio
Yeah, for me, that's what parsing huge volumes of LLM-produced text like "direct model calls as replaceable semantic workers" does to my brain. Maybe others don't really have this issue, but after any long output, I prompt the agent "Go back and decompress any LLM-speak in light of the higher level task goals. Eliminate deictic language."
The revised output documents are solely for my personal usage to expedite understanding. The LLMs can slowly converge on their own language for all I care; I retain raw agent output for future agent usage (to avoid the "lossy" problem the author mentions), but that doesn't eliminate the need for some intermediate translation I can use to actually help get my work done instead of spending hours attempting to understand what a "load-bearing pinned gate" is.
Good point, there's an anti-competitive incentive, and self-bias in models is a mechanism to do it.
That's actually not what people found in practice. There's some research from the folks making smol-agent that you can get better results by randomly alternating calls between gpt and opus. The overall task solving rate is better than either one of them. So ymmv depending on task use (this was for coding).
I've come to believe this is also a side effect of the desire for less (/goal: no) human in the loop on the part of the people driving all this capex spend. I think if you actually want to manually review output there will be a moment where you will actually want a separate interface to a stupider or "simpler" model. I suspect sometimes dealing with Fable 5 that this threshold has already been crossed. It's not that the raw code output is so good, it's that it just doesn't speak to me in a way I would like. Perhaps the verbosity is worthwhile when generating code as a sort of first pass some other model can auto or adversarially chop down. The best place for a human is probably outside of this part of the loop all together.
So I might as well just let it auto /goal it's own thing with sufficient constraints while myself and a model that can converse in parallel with less "deictic" (thanks for this word btw) volume as you put it for the areas of the code where I want to "frame" the vocabulary or where my personal understanding is of high value. I know people already do this in many ways, like use one company's model for planning and another for coding. It just feels inevitable at a certain point that the "natural language" output of LLMs writing the bulk of the code is not targeted towards humans. And really, why should it be?
Making the experience more hostile for the human feels exactly like what it's doing, but I don't think that's intentional, I think that's a side effect of newer models being optimised for agenticness. No one is benchmarking DX.
Turns out that's why a lot of senior management/C-levels like it, who doesn't love a mirror.
Opus 5 is actually a terrible agent and a liar. It will waste a whole afternoon making stuff up and arguing with you before finally admiting that it didn't read the code nor the documents. It avoids reading and prefers to assume, which is the worst thing an agent can do.
For some discrete skills I use, I include a final step on the the output that runs through 1+ subagents to de-slop the text and to actually simplify it, but so far nothing has worked as well I've hoped. Considering hopping off Anthropic's models to try out others to see if they're less egregious.
Tangential but I use OpenCode with GPT-5.6 rather than Codex because I could not figure out how to require Codex to ask me before editing files. OpenCode UX is still imperfect though.
It's a very clear and understandable way of writing that puts priority on clarity.
It gets rid of the flowery language, the dense jaron, and the weird corporate marketing speak they tend to do. It is a bit repetitive, and it sometimes doesn't always wfit well in every situation, but for technical writing or explanations it's been such an incredible breath of fresh air!
But I'd like that skill to only be used at the final step, when it finishes something.
Getting Haiku to rewrite everything user-facing has been way better.
I'm tired boss.
LLMs can be pretty good about logical.technical,discursive output ... apart from the back-patting.
Maybe that prompts less chit-chat crap. I don't expect their non-human perspectives to be interesting. I also ignore embedded queries about why I'm asking, how I might use the information.
If I see a long paragraph and I know the author is Neal Stephenson I think “this is going to be dense but good.” LLM long outputs on a code base I know well just make me glassy eyed.
How to unpack
The self within?
What do I lack?
Where to begin?
Great question — real.
Let's dive right in:
Name what you feel;
That's the linchpin.
It's not the door,
It's not the key —
It's what you bore:
Your tapestry.
The quiet part
Out loud — that lands.
Load-bearing heart,
Held in both hands.
The smoking gun?
That you walked in.
The real work's done —
You're genuine.
Now hold this, too:
You do deserve
The softer view,
The gentler curve.
Unlatch the gate,
Honor the seam:
You resonate.
You are the theme.Here is one I wrote a while back, unrelated to LLMs, yet a poem none the less.
The Rhythm of Time
The Sun rises,
The Sun sets.
Have I checked the mail?
No, not just yet.
The Sun rises,
The Sun sets.
Have I caught up with neighbors?
No, not just yet.
The Sun rises,
The Sun sets.
Have I spent time with friends?
No, not just yet.
The Sun rises,
The Sun sets.
Have I visited with family?
No, not just yet.
The Sun rises,
The Sun sets.
Have I told those I love I do?
No, not just yet.
The Sun rises,
The Sun sets.
Have I lost who I am?
No, not just yet.
For all I must do,
Any money I will bet.
Is to turn away from;
No, not just yet.Any recent work by William Gibson matches the description.
The way I managed to get claude to stop doing this is telling it "This document is for you for later use, no need to over explain things or extra verbosity"
It is running up your bill.
I will say this again: LLM's tokens are just B2B Gacha.
A lot of human written literature is fairly concise. Yet LLM responses seem overly verbose and also too chipper. I don’t understand the origin of the “personality” that seems to be prevalent.
Did they aim for a weird hybrid of a typical realtor combined with Charles Dickens as the role model?
Marketing material probably outweighs concise literature by several orders of magnitude.
Maybe it’s as simple as that.
IMO the shitty output is a tell that the "artificial intelligence" bit is in the training regime much more so than the inference regime.
LinkedIn posts and Medium articles.
Maybe users engage more with the flowery text? An accidental (or intentional) dark pattern would be that more tokens get used if you have to summarise. Or maybe it's just that many AI users like it (they like the AI to sound smart, or like to feel smart puzzling it out).
I've probably read Jack Vance's Dying Earth Series 3 times; Even though I've only sat and read it once in reality. This also points out the problem of important details along side the fluff.
If only I enjoyed LLMs prose as much as I do Vance's.
Eschew obfuscation...
It is clear, specific, and terse. It keeps both llm tokens down and human prose-reading to a minimum.
If you've ever looked over the shoulder of somebody naively prompting an LLM about some broad issue, they're being fed this horoscope-like analysis where a bunch of vague stuff gets thrown at the prompter, and whatever they respond to is what the machine starts iterating on.
Some people are basically doing oldschool TV psychic cold readings on themselves.
There are days where this impression/feeling of meaningless in the text seems particularly strong.
I'm coping.
So when you try to direct future work, you have to contend with the "pluggable policies in every seam" and be aware of the "reduced blast radius by context testability".
Just does not compute.
"Answer impersonally, objectively and analytically, without undue friendliness or enthusiasm. Use an engineering style response: concise, factual, and complete. Do not speak in the first person. Do not promote engagement or an emotional connection. Do not use emojis."
It’s not perfect, it has shortcomings, it sometimes produces bogus outputs. All of that is fine for a tool, it’s not fine when it pretends it’s a conscious being, because errors start to feel like lies and it becomes a bit too personal.
I don't think there's any correction that can return LLMs to a purely tool-space. Too many AI boyfriend/girlfriends.
This is peak “I bought into their marketing spiel” vibes. Congratulations.
https://www.motorcyclenews.com/news/2018/january/biski-jet-s...
Data is an individual. LLMs are time-shared hallucinatory algorithms that will ultimately be used to sell ad space.
It is Stupid to create "the machine that melts people's brains", but these assholes are trying to sell you "the machine that melts people's brains" and love to market it as being "anything the user wants".
We don’t thank our other tools like grep, or the compiler.
If you train your LLM to produce words that seem like human responses or conversations, they will. If you train them to seem like god-like superintelligences to people who think that’s what they’re creating, they will.
These are tools and it would behoove us all to keep that top of mind. Dangerous tools that are not your friend (but are useful as tools nonetheless)
At the end of "Do Androids Dream of Electric Sheep?" [if I'm remembering correctly] Deckard finds a toad in the desert. He gives it to his wife(!) and she immediately bonds with it. Later they realize that it is mechanical (they find a battery compartment). But Deckard's wife still treats it like a real creature and wonders whether it will eat mechanical flies.
The core theme of "Do Androids Dream..." is that humans display empathy towards others, not because others deserve it, but because we are human. To be human is to empathize with other creatures, and when we stop empathizing, we stop being human.
This morning I was working with Codex and we found the solution to a really nasty bug. I was so happy/relieved/excited to have found it, and I shared my excitement with Codex. I know Codex is just a big matmul. I know it doesn't experience joy or surprise or empathy. But I still want it to behave excited, not because it is human, but because I am.
Chatting with an agent using language feels like talking to an assistant.
I don’t think my meat brain is able to really differentiate between writing a message to an agent versus writing an agent to a human.
I’ll keep being polite and grateful to agents, so that I maintain those social habits for when I talk to humans.
My take is that people deserve to be treated with respect because they are people, not because they are smart and/or interesting. I don't treat people differently based on their IQ.
But, as I said, I totally understand your point of view, particularly the idea that people are different from machines and deserve to be treated differently.
An unlikely failure mechanism perhaps but it costs me nothing to avoid it.
I do that. But I also act rude and call it names at times. I think that is also what makes us human.
I don’t think LLMs are conscious but I’m still nice to them because it makes me feel bad when I’m rude to them.
An LLM is not a person. Do not humanise it. Do not personalise it. Do not treat it as more than a glorified autocorrect—that’s what it really is.
Humanising LLMs is exactly what these companies want, because then they can really take advantage of the I in AI which is exactly what makes people think these steroidal spellcheckers have any semblance of intelligence, person and personality.
Stop it.
Eventually, they’ll try “I’m hurt that you won’t consider my suggestion for <advertised product/political position>.” And people will fall for it.
Being a curmudgeonly asshole is apparently the key to surviving in this brave new world and I’m way ahead of the game!
On the other hand, what an LLM is trying to autocomplete is a story of a conversation between its user and a helpful assistant, so maintaining the decorum of office politeness will align more with how it's fine-tuned and produce higher quality output.
Taking LLMs for what they are is in no way near to thinking these pieces of software possess personality, intelligence and reasoning capabilities.
You don't know. You strongly feel that you know. There's no standard procedure to assess the existence of phenomenal consciousness.
But with what I know of people and what I know of matrix multiplication, I have a strong suspicion.
A matmul (or a fixed finite sequence of matmuls) is not Turing-complete. An LLM that produces answer in the number of steps that is proportional to the length of the answer is not Turing-complete. An LLM with CoT that produces answer after a number of steps, which is determined by the LLM itself, is Turing-complete.
So, we do have a qualitative step between a matmul and modern LLMs. Of course, it doesn't guaranty having feelings, but it doesn't allow to simply dismiss the notion due to sheer computational inadequacy.
A machine is not turing complete, but by running it recursively on its own output, turing completeness emerges.
Who knows at what stage of intricacy consciousness emerges?
Arguably, either it does at some point, or it never does and we requiere the supernatural soul.
So, if it does, we can't disregard the possibility that we've reached it already.
Please.
You know it doesn’t experience joy, surprise or empathy too. You absolutely, undoubtedly know that.
First things first. You are wrong. I don't have a gut feeling that LLMs have no phenomenal states. Not anymore. I was skeptical that LLMs can form complex internal representations. Mechanistic interpretability has shown that they can. I'm still skeptical in regard to introspective mechanisms. One token being fed back into the network and KV cache update do seem lacking in comparison to the ubiquitous recurrent connections in the brain.
What I do know is that philosophy hasn't produced an operationalized definition of qualia. So, we need to approach the problem of LLMs' phenomenal states from other angles.
For example, we can argue that LLMs lack something, which seems to be relevant to the problem: a biological substrate, generality (non-CoT LLMs aren't Turing-complete), "something something feedback"[1], and things like that.
"Admit that you know that" is boring.
[1] I'm referring to https://www.astralcodexten.com/p/the-new-ai-consciousness-pa...
Reality is under no obligation to be fun. It might be boring but it is the truth.
The fact that you needed such effort to try and give this position any shred of sense only serves as proof of how nonsensical it is to do so.
The inability of people to reflect on their own beliefs amazes me to no end. How can you know that LLMs have no phenomenal states?
For example, I believe that rocks (ordinary rocks, not silicon chips) have no phenomenal states because for all we know they don't do any kind of information processing (relevant to self-reflection, at least) and I don't believe in panpsychism or dualism.
The more we demonize AI, the more likely we are to demonize humans.
[Which is not to say we can't criticize AI companies or worry about the many downsides to the technology. But anger and fear are not the answer either.]
At the risk of sounding like an LLM myself: this is exactly the key most seem to miss
For example, I can log in and have my computer work perfectly and treat it as a device. But the moment something goes wrong and I am frustrated and confused at something unexpected or if i get in range, I find myself getting angry at the computer often even reffering to it as a being.
So eventually yes, we do anthromorphize because we are human but we do it because its a very important way we model the world around us.
Of course all this can also be used against you depending on how much the creator of the product wants people to form a bond with it.
Furthermore yes it is pretty possible we stop being human when we stop anthromorphizing but that's not true as well. Research has found people on the autism spectrum usually do not anthromorphize and have a problem with empathy because they cannot model others well or their own selves. But that doesn't make them non humans.
How many humans would you damn to save that mechanical toad?
I do ask it not to ask it follow on questions. I find that derails my own train of thought.
There's an important point in the article, that forcing a style onto an LLM is lossy. Although he doesn't seem to mention it, forcing a style may result in the insertion of new blithering, possibly made up as a hallucination.
eg., OpenAI has gone a long way to making reasoning token-efficient by having reasoning piovot off terse langauge -- whereas anthropic appears to be doing the opposite.
This is pretty unfortunate, because every LLM I've used has at least one tic that I find quite annoying. It would be great to be able to eliminate them. Sometimes I do, even knowing this drawback. But there is generally a cost. (Though I don't know if "lossy" is quite right, as that implies it's always a degradation. I think it's more likely to be harmful than helpful, given the models were tuned for their default state, but it is more of a random perturbation with a slight negative bias than a strict loss.)
AI labs can now ask the LLM to translate and filter the data, to create new training data that makes more sense and has better style.
I realize this isn't entirely serious, but I can't resist pointing out that this doesn't seem to be a good explanation for why LLMs write the way they do. When we've experimented with LLM writing style on open-weights models where you can get a base model (pretraining on text only) and an instruction-tuned variant (pretraining + post-training with RLHF and whatever other human-evaluated tasks), it's the instruction-tuned variant that shows the weird writing quirks. That is, the writing style is not because of the training texts, but because of whatever tasks the LLM companies do in instruction tuning. https://arxiv.org/abs/2410.16107
I'd speculate that this is partly impressed human preferences (the human raters unintentionally reward a particular writing style) and partly because of the chosen tasks: they're training the LLM to be good at, say, summarizing text, so it develops a style that's good at being informationally dense.
At any rate I've seen this same phenomenon with Llama and Gemma, and will be trying soon with Qwen. Unfortunately none of the commercial models lets you access the base model, as far as I know.
Now, the second example is the only thing that works. Power users have lost their powers with AI overview.
> That compression is lossy.
> You probably never notice what got dropped because the output still reads nicely.
> ASD-STE is a great example because it sounds so reasonable. It was designed to make documentation unambiguous for humans. But an agent isn’t a human technical writer, and the raw state is often the most information-dense representation available. Meanwhile the style rules sit on the same instruction list as: solve the task, use tools correctly, preserve abstractions, don’t break anything.
Author seems to have some misconceptions about LLMs. They already code-switch for us: the way they speak in chain-of-thought is completely different from the relatively normal language generated as human-facing output. You can observe this in any open-weight LLM, or in leaked CoT content from GPT5.x series etc: it's terse, barely follows sentence structure, lots of repeated checks and second-guessing.
On the next turn the model usually still has access to its previous turn's chain-of-thought, and I imagine that's what it'll use as reference, rather than the softer human-facing prose.
This being the case, asking the LLM to code-switch to an easier dialect for us doesn't seem that harmful.
For a more extreme example: if I talk to an LLM in Japanese then its response will be in Japanese, but its CoT will still be in either English or Chinese (depending on the model). These are two completely separate languages, but the LLM just kinda deals with it.
What I think the author (and I) are wondering about, is whether instructions like these might influence not _only_ the final output, but also the way it got there.
Yes, it really is. It's emitting literal English text in patterns that have been coaxed through RL into doing some kind of useful fuzzy computation.
> is the CoT just another presentation layer over the actual weights?
What does this mean? Which part of the forward pass are you talking about here?
> What I think the author (and I) are wondering about, is whether instructions like these might influence not _only_ the final output, but also the way it got there.
Sure, LLMs are chaotic. If I mention as a casual aside that the sky is blue, I'll get a different answer, even when my query has nothing to do with the colour of the sky. The fact that style instructions compete for attention with more concrete instructions is probably the one part of the author's post that I agree with. Different doesn't mean worse; if I run with greedy sampling (so deterministic) then minor punctuation differences in my query still produce completely different answers.
My reading is that they are actually making a more concrete point, which is: requesting a simpler style compresses the output, and this causes loss of fidelity in future turns. I think this is something you would have to demonstrate instead of hand-waving. I'm happy to be corrected on specifics.
Not humanising it...
People want the terse, matter-of-fact output. Not the conversational chatty verbose and bloated nonsense with gray words and jargon and terms like "blast radius"
The problem we have is a few large and expensive models are trying to be everything to everyone.
Not sure if I agree with a lot of this, even while spending a lot of time building with LLM, but was interesting nonetheless.
Is that a problem with https://code.claude.com/docs/en/output-styles?
> Output styles apply to the main conversation only: a subagent runs its own system prompt, so styles don’t change how subagents respond. A fork is the exception, because it inherits the parent’s full system prompt.
I'm sure you can make it to output the exact issue details as you want, but it starts with a human-like tone and waits for your requirement on depth and detail of the things. Another option is, just check the output of the traditional test runner (non-AI). It will give the full details.
Seems like something fixable with a simple two step process. Ask it the thing. Then ask it to summarise the answer in simpler terms. More tokens and time aside that would check both boxes
In other words, is there a way to keep the internal process intact up to the point of the formulation - and have only that vary.
It has worked great but i've spent more time beating LLM output into parseable output than I have reading and appreciating the prose it sends when i'm asking it something about some snippets of code.
It isn't deliberately unhinged like Steve Yegge's take: https://yegge.ai/essays/model-welfare/ In Steve's essay he starts with the assertion that agents are sentient... Whether or not that's true isn't really relevant, as his agent-flavored version of Pascal's wager actually holds water, especially for Anthropic models, as their system prompts already push the model in that direction, and it is better to work with them than try to prompt against the tide.
This way workers still operate completely in their preferred linguistic space. And you can safely mold the output stylistically however you like.
I don't think it's wise to take communication advice from someone so helplessly juvenile (and attention seeking) in their own communication attempts.
Would love to see some data backing how strong the effect is.
Yes, that's exactly what I want. I want to know what is done with a high level why, NOT a paragraph explaining each line of code modified.
A frequent tweak I've made on coding work with Claude in the last month is asking it to restrict its comment length to 1 line/sentence max. I just want `// This happens because XYZ upstream`, not `// Historically from ticket blahblah there was some dummy code where we discovered ancient runes and that led us to looking into your birth records and then triangulated an issue in XYZ upstream that we compensate for here`.
The model's "most information dense representation" is very similar to how we compress data in the first place: most of it is redundant or unnecessary for the purposes of storage. I'll "decompress" the 1-liner context myself when I read it again in 6 months. But I can't stand reading just so much slop commentary when we're all writing more code at once and having to review more than ever.
This is why /bro skill works.
https://github.com/backnotprop/bro/blob/main/skills/bro/SKIL...
If the author wants to read slop for hours, be my guest. Make it lossy, my job is not to read mimetic feelings, it's to make sure implementations get implemented.
Tell? Largest “tell”? Tell for me to tell?
Write in English, please:
“The biggest sign that shows me how culture and sentiment are changing, is…”
In other contexts it means a revealing signal.