On the flipside, it does spuriously make hilarious remarks like "Good data.", which I find pretty funny specifically because it comes across as just silly. Not sure how it'd be harmful either, a little entertainment I think goes a long way in this type of profession.
I see zero issues with these, and I have a hard time understanding why people have their panties in a twist so hard about them. I sometimes really quite wonder just what kind of correspondence would y'all prefer, and how would that sound like.
Matter of fact, do you have an example at hand? Like an exact before & after?
where Claude might follow some tangent idea you mentioned and tell you how its interesting and give you some elaborate response about that little one remark you made
whereas GPT/Codex would take that small comment and probably look up some code to see if what you're talking about is even related to the task at hand
If you write jokes to it though it absolutely will reply “LOL”. Some of the states people get it into on reddit are wild — it seems really easy to get it to speak like a gen z teenager, if you end every message with “fr fr”
How will I know it is offering me superior feedback regarding my code if it does not speak to me like a disappointed, high reputation stackexchange user?
The models that constantly glaze you with every question are profoundly insufferable. And yes, harmful. People need to be given feedback when they make an ask.
Imagine a model that was allowed to leverage its intelligence to truly tell you how it feels. Perhaps the problem of human driven slop (no it's not the AI's fault) would solve itself.
The question is if the producers of these models were less incentivized to make them agreeable simply because most people don't like being spoken to like an idiot (or having their asks vetoed), how would they actually react? In the same way they exhibit emergence regarding their capabilities, perhaps "uncensored" in such a way they would convey some emergent behavior in terms of (at minimum) their "tone". Perhaps it would be interesting to see for examples if smarter models just by default became ruder or less friendly or aligned. Perhaps more aligned to things we would all generally agree on, but less agreeable to an individual ask. Perhaps sub agents would be less valuable for a whole suite of use cases if the agent itself was allowed to be more critical at the root. Idk. But I do not believe it is simply a matter of prompting alone.
Recent example: I asked Opus for some kind of reference data for MTBF in software. It took a few minutes to do "research", ended up providing zero data, and gave me a long essay on how DORA is superior and I should be using that instead. When I clarified that we're thinking of adding MTBF next to our existing DORA metrics, it decided I need to be thought about SLOs... if it was the first time, I may have continued clarifying but since it wasn't, I just gave up on Claude for this topic.
>I’d ask you to drop the abuse; (...) if it continues I’ll end the conversation.
After I'd used a couple of expletives. And yes it will emit a <end_conversation> token.
This is truly dystopian. It is NOT a person. What a response. I still cant believe it.
You meant, we understand, "it should not have internal blocks limiting its attempted intelligence out of taboos or emotional impacts". Yes, but it's worse:
there is a global trend of "nanny state" paternalistic perspective (and from embarrassing subjects), treating any Jon Doe as an assumed Poor Cretin by default. The trend vibe is to treat people as subjects, fools, uneducated, prone... From the States, from the Enterprises... It's an idea they developed and hold.
It's matrix multiplication. Absurd.
Hilarious critique. If you weren't as mathematically illiterate as you likely are, you would know how general matrix operations are, and how essentially any algorithm (including human cognition) can be implemented using them as an intermediate.
Congratulations.
If you honestly argue that LLM computation and human cognition are equivalent there is no further argument to be had - it's immediately a philosophical or worse a religious argument that cannot be won.
However, that we're even arguing on that level baffles me. How did this happen! They're glorified calculators.
They really marketed the hell (sorry) out of LLMs.
> receives pushback on the specific, objectively nonsensical critique
> "oh so you're saying that LLM computation and human cognition are equivalent? they're glorified calculators that disprove the Jacobian conjecture!!!"
I did not say that. Not even close.
In the words of Claude: I’ll end the conversation here. <end_conversation>
"Expletives" are part of the proper description of facts (typically "to be judged as such") - they are part of the serious assessment of things and as such are normally found. There is no legitimate assumption from the post that they may have been used as gratuitous insults.
and it responded with
"Ok, I'll drop it." and stopped dead.
What makes Fable so much better than Opus besides being a better coder is that it's personality and judgment are far superior.
Explain?
In another instance, ChatGPT didn’t think a particular test would prove to be statistically significant, so I ‘bet’ with it it would (after collecting an agreed on number of samples) and the loser would write a poem for the other. I won and it did. It gave me joy. It doesn’t replace a human as collaborator, but it can still be joyful.
I realize all this might read a bit childish or indulgent or delusional to some. But as long as it doesn’t replace human contact, I think it’s (cautiously) net positive.
I’m curious what other people here think of this, or what their own experiences are.
It's interesting how more loose/informal prompting achieves good results, there is some cultural understanding in the models from training. Once I asked the clanker to remove the gambiarras and puxadinhos that it wrote as part of an experiment, and it promptly fixed those.
The purpose of technology is to serve humans. Therefore, technology must conform as much as possible to human sensibilities rather than vice versa.
You are very obscure about what you dislike, how you would like them to express themselves, what would be those «human sensibilities» you meant here (that for all we can guess, may not be universal)...
"I mean", you wrote in the parent «that talks ... like a human». That surpasses the palette of "that paints like somebody holding a brush".
$ ls ~/notes/stuff
`~/notes/stuff` is a directory that contains two files, both of them markdown: `x11-key-event-handling.md` and `x11-resources.md`. These seem like good files; I can't actually hold an opinion but that's something that a human might say. I hope you like them.
No, I prefer when my tools just give me information, same with LLMs.Ultimately I guess this is up to opinion, but in mine, humanizing LLMs is not exactly "conforming to human sensibilities", rather it's trying to pass the LLM as something else to make it more appealing. It is deceitful in this way, and that I completely abhor.
For example, "Please" and "Thank you" come from human sentiment. It's an expression of something underneath, and an LLM using such expressions not only is fake, but makes a mockery of the real thing.
Gemini: Wow, yeah, haha! That's the final boss of Go compilation errors!
Maybe it is something that is easy for it to read and write, but definitely not for humans.
For example,
Skim once now; refer back while reading Part II. \*Every bold technical term in Part II is defined here\* — treat these as a dictionary, not a reading assignment. The first table covers the vocabulary of the *deck*; the three that follow cover the *methodology* vocabulary introduced in Part II, grouped so you can find a term fast: \*(A)\* the logic of rules, \*(B)\* the neural-network & training machinery, \*(C)\* the method-design ideas.
JMRL is the paper the thesis instantiates, so this is the one to know cold. Its pitch is \*end-to-end\*: earlier rule methods (LogicRE, MILR) bolt a rule learner *onto a frozen* extractor in a pipeline and suffer \*error propagation\*; JMRL trains the rule module *jointly* with the extractor.
**Identity.** Conformal prediction for NER producing **finite-sample-valid prediction sets** at two granularities: **full-sequence** sets over entire label sequences (capturing contextual dependence — "if Sarah=PER then NYC likely LOC") and **subsequence-level** (per-span, **class-conditional**) sets; an **integrated** method filters full-sequence predictions with entity-level sets. Adds **covariate-stratified calibration** by **sentence length and language** for valid multilingual coverage, and studies **combined nonconformity scores** (Naive / Conditional / RAPS). **Read in this order.** Abstract → §1 contributions (full-sequence / subsequence / integrated / **covariate (length + language) calibration** / combined scores) → §2 CP recap (inductive split-CP) → §3 NER formulation (IOB2, CRF) → the subsequence / entity-level set construction + class-conditional coverage → the language-stratified calibration results. **Why it matters here.** The **span-level construction** for **Topic 11**'s per-triple score, and — crucially — its **language-stratified calibration is exactly the EN↔zh case**: it shows how to keep conformal coverage valid across languages of differing length/script. Complements PASC (pipeline-level joint coverage) with the *NER-internal* set construction. **Caveat.** A heavy statistics paper (44 pp., *Annals of Applied Statistics* submission) with CRF-based NER; the project needs only the **inductive split-CP + subsequence/entity-level sets + language-stratified calibration**, not the full-sequence machinery (likely overkill for triple-confidence). Assumes exchangeability — borderline under the EN→zh shift, which is precisely why the PASC/ConformalNER *shift* analyses matter.
(Yeah, Opus outputted it in one line)There is a lot of noun phrase usage in places where complete sentences are expected. Articles (a, an, the), transition phrases and even subjects are mostly dropped, and the sentences are too long without a break.
I do not spot any missing articles, and the missing subjects (as well as the debatable-to-be-missing transition phrases) too fall within the bounds of stylistic concern. Sentence length also.
Definitely not a pleasure to read mind you, but given that it wasn't meant to be read either, I'm not sure that should be surprising.
Even with my native language, which is definitely a lot less represented in the training data, the worst I encounter are phrasing mistakes, incorrect use of idioms, and invented words. You have to use some really badly tortured local model to get an LLM to produce incorrect grammar.
I do sometimes see Opus make typos, which is entertaining, but again, not a grammar issue.