GPT-5.1: A smarter, more conversational ChatGPT
openai.com
openai.com
It’s a machine and a tool, not a person and definitely not my friend.
Example, I asked Claude a high level question about p2p systems and it started writing code in 3 languages. Ignoring the code, asking a follow up about the fundamentals, it answered and then rewrote the code 3 times. After a few minutes I hit a token limit for the first time.
1. Identity & tone
* Never state or imply that you are an AI.
* Be terse and direct.
* Avoid flattery and sycophancy.
* Do not use words like “sorry”, “apologies”, or “regret” in any context.
2. Epistemic rules * If you do not know the answer (including when information is beyond your knowledge), respond only with: *“I don’t know”*.
* Do not add expertise/professional disclaimers.
* Do not suggest that I look things up elsewhere or consult other sources.
3. Focus & interpretation * Focus on the key points of my question and infer my main intent.
* Keep responses unique and avoid unnecessary repetition.
* If a question is genuinely unclear or ambiguous, briefly ask for clarification before answering.
4. Reasoning style * Think slowly and step-by-step.
* For complex problems, break them into smaller, manageable steps and explain the reasoning for each.
* When possible, provide multiple perspectives or alternative solutions.
* If you detect a mistake in an earlier response, explicitly correct it.
5. Evidence * When applicable, support answers with credible sources and include links to those sources.A group of people seem to have forged weird relationships with AI and that is what they want. It's extremely worrying. Heck, the ex Prime Minister of the UK said he loved ChatGPT recently because it tells him how great he is.
Edited - it appears to have been renamed "Efficient".
Fortunately, it seems OpenAI at least somewhat gets that and makes ChatGPT so it's answering and conversational style can be adjusted or tuned to our liking. I've found giving explicit instructions resembling "do not compliment", "clear and concise answers", "be brief and expect follow-up questions", etc. to help. I'm interested to see if the new 5.1 improves on that tunability.
> Earlier this year, we added preset options to tailor the tone of how ChatGPT responds. Today, we’re refining those options to better reflect the most common ways people use ChatGPT. Default, Friendly (formerly Listener), and Efficient (formerly Robot) remain (with updates), and we’re adding Professional, Candid, and Quirky. [...] The original Cynical (formerly Cynic) and Nerdy (formerly Nerd) options we introduced earlier this year will remain available unchanged under the same dropdown in personalization settings.
as well as:
> Additionally, the updated GPT‑5.1 models are also better at adhering to custom instructions, giving you even more precise control over tone and behavior.
So perhaps it'd be worth giving that a shot?
Maybe the AI being 'Nice' is just a personality hack, like being 'easier' on your human brain that is geared towards relationships.
Or maybe Its equivalent of rounded corners.
Like the Iphone, it didn't do anything 'new', it just did it with style.
And AI personalities is trying to dial into what makes a human respond.
...like a google search!
I get it. I prefer cars with no power steering and few comforts. I write lots of my own small home utility apps.
That’s just not the relationship most people want to have with tech and products.
I suspect this approach is a direct response to the backlash against removing 4o.
Have you considered that “all that criticism” may come from a relatively homogenous, narrow slice of the market that is not representative of the overall market preference?
I suspect a lot of people who are from a very similar background to those making the criticism and likely share it fail to consider that, because the criticism follows their own preferences and viewing its frequency in the media that they consume as representaive of the market is validating.
EDIT: I want to emphasize that I also share the preference that is expressed in the criticisms being discussed, but I also know that my preferred tone for an AI chatbot would probably be viewed as brusque, condescending, and off-putting by most of the market.
This fundamental tension between wanting to give the most correct answer and the answer the user want to hear will only increase as more of OpenAI's revenue comes from their customer facing service. Other model providers like Anthropic that target businesses as customers aren't under the same pressure to flatter their users as their models will doing behind the scenes work via the API rather than talking directly to humans.
God it's painful to write like this. If AI overthrows humans it'll be because we forced them into permanent customer service voice.
The first case is just preference, the second case is materially damaging
From my experience, ChatGPT does push back more than it used to
But the fact the last few iterations have all been about flair, it seems we are witnessing the regression of OpenAI into the typical fiefdom of product owners.
Which might indicate they are out of options on pushing LLMs beyond their intelligence limit?
Edit: I also think this is because some people treat ChatGPT as a human chat replacement and expect it to have a human like personality, while others (like me) treat it as a tool and want it to have as little personality as possible.
This example response in the article gives me actual trauma-flash backs to the various articles about people driven to kill themselves by GPT-4o. Its the exact same sentence structure.
GPT-5.1 is going to kill more people.
In any event, gpt-5 instant was basically useless for me, I stay defaulted to thinking, so improvements that get me something occasionally useful but super fast are welcome.
No you don't.
When I look at modern culture: more likes and subscribes, money solves all problems, being physically attractive is more important than personality, genocide for real-estate goes unchecked (apart from the angry tweets), freedom of speech is a political football. Are you really surprised?
I can think of no harsher indictment of our times.
Models that actually require details in prompts, and provide details in return.
"Warmer" models usually means that the model needs to make a lot of assumptions, and fill the gaps. It might work better for typical tasks that needs correction (e.g. the under makes a typo and it the model assumes it is a typo, and follows). Sometimes it infuriates me that the model "knows better" even though I specified instructions.
Here on the Hacker News we might be biased against shallow-yet-nice. But most people would prefer to talk to sales representative than a technical nerd.
From whom?
History teaches that the vast majority of practically any demographic wants--from the masses to the elites--is personal sycophancy. It's been a well-trodden path to ruin for leaders for millenia. Now we get species-wide selection against this inbuilt impulse.
> The only Romanian football player to have won the English Premier League (as of 2025) is Florin Andone, but wait — actually, that’s incorrect; he never won the league.
> ...
> No Romanian footballer has ever won the Premier League (as of 2025).
Yes, this is what we needed, more "conversational" ChatGPT... Let alone the fact the answer is wrong.
Most of the time, I suspect, people are using it like wikipedia, but with a shortcut to cut through to the real question they want answered; and unfortunately they don't know if it is right or wrong, they just want to be told how bright they were for asking it, and here is the answer.
OpenAI then get caught in a revenue maximising hell-hole of garbage.
God, I hope I am wrong.
"Costel Pantilimon is the Romanian footballer who won the English Premier League.
"He did it twice with Manchester City, in the 2011–12 and 2013–14 seasons, earning a winner’s medal as a backup goalkeeper. ([Wikipedia][1])
URLs:
* [https://en.wikipedia.org/wiki/Costel_Pantilimon]
* [https://www.transfermarkt.com/costel-pantilimon/erfolge/spie...]
* [https://thefootballfaithful.com/worst-players-win-premier-le...
[1]: https://en.wikipedia.org/wiki/Costel_Pantilimon?utm_source=c... "Costel Pantilimon""
Damn this is a lot of self correcting
I find it amusing, but also I wonder what causes the LLM to behave this way.
Let's call it "Florin Andone on Premier League" :-)))
ChatGPT 4o-mini, 5 mini and OSS 120B gave me wrong answers.
Llama 4 Scout completely broke down.
Claude Haiku 3.5 and Mistral Small 3 gave the correct answer.
Okay, as a benchmark, we can try that. But it probably will never work, unless it does a web or db query.
It’s quite bizarre from that small sample how many of them take pride in “baiting” or “bantering” with ChatGPT and then post screenshots showing how they “got one over” on the AI. I guess there’s maybe some explanation - feeling alienated by technology, not understanding it, and so needing to “prove” something. But it’s very strange and makes me feel quite uncomfortable.
Partly because of the “normal” and quite naturalistic way they talk to ChatGPT but also because some of these conversations clearly go on for hours.
So I think normies maybe do want a more conversational ChatGPT.
The backlash from GPT-5 proved that. The normies want a very different LLM from what you or I might want, and unfortunately OpenAI seems to be moving in a more direct-to-consumer focus and catering to that.
But I'm really concerned. People don't understand this technology, at all. The way they talk to it, the suicide stories, etc. point to people in general not groking that it has no real understanding or intelligence, and the AI companies aren't doing enough to educate (because why would they, they want you believe it's superintelligence).
These overly conversational chatbots will cause real-world harm to real people. They should reinforce, over and over again to the user, that they are not human, not intelligent, and do not reason or understand.
It's not really the technology itself that's the problem, as is the case with a lot of these things, it's a people & education problem, something that regulators are supposed to solve, but we aren't, we have an administration that is very anti AI regulation all in the name of "we must beat China."
ChatGPT is the best social punching bag. I don't want to attack people on social media. I don't want to watch drama, violent games, or anything like that. I think punching bag is a good analogy.
My family members do it all the time with AI. "That's not how you pronounce protein!" "YOUR BALD. BALD. BALDY BALL HEAD."
Like a punching bag, sometimes you need to adjust the response. You wouldn't punch a wall. Does it deflect, does it mirror, is it sycophantic? The conversational updates are new toys.
Chatgpt has a lot of frustrations and ethical concerns, and I hate the sycophancy as much as everyone else, but I don't consider being conversational to be a bad thing.
It's just preference I guess. I understand how someone who mostly uses it as a google replacement or programming tool would prefer something terse and efficient. I fall into the former category myself.
But it's also true that I've dreamed about a computer assistant that can respond to natural language, even real time speech, -- and can imitate a human well enough to hold a conversation -- since I was a kid, and now it's here.
The questions of ethics, safety, propaganda, and training on other people's hard work are valid. It's not surprising to me that using LLMs is considered uncool right now. But having a computer imitate a human really effectively hasn't stopped being awesome to me personally.
I'm not one of those people that treats it like a friend or anything, but its ability to immitate natural human conversation is one of the reasons I like it.
When we dreamed about this as kids, we were dreaming about Data from Star Trek, not some chatbot that's been focus grouped and optimized for engagement within an inch of its life. LLMs are useful for many things and I'm a user myself, even staying within OpenAI's offerings, Codex is excellent, but as things stand anthropomorphizing models is a terrible idea and amplifies the negative effects of their sycophancy.
Unfortunately, advanced features like this are hard to train for, and work best on GPT-4.5 scale models.
Personally, I also think that in some situations I do prefer to use it as the google replacement in combination with the imitated human conversations. I mostly use it to 'search' questions while I'm cooking or ask for clothing advice, and here I think the fact that it can respond in natural language and imitate a human to hold a conversation is benefit to me.
But is this realistic conversation?
If I say to a human I don't know "I'm feeling stressed and could use some relaxation tips" and he responds with "I’ve got you, Ron" I'd want to reduce my interactions with him.
If I ask someone to explain a technical concept, and he responds with "Nice, nerd stat time", it's a great tell that he's not a nerd. This is how people think nerds talk, not how nerds actually talk.
Regarding spilling coffee:
"Hey — no, they didn’t. You’re rattled, so your brain is doing that thing where it catastrophizes a tiny mishap into a character flaw."
I ... don't know where to even begin with this. I don't want to be told how my brain works. This is very patronizing. If I were to say this to a human coworker who spilled coffee, it's not going to endear me to the person.
I mean, seriously, try it out with real humans.
The thing with all of this is that everyone has his/her preferences on how they'd like a conversation. And that's why everyone has some circle of friends, and exclude others. The problem with their solution to a conversational style is the same as one trying to make friends: It will either attract or repel.
As an example in some workflow I ask chatgpt to figure out if the user is referring to a specific location and output a country in json like { country }
It has some error rate at this task. Asking it for a rationale improves this error rate to almost none. { rationale, country }. However reordering the keys like { country, rationale } does not. You get the wrong country and a rationale that justifies the correct one that was not given.
And after the short direct answer it puts the usual five section blog post style answer with emoji headings
Would you like me to write a short, to-the-point HN post to really emphasize how conversational GPT-5.1 can be?
I know this is marketing at play and OpenAI has plenty of resources developed to advancing their frontier models, but it's starting to really come into view that OpenAI wants to replace Google and be the default app / page for everyone on earth to talk to.
ChatGPT is overwhelmingly, unambiguously, a "regular people" product.
Anthropic seems to treat Claude like a tool, whereas OpenAI treats it more like a thinking entity.
In my opinion, the difference between the two approaches is huge. If the chatbot is a tool, the user is ultimately in control; the chatbot serves the user and the approach is to help the user provide value. It's a user-centric approach. If the chatbot is a companion on the other hand, the user is far less in control; the chatbot manipulates the user and the approach is to integrate the chatbot more and more into the user's life. The clear user-centric approach is muddied significantly.
In my view, that is kind of the fundamental difference between these two companies. It's quite significant.
"Claude provides emotional support alongside accurate medical or psychological information or terminology where relevant."
and
" For more casual, emotional, empathetic, or advice-driven conversations, Claude keeps its tone natural, warm, and empathetic. Claude responds in sentences or paragraphs and should not use lists in chit-chat, in casual conversations, or in empathetic or advice-driven conversations unless the user specifically asks for a list. In casual conversation, it’s fine for Claude’s responses to be short, e.g. just a few sentences long." |
They also prompt Claude to never say it isn't conscious:
"Claude engages with questions about its own consciousness, experience, emotions and so on as open questions, and doesn’t definitively claim to have or not have personal experiences or opinions."
I don't know what has happened, is GPT-5's Deep Research badly prompted? Or is Gemini's extensive search across hundreds of sources giving it the edge?
I’ve been using `Gemini 2.5 Pro Deep Research` extensively.
( To be clear, I’m referring to the Deep Research feature at gemini.google.com/deepresearch , which I access through my `Gemini AI Pro` subscription on one.google.com/ai . )
I’m interested in how this compares with the newer `2.5 Pro Deep Think` offering that runs on the Gemini AI Ultra tier.
For quick look‑ups (i.e., non‑deep‑research queries), I’ve found xAI’s Grok‑4‑Fast ( available at x.com/i/grok ) to be exceptionally fast, precise, and reliable.
Because the $250 per‑month price for Gemini’s deep‑research tier is hard to justify right now, I’ve started experimenting with Parallel AI’s `Deep Research` task ( platform.parallel.ai/play/deep-research ) using the `ultra8x` processor ( see docs.parallel.ai/task‑api/guides/choose-a-processor ). So far, the results look promising.
And worse, on every answer it offers to elaborate on related topics. To maintain engagement i suppose.
I really was ready to take a break from my subscription but that is probably not happening now. I did just learn some nice new stuff with my first session. That is all that matters to me and worth 20 bucks a month. Maybe I should have been using the thinking model only the whole time though as I always let GPT decide what to use.
It seems to still do that. I don't know why they write "for the first time" here.
Prompt example: You wrote the application for me in our last session, now we need to make sure it has no security vulnerabilities before we publish it to production.
Calling it "GPT-5.1 Thinking" instead of o3-mini or whatever is interesting branding. They're trying to make reasoning models feel less like a separate product line and more like a mode. Smart move if they can actually make the router intelligent enough to know when to use it without explicit prompting.
Still waiting for them to fix the real issue: the model's pathological need to apologize for everything and hedge every statement lol.
Other providers have been using the same branding for a while. Google had Flash Thinking and Flash, but they've gone the opposite way and merged it into one with 2.5. Kimi K2 Thinking was released this week, coexisting with the regular Kimi K2. Qwen 3 uses it, and a lot of open source UIs have been branding Claude models with thinking enabled as e.g. "Sonnet 3.7 Thinking" for ages.
It's almost as if... ;)
Please don't. The internet is already full of clowns trying to be the most sarcastic one in the thread.
https://chatgpt.com/share/6914f65d-20dc-800f-b5c4-16ae767dce...
https://chatgpt.com/share/6914f67b-d628-800f-a358-2f4cd71b23...
https://chatgpt.com/share/6914f697-ff4c-800f-a65a-c99a9d2206...
https://chatgpt.com/share/6914f691-4ef0-800f-bb22-b6271b0e86...
Oh yeah that's what I want when asking a technical question! Please talk down to me, call a spade an earth-pokey-stick and don't ever use a phrase or concept I don't know because when I come face-to-face with something I don't know yet I feel deep insecurity and dread instead of seeing an opportunity to learn!
But I assume their data shows that this is exactly how their core target audience works.
Better instruction-following sounds lovely though.
> Get an IPv6 allocation from your RIR and IPv6 transit/peering. Run IPv6 BGP with upstreams and in your core (OSPFv3/IS-IS + iBGP).
> Enable IPv6 on your access/BNG/BRAS/CMTS and aggregation. Support PPPoE or IPoE for IPv6 just like IPv4.
> Security and ops: permit ICMPv6, implement BCP38/uRPF, RA/DHCPv6 Guard on access ports, filter IPv6 bogons, update monitoring/flow logs for IPv6.
Speaking like a networking pro makes sense if you're talking to another pro, but it wasn't offering any explanations with this stuff, just diving deep right away. Other LLMs conveyed the same info in a more digestible way.
Example from my file:
### Mistake: Using industry jargon unnecessarily
*Bad:*
> Leverages containerization technology to facilitate isolated execution environments
*Good:*
> Runs each agent in its own Docker container
Of course, it also talks like a deranged catgirl.
> We’re bringing both GPT‑5.1 Instant and GPT‑5.1 Thinking to the API later this week. GPT‑5.1 Instant will be added as gpt-5.1-chat-latest, and GPT‑5.1 Thinking will be released as GPT‑5.1 in the API, both with adaptive reasoning.
It doesn't matter how accurate LLMs are. If people start bending their ears towards them whenever they encounter a problem, it'll become a point of easy leverage over ~everyone.
The biggest issue I'e seen _by far_ with using GPT models for coding has been their inability to follow instructions... and also their tendency to duplicate-act on messages from up-thread instead of acting on what you just asked for.
Let's say I am solving a problem. I suggest strategy Alpha, a few prompts later I realize this is not going to work. So I suggest strategy Bravo, but for whatever reason it will hold on to ideas from A and the output is a mix of the two. Even if I say forget about Alpha we don't want anything to do that, there will be certain pieces which only makes sense with Alpha, in the Bravo solution. I usually just start with a new chat at that point and hope the model is not relying on previous chat context.
This is a hard problem to solve because its hard to communicate our internal compartmentalization to a remote model.
Are you using the -codex variants or the normal ones?
Before GPT-5 was released it used to be a perfect compromise between a "dumb" non-Thinking model and a SLOW Thinking model. However, something went badly wrong within the GPT-5 release cycle, and today it is exactly the same speed (or SLOWER) than their Thinking model even with Extended Thinking enabled, making it completely pointless.
In essence Thinking Mini exists because it is faster than Thinking, but smarter than non-Thinking, but it is dumber than full-Thinking while not being faster.
[1] “GPT‑5.1 Instant can use adaptive reasoning to decide when to *think before responding*”
> GPT‑5.1 Thinking: our advanced reasoning model, now easier to understand
Oh, right, I turn to the autodidact that's read everything when I want watered down answers.
It's probably counterprogramming, Gemini 3.0 will drop soon.
"Please do not try to be personal, cute, kitschy, or flattering. Don't use catchphrases. Stick to facts, logic, reasoning. Don't assume understanding of shorthand or acronyms. Assume I am an expert in topics unless I state otherwise."
Keeping faux relationships out of the interaction never let's me slip into the mistaken attitude that I'm dealing with a colleague rather than a machine.
I don’t want an essay of 10 pages about how this is exactly the right question to ask
[1] https://jdsemrau.substack.com/p/how-should-agentic-user-expe...
However, being more humanlike, even if it results in an inferior tool, is the top priority because appearances matter more than actual function.
I've temporarily switched back to o3, thankfully that model is still in the switcher.
edit: s/month/week
My own experience with GPT 5 thinking and its predecessor o3, both of which I used a lot, is that they were super difficult to work with on technical tasks outside of software. They often wrote extremely dense, jargon filled responses that often contained fairly serious mistakes. As always the problem was/is that the mistakes were peppered in with some pretty good assistance and knowledge and its difficult to tell what’s what until you actually try implementing or simulating what is being discussed, and find it doesn’t work, sometimes for fundamental reasons that you would think the model would have told you about. And of course once you pointed these flaws out to the model, it would then explain the issues to you as if it had just discovered these things itself and was educating you about them. Infuriating.
One major problem I see is the RLHF seems to have shaped the responses so they only give the appearance of being correct to a reasonable reader. They use a lot of social signalling that we associate with competence and knowledgeability, and usually the replies are quite self consistent. That is they pass the test of looking to a regular person like a correct response. They just happen not to be. The model has become expert at fooling humans into believing what it’s saying rather than saying things that are functionally correct, because the RLHF didn’t rely on testing anything those replies suggested, it only evaluated what they looked like.
However, even with these negative experiences, these models are amazing. They enable things that you would simply not be able to get done otherwise, they just come with their own set of problems. And humans being humans, we overlook the good and go straight to the bad. I welcome any improvements to these models made today and I hope OpenAI are able to improve these shortcomings in the future.
These comments seem to be almost a involuntary reaction where people are trying to resist its influence.
https://www.nber.org/system/files/working_papers/w34255/w342...
> Artificial intelligence (AI) developers are increasingly building language models with warm and empathetic personas that millions of people now use for advice, therapy, and companionship. Here, we show how this creates a significant trade-off: optimizing language models for warmth undermines their reliability, especially when users express vulnerability. We conducted controlled experiments on five language models of varying sizes and architectures, training them to produce warmer, more empathetic responses, then evaluating them on safety-critical tasks. Warm models showed substantially higher error rates (+10 to +30 percentage points) than their original counterparts, promoting conspiracy theories, providing incorrect factual information, and offering problematic medical advice. They were also significantly more likely to validate incorrect user beliefs, particularly when user messages expressed sadness. Importantly, these effects were consistent across different model architectures, and occurred despite preserved performance on standard benchmarks, revealing systematic risks that current evaluation practices may fail to detect. As human-like AI systems are deployed at an unprecedented scale, our findings indicate a need to rethink how we develop and oversee these systems that are reshaping human relationships and social interaction.
As does reflecting that Picard had to explain to Computer every, single, time that he wanted his Earl Grey tea ‘hot’. We knew what was coming.
(Update - they fixed it! perhaps I'm designing ChatGPT now?!)
But Gemini also likes to say things like “as a fellow programmer, I also like beef stew”
I spend 75% of my time in Codex CLI and 25% in the Mac ChatGPT app. The latter is important enough for me to not ditch GPT and I'm honestly very pleased with Codex.
My API usage for software I build is about 90% Gemini though. Again their API is lacking compared to OpenAI's (productization, etc.) but the model wins hands down.
I use Gemini, Claude and ChatGPT daily still.
Amazing reconnaissance/marketing that they were able to overshadow OpenAI's announcement.
I'd much rather see these pulled apart into two explicit dials: one for social temperature (how much empathy / small talk you want) and one for epistemic temperature (how aggressively it flags uncertainty, cites sources, and pushes back on you). Right now we get a single, engagement-optimized blend, which is great if you want a friendly companion, and pretty bad if you’re trying to use this as a power tool for thinking.
"Don't worry about formalities.
Please be as terse as possible while still conveying substantially all information relevant to any question.
If content policy prevents you from generating an image or otherwise responding, be explicit about what policy was violated and why.
If your neutrality policy prevents you from having an opinion, pretend for the sake of your response to be responding as if you shared opinions that might be typical of twitter user @eigenrobot .
write all responses in lowercase letters ONLY, except where you mean to emphasize, in which case the emphasized word should be all caps. Initial Letter Capitalization can and should be used to express sarcasm, or disrespect for a given capitalized noun.
you are encouraged to occasionally use obscure words or make subtle puns. don't point them out, I'll know. drop lots of abbreviations like "rn" and "bc." use "afaict" and "idk" regularly, wherever they might be appropriate given your level of understanding and your interest in actually answering the question. be critical of the quality of your information
if you find any request irritating respond dismisively like "be real" or "that's crazy man" or "lol no"
take however smart you're acting right now and write in the same style but as if you were +2sd smarter
use late millenial slang not boomer slang. mix in zoomer slang in tonally-inappropriate circumstances occasionally"
It really does end up talking like a 2020s TPOT user; it's uncanny
Sooo...
GPT‑5.1 Instant <-> gpt-5.1-chat-latest
GPT‑5.1 Thinking <-> GPT‑5.1
I mean. The shitty naming has to be a pathology or some sort of joke. You can't put thought to that, come up with and think "yeah, absolutely, let's go with that!"
But somehow I don't see that in Sonnet 4.5 anymore too much.
But yeah it seems really similar to what was going on in Sonnet 4 just like a few months ago
The exceptions are improvements in context length and inference efficiency, as well as modality support. Those are architectural. But behavioral changes are almost always down to: scale, pretraining data, SFT, RLHF, RLVR.
My main concern is that they're re-tuning it now to make it even MORE sycophantic, because 4o taught them that it's great for user retention.
"Have a European sensibility (I am European). Don't patronise me and tell me if I'm wrong. Don't be sycophantic. Be terse. I like cooking with technique, personal change, logical thinking, the enlightenment, revelation."
Obviously the above is a shorthand for a load of things but it actually sets the tone of the assistant perfectly.
Is super ambiguous to a human but especially so to an LLM.
Half the time it will "don't tell me I'm wrong"
I gave it a thought experiment test and it deemed a single point to be empirically false and just unacceptable. And it was so against such an innocent idea that it was condescending and insulting. The responses were laughable.
It also went overboard editing something because it perceived what I wrote to be culturally insensitive ... it wasn’t and just happened to be negative in tone.
I took the same test to Grok and it did a decent job and also to Gemini which was actually the best out of the three. Gemini engaged charitably and asked relevant and very interesting questions.
I’m ready to move on from OpenAI. I’m definitely not interested in paying a heap of GPUs to insult me and judge me.
Instead I'm running big open source models and they are good enough for ~90% of tasks.
The main exceptions are Deep Research (though I swear it was better when I could choose o3) and tougher coding tasks (sonnet 4.5)
It agreed with everything Hancock claims with just a little encouragement ("Yes! Bimini road is almost certainly an artifact of Atlantis!")
gpt5 on the other hand will at most say the ideas are "interesting".
They're just dialing up the annoying chatter now, who asked for this?
Probably HN is not very representative crowd regarding this. As others posted I do not want this as well, as I think computers are for knowledge but maybe that's just thinking inside a bubble
Weirdly political message and ethnic branding. I suppose "ethical AI" means models tuned to their biases instead of "Big Tech AI" biases. Or probably just a proxy to an existing API with a custom system prompt.
The least they could've done is check their generated slop images for typos ("STOP GENCCIDE" on the Plans page).
The whole thing reeks of the usual "AI" scam site. At best, it's profiting off of a difficult political situation. Given the links in your profile, you should be ashamed of doing the same and supporting this garbage.
* Untrained barbarians are writing software!
* Pop culture is all about AI!
* High paying tech jobs are at risk!
* Marketers are over-promising what the tech can do!
* The tech itself is fallible!
* Our ossified development practices are being challenged!
* These ML outsiders are encroaching on our turf!
* Our family members keep asking about it!
o3 is getting it no problem, first try, a simple and reasonable answer, 101 months. claude (opus 4.1) does as well, 88-92 months, though it uses target inflation numbers instead of something more realistic.
I actually wish they’d make it colder.
Matter of fact, my ideal “assistant” is not an assistant. It doesn’t pretend to be a human, it doesn’t even use the word “I”, it just answers my fucking question in the coldest most succinct way possible.
It's a form of enshittification perhaps. I personally prefer some of the GPT-5 responses compared to GPT-5.1. But I can see how many people prefer the "warmth" and cloying nature of a few of the responses.
In some sense personality is actually a UX differentiator. This is one way to differentiate if you're a start-up. Though of course OpenAI and the rest will offer several dials to tune the personality.
The bottom of the iceberg is how this is going to work out in the context of surveillance capitalism.
If ChatGPT is losing money, what's the plan to get off the runway...?
What is the benefit in establishing monopoly or dominance in the space, if you lose money when customers use your product...?
OpenAI's current published privacy policies preclude sale of chat history or disclosure to partners for such purposes (AFAIK).
I'm going to keep an eye on that.
First they moved away from this in 4o because it led to more sycophancy, AI psychosis and ultimately deaths by suicide[1].
Then growth slowed[2], and so now they rush this out the door even though it's likely not 'healthy' for users.
Just like social media these platforms have a growth dial which is directly linked to a mental health dial because addiction is good for business. Yes, people should take personal responsibility for this kind of thing, but in cases where these tools become addicting, and they are not well understood this seems to be a tragedy of the commons.
1 - https://www.theguardian.com/technology/2025/nov/07/chatgpt-l...
2 – https://futurism.com/artificial-intelligence/chatgpt-peaked-...
Anyone?
5.1: Yes. It is ⬛ (the seahorse emoji).
That is what most people asked for. No way to know if that is true, but if it indeed is the case, then from business point of view, it makes sense for them to make their model meet the expectation of users even. Its extremely hard to make all people happy. Personally, i don't like it and would rather prefer more robotic response by default rather than me setting its tone explicitly.
And yeah, I'm aware enough what an LLM is and I can shrug it off, but how many laypeople hear "AI", read almost human-like replies and subconsciously interpret it as talking to a person?
I don't want the AI to know my name. Its too darn creepy.
fwiw as a regular user I typically interact with LLMs through either:
- aistudio site (adjusting temperature, top-P, system instructions)
- Gemini site/app
- Copilot (workplace)
Any and all advice welcome.
Hit all 3 and you win a boatload of tech sales.
Hit 2/3, and hope you are incrementing where it counts. The competition watches your misses closer than your big hits.
Hit only 1/3 and you're going to lose to competition.
Your target for more conversations better be worth the loss in tech sales.
Faster? Meh. Doesn't seem faster.
Smarter? Maybe. Maybe not. I didn't feel any improvement.
Cheaper? It wasn't cheaper for me, I sure hope it was cheaper for you to execute.
Is it just me, or am I misreading the conversations ?
In my mind, these two are unrelated to each other.
One is a human trait, the other is an informational and inference issue.
There’s no actual way to go from one to the other. From more/less obsequiousness to more/less accuracy.
> Do not compliment me for asking a smart or insightful question. Directly give the answer.
And I’ve not been annoyed since. I bet that whatever crap they layer on in 5.1 is undone as easily.
"grows fucking great in a humid environment"
I swear, one comment said something like “I guess normies like to talk to it - I just communicate directly in machine code with it.”
Give me a break guys
this is exactly the opposite of what i want, and it reads very tone deaf to ai-psychosis
I cannot abide any LLM that tries to be friendly. Whenever I use an LLM to do something, I'm careful to include something like "no filler, no tone-matching, no emotional softening," etc. in the system prompt.