They all have various strengths and weaknesses. My favorite is still ChatGPT, then Gemini/Claude, then Grok.
Grok often feels 1-2 generations behind the competition in general use, but it has three things that I love:
1. It seems to be the best at understanding current events. Maybe due to X integration, or some other tool call optimization in the backend? I don't know, but I often ask about things going on, and the other models have outdated info, give unhelpful answers, etc.
2. It is generally the least sycophantic for personal things. Anthropic is getting here too. ChatGPT and Gemini are working on this, but previous models in those families would almost never say anything negative about what I am doing. Sometimes I need career advice, personal advice, etc and I like the tone of how it responds. I think Claude will be caught up soon.
3. For professional work, there are certain topics that other models would refuse to engage with. At my last company we had an enormous amount of legal users. When a deposition would need a summary on certain topics, most models would refuse. Grok would not. I understand the need for safety and I don't blame the other model providers, but for some professional use cases you NEED a model that is capable of handling sensitive subjects.
You don't think people talking about the car doing things has anything to do with anthropomorphising the car?
> avoid changing any of the circumstances that cause the behavior
The normalisation of unsafe driving is the circumstance that causes the behaviour. Just look at how the cultural shift in how drink-driving is perceived over the last few decades has changed the rate of it happening.
Grok used to be really really bad ~8 months ago or so, but it's gotten better.
ChatGPT team needs to turn down the 'disagree just because' factor by a lot.
As soon as we hand over searching out information to social media algorithms and LLM tools, we abandon our ability to see reality outside our direct vision.
Grok's ownership has already demonstrated capacity to influence major world elections and other events. You cannot trust it with this sort of information gathering and reporting.
That makes sense, but occasionally you ask about an issue where it's clearly received political instruction from the commissar and it acts totally lobotomized. But it's true that Gemini will often blithely state that something could never happen and you'll say "what do you mean, that just happened" and then it comes back apologizing after running a Web search.
I almost exclusively use claude for all my professional and private needs. In my experience it's really good at adhering to my wishes in regards to sycophancy and pushing back. If you really want to you can tell it to systematically push back on anything where pushback makes sense until it continues with the flow of conversation.
In my first therapy session, the answers were too long and contained multiple questions, spawning multiple threads of conversation. I told it to tone it down and only ever ask one question back, maybe two, if they are related. The answers got too short. I told it to make them "slightly longer" again and reached a sweet spot.
The conversation is yours to form! You need to find the "system prompts" and guidelines to give it that work for you.
I guess the benchmarks disagree, but whenever I need to find specific information that does not easily show up with a web search, I try chatgpt, gemini and grok. Grok surfaces what I was looking for more often than the others.
Things like "find the github repo from 2017 that does $vague_thing".
Come on, the most logical thing is that Musk overestimated the compute he needs and got lucky with the secondary usage of it.
As soon as the IPO is done and if it didn't fail, he will buy curser and try to push again if he hasn't given up on it.
He also needs some compute for the robotics stuff and for Tesla in-car entertainment and for training FSD.
So they’re cutting edge in that way.