Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index
artificialanalysis.ai
artificialanalysis.ai
I found it has improved my productivity and output over claude where it felt like claude was giving me work to do. furthermore with the recent claude watermarking thing, I'd rather use grok or openai.
If anyone is curious download grok cli and throw a couple of prompts at it. you'll be surprised.
I am curious, is that a plugin, or skill?
He's become a toxic "brand".
but what's to validate any given population as centrist anyway
they suggested the ai opinion drift was caused by internet demographics not directly reflecting actual population i.e. California publishes more etc
[1] https://futurism.com/future-society/elon-musk-ai-woke-free-o...
My token spend at api rates is about $3000 usd a month recently.
On a similar note, how much Elon hate is about his politics vs his trillionaire status vs what I like to call “watercooler hate” where folks simply parrot the loudest opinion in order to be accepted into the group?
On a final note, my son is in primary school and recently brought up in a dinner time discussion that “Elon is really bad” - this is a kid who has no social media (unlike some of his peers who are already on TikTok) and doesn’t watch traditional media.
He's also publicly and militantly supported quite a few far right political parties throughout Europe, particularly those who have a long track record on race-based topics.
Beside don't read too much in to benchmarks. They are alrrady ruined by Goldhart principle. Tgese models have already seen most of the data.
ChatGPT is very long-winded, sometimes producing multiple bullet point lists for a simple answer. Claude is full of mannerisms: 'not merely x, but y', 'Here's where it gets interesting', 'the real question is', etc.
Claude suggests something, then later on suddenly it's something I wanted all along. ChatGPT is far superior to Claude. I will try Grok.
Right now, I find that Grok offers better value, uses fewer tokens per turn, and makes better code. I haven’t tried Cursor because I don’t want to change editors again, but maybe I should try it…
https://artificialanalysis.ai/models/grok-4-6#token-use
So depending on how you want to define "token efficiency", Grok is either tied with OpenAI, or in the lead.
[1] Though I grant that 4.6 appears to be wordier, on the order of Terra max.
https://github.com/features/copilot/plans
https://github.blog/changelog/2026-08-06-kimi-k3-is-now-avai...
Have they converted entirely to transparent API rates + base allocation now? One of the reasons I left was that if I was going to be billed at API rates anyway, I'd just rather use the APIs. The value proposition still sucks for individuals now, when the other major providers are bundling at below-API rates.
On the OpenCode Go workspace there is a huge box:
Providers
Control which providers are used for routing.
Enable models hosted in China <toggle-button>
If you turn this off, DeepSeek Flash & Pro both stop working: Error: Provider request failed with HTTP 403: The latest version of this model is only available hosted in China and requires explicit opt in: https://opencode.ai/workspace/wrk_msh_is_a_liar_stop_bullshitting_please_be_real_example_url/go
Kimi K3, MiMo V2.5 Pro, MiniMax M3, Qwen3.8 Max all work (remarkably!) without the box ticked! That was a strong surprise for me.My request stands: I would like to have up-front information on what models only have Chinese providers. This information is not available except by buying a plan and experimenting, currently. And I had to try each one to find out, which I wish could be avoided. I want to end this thread where I started it (now that I have done the work to cut through the din and noise and misinformation), with my original request: please OpenCode Go make the information about which providers have Chinese-only/non-China hosting readily available on your website.
Apparently 5x usage when using Kimi Code too.
If you willing to share to no zdr, meta is waaaaaay cheaper vs Kimi.
With recent offerings from spacex and meta , I hardly imagine why would you pay money to any Chinese vendor it’s not as cheap and it’s not as intelligent neither .
Maybe deepseek is an exception , but it’s only good for narrow use cases that probably goes into modal.com and other gpu + fine tune me easy vendors , not vanilla dumb but cheap model .
I'm willing to pay 2x for a 10% smarter model. Intelligence matters that much (because 10% smarter probably saves, on average, several hours of human time).
I haven't tried so it's pure speculation based on benchmarks, but I'd assume Grok 4.6 is around Opus 4.8 in real world use, but clearly below Opus 5.
For building a full stack custom CRM and media pipeline tool with video conversion, transcription, and indexing. Supabase, AWS, Meili, NextJS, GCS - lots of surfaces and planes.
4.8 basically couldn't do it, I abandoned the project as the fallback was, "current business processes".
With F5 it's been 4 weeks and almost ready for production release.
It’s not as quite as smart as opus 4.8 but it’s close and x4 the cheaper.
Until it doesn't...
Honestly, this entire OpenAI reset credit fiasco this past week has convinced me to rip off the Codex and Claude Code bandaids and start building my own proper Pi Coding Agent running models that I select and pay for on openrouter.
And I am feeling a lot better about it now that I've finally got it working.
Deepseek V4 Flash 0731 is surprisingly capable and cheap. [0]
Checkout pi coding agent. You can create as many different sub-agents as you wish, to specialize and understand and tackle or pass off any problem you like. It's refreshing, really. I feel like a coder in control again.
Huh, what's happened? I'm on the 20x plan and haven't noticed any fiasco, what went down exactly?
But since you don't know when resets are coming, it becomes kind of frustrating trying to game it out. One week they're resetting like crazy so you're trying to burn tokens as fast as possible. The next week they're not resetting at all so you have to adjust your workflow to be more conservative.
This whole thing sounds crazy to me, why are you so focused on making sure you hit 0% usage left when it's supposed to reset? Why can't you just use what you have and if it resets, it resets, and if it doesn't, it doesn't?
I've calculated I can spend about ~12% of usage every day, more or less, this is my "budget". If it resets, then the "12% per day" gets reset for that day, but that's it. Sounds crazy to me that I'd "invent work out of nowhere" just to spend more usage, why on earth would I adopt such a workflow? Sounds like you're burning tokens just to burn tokens???
Sure, in a video game where you have one attribute and you can minmax a strategy just focusing on that, but that's not how real-life works.
You can't just stack pending work on top of each other, expect yourself to be able to stay equally on top of everything and have the same results as if you didn't. With this comes the consideration about the tradeoffs of "produce mediocre but large body of works" vs "produce high quality but small body of works".
Who's to say what's more "rational" or not, it's not a straight-forward calculation which culminates in "Must consume all available usage to maximize resource usage" like some robot, as we are not.
It's quite literally inventing work out of nowhere as you wouldn't put the agent to do that and forcing yourself to be conscious about that work until it completes, unless you actually had the usage available. Asked another way, wouldn't your workflow clearly change if you had unlimited usage available? You'd probably attack tasks/problems that you didn't consider actually spending time/energy solving.
Not really. The AI will output the same sort of code at a certain skill level regardless of speed, it's not a human so the above is a false dichotomy. Also, it's sort of strange that you're saying it's inventing work out of nothing, you've never heard of a backlog? In companies that can grow very long and the more usage means the more that can be tackled.
> You'd probably attack tasks/problems that you didn't consider actually spending time/energy solving.
Yes? But that's probably after the backlog is complete unless they are truly low hanging or high priority fruit. So not sure how that reasons with your point.
It'll output the same code given the same prompts yes, but you don't just accept whatever it puts out, it requires iterations before it's actually ready to be committed as none of the agents write perfect code on their first try. So, it's not a "false dichotomy", I'm just looking at larger things than "LLM does inference"
I don't get the point of this. We all seem to agree that these companies have almost no moat, if one stops being a good deal, you can switch to another. That doesn't invalidate the existence of a deal that is currently good.
My point was that chasing deals like this is just kicking the can down the road. You're going to have to reckon with harsh price increases sooner or later.
So I have resolved to avoid that future-dreading and fixed it, basically.
Seems weird to me to not take advantage of the great deal the frontier labs are currently giving for subscription pricing when there's such an easy fallback in the worst case.
the rest of the comment is about moving from expensive switching costs of what harness you are using to having the difference be closer to changing some config in openrouter
https://www.theguardian.com/technology/2026/jan/15/elon-musk...
It might or might not work long-term, but I wouldn't count on courts to help.
This is a myth.
I’m confused by how your reply makes sense in the context of the parent comment. They weren’t stating a preference, they were linking to simple facts.
Please do better.
They claim that spacex has competitive advantage. They take no stance on condoning spacex’s behavior. Personally I strongly dislike musks behavior, but I appreciate discussion of competitive advantages/disadvantages, independent of moral views. I liked the thread, it adds new info to the conversation.
No, they're not. You're making up things and pretending that they said them.
> Please do better.
You're badly breaking the HN guidelines. Please review them: https://news.ycombinator.com/newsguidelines.html
This is factually incorrect. I did not say anything "snarky" or "mean". I stated facts and pointed out that the poster was breaking the guidelines.
They obviously have a huge token cost advantage of the AI labs they are renting compute to, at least for now while they can charge current crazy rates for GPU compute.
they just increased cache read from 0.30 to 0.50 - this has the biggest impact on agentic coding. Elon companies have the most expensive everything:
xAI sub: $30 when other starts at $20, pro like sub for $300 where other charge $200.
Expensive electric cars, powerwalls, solar roofs when competetive products/better are cheaper.
The Model 3 and Model Y became the highest selling EVs of all time because they were the first below $50K to have long-range and be worth buying.
Until a few years ago, every other sub-$50K EV absolutely sucked.
And then Musk totally abandoned Tesla's original brilliant game plan of using the luxury models to find actual low-cost models (which $50K is not), and completely ceded the future EV market to Chinese companies that understand how to make a better car than Tesla for less money. BYD sells more cars than Tesla.
Model 3 starts at $37K, and Model Y at $40K.
They are #1 and #2 best-selling EVs of all time! So clearly they are affordable.
>BYD sells more cars than Tesla.
Sure, but only by 20%, and only because it also sells low-cost EVs that have less range than anything Tesla sells.
If you exclude BYD EVs with lower range than any Tesla, Tesla outsells BYD by 50%.
Even if you don't, second place (which Tesla still holds) is really not "completely ceded"!
This is meaningless because you're not controlling for amount of subsidized usage or model quality.
To say that I somewhat doubt this would be an understatement.
Google would like a word. Also, Microsoft.
It's good to have model diversity. When I run a task across Sol, Terra, and Luna, I get variations of the same thing with diminishing quality. It makes the lineup pointless. Ditto for Anthropic. Gemini-3.6-Flash and 3.1 Pro genuinely behave differently. Opus 5 and Fable are.. cousins.
I find that when I want to test a complex creative challenge, having 4 "families" to choose from makes the experience interesting since they will excel in different areas.
Grok might implement unique lighting, Opus, elegant primitives, Sol, accurate snowfall in one pass, Gemini, silky movement. Combined, you can pick and choose best.
For what its worth, Grok always feels "messy" but finishes. Grok 4.6 though is no longer "smart and fast". It's about as fast as Sol though.
A big improvement I noticed in 4.6 was tool use for verification. Previously, Opus/Fable were the only models to consistently screenshot things that they can't directly interact with easily. Now Grok is probably right behind them, perhaps tied with Sol on propensity to verify visually. Grok 4.5 notably did not do this often.
In my experience in heavy coding sessions most pricing is just cache read and cache write like 80% of my token bill.
The rental deal can be terminated by either side with 90 days notice, and presumably Musk would do so if he needed the compute or generally thought it advantageous to do so. For now he doesn't need the compute.
The rental deal may also have been at least in part to juice the SpaceX IPO and to help Anthropic stick it to his enemy OpenAI.
Deals like this that look awkward from the outside but are mutually beneficial to both participants exist everywhere.
Feedback: I'd like Grok to have more connectors (I see OpenAI just added Apple Health, that would be nice to have, and I wish it could read my Onenote notebooks) and for existing ones to be improved. I gave it access to my gmail and asked it "what was my last electricity bill?". It failed to find it, even when I told it the exact subject line to search for. Something about not getting any data back when trying to get the email contents.
Does xAI have plans to do Mac/iOS apps with cloud environments? When can we expect them?
All of these also enable multi vendor LLMs.
Pretty good bang for the buck.
here's Stanford HAI's graph on the carbon emitted from model training per model:
https://spectrum.ieee.org/media-library/chart-showing-estima...
note that Grok's training, thanks to its portable gas generators that are magnitudes less efficient than even other integrated, permanent gas turbines, means the training for this model is dramatically less efficient than models like DeepSeek
a lot of the CO2 emission debate on AI is overblown but it's accurate for Grok
Grok is 3x+ faster than Claude and I can't tell the diff in engineering work quality. As an engineer, speed is important to me.
Maybe because I like to verify its outputs and spend a lot of time iterating to get better outcomes. Presumably if I just let it "do its thing" I'd burn more tokens and "get more done" but I'd lose my grasp on what's in the code base.
That makes me really skeptical of it being GPT5.6-tier, much less Fable-tier, based on some of these benchmarks alone. But I'll test here shortly.
We have Claude, ChatGPT, and Cursor with essentially no cap on spend (top guy is spending over 10K a month on AI at API prices), and he hasn't had his hand slapped.
So it's not like they are using it purely because it's cheaper.
I think people like to use it for its speaking style, pretty solid performance, and its speed.
> Save your context. Always use a subagent (Opus 5 or GPT 5.6) for performing the implementation and then review the work yourself. You are the orchestrator and coordinator it’s up to you to ensure a cohesive final result.
I’ve done more or less the same thing with other agents/versions but with Fable consuming usage credits/tokens so quickly I do it more regularly than usual. I know people who will specify Composer (to my chagrin) as the implementing agent.
Addendum: this really goes a long way, and I can use a single chat session for days before I get into the context danger zone and need to compact/summarize.
But in my experience, the overall productivity ends up similar, give you are willing to work with it in that way.
The grok build TUI harness is excellent and I really enjoyed using it.
For debugging and such I found it pretty much the same as other models.
Fable's taste in software abstraction and project planning in greenfield setups[1] is unmatched in my experience. Sol is OK. My primary use is launching tens of experiments that have to smartly use a limited pool of GPUs.
I use fable to start off the experiments, decide checkpoints, gpu alloc, where to sacrifice precision for performance, and then grok4.5 to iterate, tune, debug, eval, etc, within the abstraction and setup that fable initiated. I have fable write simple scripts that are then wrapped in skills for grok to use. Speed for that loop is extremely important for me, since I also apply human judgement there and I don't like waiting for model output.
I have tried Deepseek and such for the inner agent, but I desperately need multi-modal. Otherwise it's OK, but it tends to use tools less and rambles on and tries to reason with limited information and gets things wrong. Probably a relative la k of tool use posttraining. Gemini flash limits in google ai pro are too low for me to use to compare.
I use anthropic and openais models through grants and so can't compare subscription plan token budgets, but supergrok's budgets are satisfactory.
[1] aside, I have not yet met a model that continues off of a human codebase and actually follows the patterns reliably long term. Eventually it's all slop.
You need to manually push models to clean up the slop every now and then otherwise it becomes chaotic. And every change with LLMs is always extra lines.
I opened a developer API account, loaded 5 dollars and got the free $100s of credits for the month. Like two weeks later, xAI announced they were shutting down the subsidized credits entirely lol. Didn’t even get a full month out of it, and closed my account entirely since I sure wasn’t ever going to put another penny of my own money in.
So my personal lesson was to ignore any hype about the latest “crazy value / unbeatable / free / subsidized X, Y or Z” from anything xAI/Elon adjacent in the future.
These days I get more than enough personal usage from Codex + OpenCode Go to put up with yet another xAI/Cursor offer treadmill, especially if it involves installing new tooling to get it.
Making it open weight does the opposite – it allows traffic to not be routed through them at all.
META releases open weight models because selling access to models isn't their business model, they get optimization for free, good press with adoption of their tech etc. They don't want AI to become another iOS, they want to commoditize the layer underneath them.
GOOGLE releases open weight models because they complement the rest – from premium offering to Android/Chrome/Google Cloud etc products. They use it as part of open ecosystem / local / edge / experimentation layer. 400 million downloads with 100k community variants is nice vibrant ecosystem they have and want to have, they can capture value in several places and would prefer if devs standardize on their tooling.
OPEN AI because their PR was shit and it saved them, purely defensive move, which is interesting considering their name and initial goal.
ALIBABA because it tickles them to erase/cap profits in western labs, free optimisation/research, great PR.
DEEPSEEK they want to have day-zero support for all possible hardware platforms, they started it because Liang wanted to do it, had resources from High-Flyer to do it and he said fuck it and did it (zero commercial incentives which is so bizzare that it deserves a movie or something). He wants to do it because his destination is AGI. In that sense he's more original OpenAI than OpenAI.
MISTRAL because they want to be to AI what RedHat is to Linux.
NVIDIA because they sell GPUs.
...the list is long, there are few dozens of companies including Microsoft, IBM, AI2, Databricks, Snowflake, xAI, EleutherAI, Hugging Face etc. that release open weight models, there are thousands of models and tens of thousands of fine tunes across across text, image, video, audio etc.
In general half life of any model is short. What is more lasting is ecosystem around it, releasing new generation as open doesn't give away technical advantage.
They give away asset for which strategic value is deprecating rapidly and in exchange they get developers, mindshare, tooling, optimizations, integrations, research, standarisation etc. and put pressure on competitors destroying their margins.
It's not an absurd way of thinking, it may seem like irrational move but let's wait and see – imho labs like Anthropic can't sustain long term this kind of pressure and will eventually collapse – regardless of the fact that currently they look like strongest player that can't be touched, the whole thing they have holds on thin, fragile support that gets eroded.
we talked about Chinese models. Yes, erase/cap profits, PR, and give clients peace in mind that model access won't disappear. So, it is infiltraiting markets.
Motivations are more complex than that, if they wanted to achieve that they'd follow approach taken by some western labs to release weaker models only as open keeping frontier behind APIs.
Look at ie. OpenRouter you'll see how many providers there are and how much traffic they get.
If the goal was data, opening weights would be the worst available way to achieve it: a cheap, closed API would capture 100% of traffic (ie Anthropic style), weights can only lose share from there.
With open weights it's net loss of active users of your official api – you're loosing users to self hosting and dozens of providers.
The thing is that open weights create permanent exit that closed models do not have, inference is commoditized immediately, people choose open weight models specifically for data privacy (and stuff like soc2/hipaa compliance) and if somebody wants convenience they go to claude/openai and friends anyway.
Also chinese labs are not uniform with their approach just as western labs are not.
Personally - and I know I'm not alone with this sentiment based on comments I see on this site - I wouldn't touch Grok no matter how good or cheap it is. I don't trust Elon and I don't want to give another dollar to the world's richest person who turns around and uses the money to interfere with elections. The guy I know uses it for essentially the same reason I won't use it.
thats a very dumb reason considering all rich people do it, most are just not as open about it as Musk
It's a really, really easy line to draw in the sand: don't support openly corrupt individuals.
It's absolutely true there's money in politics. To call all money in politics equally corrupt because Bernie got a dollar to have dinner with someone vs Elon effectively directly buying votes....
A complete lack of nuance here. And I wouldn't be surprised if corporations / super rich WANT you to think like that. The more defeatist the mentality becomes the more we just accept whatever they do next.
That's before we even start talking about the models themselves. He claims to want "unbiased" models, but he very clearly has a distorted view of the world and has repeatedly demonstrated a desire and willingness to bend the world to his will. I don't want to use a model that is so obviously suspect. Not to mention, his models repeatedly produce racist, Nazi-like propaganda.
IMO, he is, at best, a clueless amateur masquerading as an expert and running into problems a more careful person manages to mostly avoid. At worst... well, you get the picture.
Elon would be in prison for SEC violations if the current administration hadn’t been elected, and that’s only the tip of the iceberg with that guy.
Elon is directly responsible for Grok becoming self-titled “MechaHitler” which the other three haven’t come close to matching yet.
Calling it "interfering with elections" is utterly bizarre to me.
Me too. The only people I ever saw using grok were using it by accident as they used copilot in auto mode and noticed some prompts were thrown it's way.
I saw far more people using Mistral than grok.
I'm not touching Grok. But there are alternatives to Anthropic and OpenAI that don't refuse to do security work.
Say more about this.
I don't give a shit about Elon's politics in the same way I don't give a shit about Dario or Altman's politics.
Nobody's profitable in this space, they can price it however they want as long as investors keep pouring money in. And SpaceX just got a lot of money poured in.
I used auto in cursor it’s much faster va Claude code and as good.
1. The CapEx play is interesting because it's not just Grok using the hardware. They have rented out hardware for others, including Google, to use. This is making xAI money.
2. It appears that Elon is building a suite of things that work together as part of the push to be multi-planetary. What AI will power the robots? I can understand the drive to have AI they can control to make sure it's appropriate for all the things they are dreaming up. This is a piece they don't want to outsource.
3. OpenAI and Anthropic models are expensive in terms of token costs. Sure, they are frontier. Neither appears to be trying to drive down expenses. This is a problem for heavy users. Companies are trying to put cost controls in place. Does the rest of SpaceX want those cost controls? Having a Frontier model that pushes the pace of driving down costs is really useful.
4. OpenAI and Anthropic are producing models with a progressive lean, according to the Neutrality Project [1]. Having a frontier model that is closer to the middle is considered a good thing by many who are noticing the bias.
These are just some of the reasons. Competition is often a good thing that drives useful change.
As for why anyone else would want it: I've found it's a good model for coding, and it sometimes catches bugs that other models (especially open source models) don't always spot.
https://www.autoblog.com/news/nissan-reports-fifth-straight-...
that alone makes it the closed source subscription i would choose. claude and openai are spying on you.
as it stands i don't have it because the reasoning is encrypted, so i feel that it still is not working for me, it's two faced.
I just assumed every model manufacturer is distilling from the frontier models. If they aren't they are definitely trying to do it.