Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users
openai.com
openai.com
So 5.6 Luna is just their next version of what they used to call 5.5 instant tier
And 5.x instant models were never much to write home about anyway so the default ChatGPT free model hasn’t been particularly distinctive since 4o
This is precisely why I do not offer anything for free. I've been there and done it back in the iPhone 3 era with a dozen apps. Learned my lesson quickly. Never again. It was a nightmare. If the product isn't good enough for someone to pay for it (that could be direct user-paid or advertiser supported) it should not exist.
Also, it is intended to be supported by ads, so it fits your ideal product description quite well despite being a scam.
Oh, please.
Don't use it then. Simple.
Sign of entitlement: You get something of value for free, bitch about limitations and make-up bogus reasons to continue using the product without paying.
If it is of value, use it and pay for it. If it isn't, do not use it and you have nothing to worry about. You are bitching about having access to untold billions of dollars of infrastructure and a massive amount of work from people who have lives and families FOR FREE. Get a grip dude.
I use it for work, extensively, and I gladly pay for the services from a number of AI providers. I would gladly pay double. It's the no-brainer of the century: I can do months of work in weeks or less for a few hundred to a few thousand dollars per month.
I was talking to someone the other day who shares their Netflix account for four other families. Same thing, although, in this case, it's not entitlement, it's theft.
They are both versions of the tragedy of the commons in various ways.
You are missing the point entirely. Let's just take one: LibreOffice. This was a FOSS project from the very start. It is intended to be used as you use it. Even then, bitching about functionality and features would be entitled. In that case, if you don't like it, stop using it or participate in the development...for free, donating your time.
ChatGPT is NOT a free product. Yes, they have a free tier. And that's fantastic. We've all used it. But this is not a FOSS company. Bitching about their free tier isn't right. Again, that's being entitled. Why do they owe you more than they are giving you? They do not. If you want more, pay for it.
Or, you could buy a bunch of powerful computers and run an open weights model on your own hardware and have free intelligence to use as much as you want.
Oh, wait, that wouldn't be free, would it? You'd have to pay thousands for the hardware.
OK, wait, how about you do that and I use your self-hosted LLM's for free and then bitch about the lack of performance.
Yup. Entitlement. Always a reason for other people to owe the entitled and never a reason for the entitled to owe them anything. And this goes way beyond what we are talking about here, far and wide.
> I'm very glad most good software creators I've come across do not have your mindset.
Here's something else you do not understand. FOSS isn't free to create, not even close. Take something like Linux. I would be comfortable estimating that Linux has probably a development cost in the one-hundred billion dollar range. I would not be surprised if was way more than that. The massive number of people who worked on Linux over the decades have lives, families, homes, responsibilities, etc. The only way Linux happened is because they are otherwise employed and donate their time (with is fantastic, I did a lot of of that when I was younger, had time and less responsibilities) or someone is explicitly paying for them to work on Linux because of the benefits they might derive from it. The only "free" part about FOSS is that users of the software don't pay for it. Don't confuse that with all software having to function this way. Like I said, OpenAI is a business. It isn't your personal charity. If you want more, pay for it.
I have not paid a single cent for any of them. Many of them I'm even free to change it however I like and even release my version as a fork.
>Even then, bitching about functionality and features would be entitled.
If you think Linux or ANY of the software I listed would be nearly as good without users "bitching" I have several bridges to sell you. The alternative is to pay a whole bunch of people to test every aspect of the software to find bugs, missing features, minor UI quirks that could be tightened, etc. FOSS only works because of us "entitled" people writing git issues or contributing code when we know how to fix the issue.
You changed my mind. I think OpenAI should pay you to use their technology for free without restrictions. It's only fair.
Live long and prosper.
Relevant. Star Trek is a post-scarcity society where people invent and create solely for the betterment of humanity. You know what else isn't "free" to create? Game mods. Yet I can go out there right now and download several hundred mods for any game that can be modded for free and most times without even a way to possibly pay the creator. I've also spent 100s of hours of my own time creating mods / helping communities on discord or reddit despite there never being a possible financial gain. I genuinely wonder what the software world / internet would be like if everyone were in it only if their time were compensated with money. Horrible place. We wouldn't be having this conversation right here right now that's for sure.
It's also 50% cheaper if you only need it to run sometime in the next 24 hours.
Flex is also easier to get caching to work, there is a little futzing around with OAI's implicit caching but if you do the upfront work you can get haiku quality responses with caching in close to realtime for 1/7th the cost and you don't need to design a polling loop to check for batch completions
I wonder if this means that Luna is more, uhm, "free tier like" in its responses? The "Instant"/"Chat" models have a pretty particular vibe.
This clearly implies that they believe ChatGPT models are AGI and are now willing to say it out loud.
Which I think is a fair interpretation of the term. They are general purpose intelligence in that you can get help from them about almost anything. They are not like narrow single purpose AI models.
I don't think we need that term to mean "can completely emulate a human" or "can do every task any human on earth can do as well as them".
It also needs to be differentiated from ASI with godlike powers many times greater than human.
It is a huge leap. Now we can start talking about "intelligence" at all - we really couldn't before. That we're still hovering barely above the starting line is a separate matter (also worth noting, of course).
They are definitely smart enough to be useful, but dumb enough in their weak spots not to deserve the "general intelligence" qualifier.
The test wasn't made to accurately measure IQs that high.
As an autistic this is painfully obvious to me but: much of human interactions operates on rules that are not only unspoken but often unacknowledged or even outright denied - not just that, but most rules are also highly contextual and rarely treated literally. E.g. corporate guidelines mostly don't exist to be followed (and following them will often result in punishment) but to be able to shift blame - but you need to know for which ones this is the case and for which ones it isn't. This is further complicated because any AI or AI vendor openly making such distinctions would be rejected - AI would not only need to understand all this nuance but also this additional meta layer.
This sounds interesting on it's own. I would be curious to hear more if you are willing to share.
A simple example would be "work to rule": in many professions work processes are heavily regulated (whether by law or by corporate guidelines) but the unspoken assumption is that you know which rules you should ignore and which ones you actually need to follow - but if you tried to find this out by asking "is this a rule I need to follow or not" you would get the clearly incorrect answer that all rules must be followed; of course if you did follow all the rules (aka "work to rule") you would be disciplined for failing to meet quotas (because you can't be disciplined for following the rules).
That's why you shouldn't listen to your pop celebrities for political advice.
Just look at some training sets to see how the sausage is made: https://huggingface.co/datasets/nickrosh/Evol-Instruct-Code-...
In the original test the evaluator knew one was a machine and one a human and could have conversations of arbitrary length.
We have one of those: Grok.
Interesting idea. "This is not dumb and biased enough, probably not a human".
LLMs still live in the uncanny valley and can be sussed out immediately.
For example, I’m extremely annoyed by the fact that offshore developers respond to me almost exclusively using text generated by Claude. You can tell immediately because they use overly descriptive techno word salad that no normal human uses unless they are trying to be ultra specific for a scientific paper - and even then it’s still too much for a real person.
5 years ago, you would not have been able to determine that this was the case, and just have assumed it's a know it all character.
Tell your offshore developers to use caveman or so
Well, the models are smart enough to point out why this is wrong.
Wasn't there a thread the other day about how very few humans on earth can understand the new math proofs?
Although "with sufficient study" vs "not even with unlimited study" are probably worth distinguishing there.
Do you perhaps have a link to it? Tried to search around but couldn’t find it. Would love to read it
https://news.ycombinator.com/item?id=49165570
> When OpenAI posted about their 10 breakthroughs, I saw lots of career research mathematicians say things mostly along the lines of “I don’t understand any of this it’s way over my head”.
https://news.ycombinator.com/item?id=49182779
> They're publishing machine checkable proofs because that's the only way they, themselves, can check them.
> They don't understand the math either.
>artificial general intelligence (AGI)—by which we mean highly autonomous systems that outperform humans at most economically valuable work
Take mathematical calculations, styrofoam, LCD screens, ice cubes, embroidery (try searching for "computer work blouses"!), texts, navigation. The thing itself costs nothing, the service and theater around it becomes everything.
OAI is not going to announce anything until their next model officially releases (allegedly later this month).
Don't you see a problem here? Terms are used to describe the world and need a semblance of stability so we don't end up in a race to the bottom just so investors can feel good.
Do you genuinely hold this position, or do you not realize how far the goalposts have shifted?
In 2022, prominent AI critic Gary Marcus offered to bet $100,000 that we wouldn't have AGI by 2029. https://garymarcus.substack.com/p/dear-elon-musk-here-are-fi... Because the definition of AGI is unclear, he defined that AGI would be achieved if an AI model could do THREE of the five following tasks:
- In 2029, AI will not be able to watch a movie and tell you accurately what is going on (what I called the comprehension challenge in The New Yorker, in 2014). Who are the characters? What are their conflicts and motivations? etc.
- In 2029, AI will not be able to read a novel and reliably answer questions about plot, character, conflicts, motivations, etc. Key will be going beyond the literal text, as Davis and I explain in Rebooting AI.
- In 2029, AI will not be able to work as a competent cook in an arbitrary kitchen (extending Steve Wozniak’s cup of coffee benchmark).
- In 2029, AI will not be able to reliably construct bug-free code of more than 10,000 lines from natural language specification or by interactions with a non-expert user. [Gluing together code from existing libraries doesn’t count.]
- In 2029, AI will not be able to take arbitrary proofs from the mathematical literature written in natural language and convert them into a symbolic form suitable for symbolic verification.
Today's AI models can do FOUR of these five.
Using 2022 goalposts, we already have AGI. We blew past these goalposts months ago, and nobody noticed.
In the meantime, it’s still very easy to differentiate between an AI and a human in a chat. You just need to know the quirks of these systems. Like counting letters, hitting the safeguards, etc.
So call me when one of them can pass the Turing test against me and then we can talk about AGI
On top of that for your strong feelings, you don't have the conviction to write down a strong definition of intelligence yourself, which allows you to accelerate the goal posts up to light speed. The fun thing about writing out a formal definition is suddenly almost everything or almost nothing, including a lot of humans, has intelligence.
Not basing intelligence on your feelings of the moment makes it a hard thing to define across everything intelligence applies to.
Ring ring I'm calling you right now. We blew past the Turing test goalpost over a year ago, using 2024 models.
https://www.ie.edu/uncover-ie/has-ai-passed-the-turing-test-...
GPT-4.5 passed the Turing Test with a 73% human rating, outscoring actual human subjects. That is, human evaluators considered the AI more human than an actual human, 73% of the time. LLaMa-3.1 was judged to be a human 56% of the time.
The 'strawberry' test was fixed years ago with the invention of CoT; models only fail that test today when thinking is disabled.
You can't be serious... Case in point - I just asked GPT Sol High to give me a weekly update of local ai changes.
Here's it's first update, which is complete and utter garbage; i.e. it's a lot of words that says absolutely nothing. That's just a random word generator; AGI? Not even remotely in the ballpark.
https://chatgpt.com/share/6a761ec1-4b38-83ea-a8f7-8f02084e63...
I assume research-level questions about scientific topics (with a focus on math and computer science) are not easy to monetize (yes, this is by far the most common kind of question to AI models for me). :-D
I can imagine some hypothetical scenarios, but I do believe that I have strong evidence that this is mostly not the case.
For example, I think I have already told the story how when I worked together with a person who is a business consultant (and thus salesman) told me that C-level executives are an incredibly easy sales target compared to me. For him, selling something to me was "ultra-hard mode", basically because I could immediately see through every sales tactic that he tried on me.
As I wrote in some parallel post:
"When I look at my history with AI chatbots, 90 % of the conversations revolve around research-level problems in math and computer science. :-D"
For quite a few of them, creating a decent answer involved quite a bit of thinking by the AI, so I assume they are not the easiest queries to answer, and the queries are too obscure to cache.
The median LLM query isn't significantly costlier than web search.
Whatever you think about Altman and Amodei, they aren't clueless idiots and if ads was all they needed they would have done that from the beginning.
No they wouldn't have. Low cost isn't no cost, R$D is expensive and there are a class of tokenmaxxing users well outside the median (e.g Agentic Coding).
> I don't know where you get that from
https://cloud.google.com/blog/products/infrastructure/measur...
https://epoch.ai/gradient-updates/how-much-energy-does-chatg...
I'm not talking about R&D.
> there are a class of tokenmaxxing users well outside the median (e.g Agentic Coding).
It wasn't a thing before last year and OpenAI wasn't making profit before that either.
> https://cloud.google.com/blog/products/infrastructure/measur...
> https://epoch.ai/gradient-updates/how-much-energy-does-chatg...
None of these support your claim.
Again, when you're here saying that the CEOs of the two biggest AI companies are basically idiots, nobody will take you seriously.
You said OpenAI and Anthropic would be profitable if the median query was as cheap as I say. I'm telling they wouldn't be because inference isn't the only cost these companies shoulder. R&D/Training is a very big part of costs. That you didn't mention it is irrelevant. R&D is the lion's share of costs - https://news.ycombinator.com/item?id=48550465
>None of these support your claim.
Yeah they do lol. Both those articles place the median query cost around the same as a Google search.
>Again, when you're here saying that the CEOs of the two biggest AI companies are basically idiots, nobody will take you seriously.
That's not what I'm saying and I don't see what's so hard to understand here.
> Research and Development: $19.18 billion. Loss from Operations: $20.92 billion
OpenAI isn't profitable even if you discount R&D entirely. Google search on the other hand has had ridiculous margin from the beginning, allowing the company to be highly profitable while financing significant R&D investment. These two businesses models are as different as it can be.
If ads were the solutions for profitability these companies would have used ads as their source of income from the very beginning.
They have a billion weekly active users, almost all free (Not Google search free. Free free) with about ~50M subscriptions. Of course they're not profitable even without R&D. Low cost isn't no cost. Take away the huge R&D and they become hugely profitable with a robust ad business like Google Search.
>If ads were the solutions for profitability these companies would have used ads as their source of income from the very beginning.
You've said this a couple times and it doesn't make any more sense the more you say it.
Just admit you're wrong about inference costs and move on.
> Take away the huge R&D and they become hugely profitable with a robust ad business like Google Search.
Is refuted by the sentence that comes immediately before :
> They have a billion weekly active users, almost all free (Not Google search free. Free free) with about ~50M subscriptions. Of course they're not profitable even without R&D. Low cost isn't no cost.
Please explain why have OpenAI been actively not trying to “hugely profitable with a robust ads business” for the past 3 and a half year then? Or at least make a profit that can cover some of there R&D expenses, instead of losing money from their operations?
> Just admit you're wrong about inference costs and move on.
Unfortunately I'm not. And again, none of your links above says otherwise.
No. "They're unprofitable while serving almost everyone for free" does not refute "they could be profitable if those users were monetized with ads." That's the entire distinction.
>why have OpenAI been actively not trying to “hugely profitable with a robust ads business” for the past 3 and a half year then?
Because "they didn't do it early enough" is not evidence. They may have prioritized growth, product adoption, subscriptions, or simply delayed ads for strategic reasons.
>Unfortunately I'm not. And again, none of your links above says otherwise.
They show that "LLM inference which is significantly costlier than web search." is unsupported. Imagine dismissing numbers over "but why no ads earlier".
Prioritizing growth? When OpenAI had saturated the market by year one and has only been losing market share for the past 18 months! Again that's basically assuming Sam Altman is a complex idiot. No board in existence will let you move away from the most obvious business model in the Silicon Valley if this model was the obvious solution you think it is.
> They show that "LLM inference which is significantly costlier than web search." is unsupported
It's not unsupported, the price per million token of frontier model is public, as well as the cost of running open source models, which gives us a rough idea of how much those frontier models cost to run, and it's significantly higher than what it costs to run a web search.
> Imagine dismissing numbers
But you're just making those “numbers” up: there's simply no cost numbers at all in these papers. They are just about energy, and you jump from there forgetting that consuming a kWh of electricity on a server from 19 years ago is substantially cheaper in capex than doing so on a H100, let alone doing so on newer hardware.
To expand: it seems inevitable that Google's SERP format will be replaced with a conversational / chatbot / agentic interface, which equalizes the playing field for all chatbot providers.
This is because you can stuff in much fewer ads into a chat interface compared to SERPs. (They could try stuffing more ads but that would likely just push users more to the competition who have a much lower baseline on which to show growth.) As such, Google would be forced to progressively nullify its own invincible firehose of ad revenue as they deprecate SERPs in favor of AI overviews.
There's no way OAI has a long term advantage over google in replacing the search engine experience.
Google's one major weakness is it will face the innovator's dilemma as their core search revenue gets cannibalized. But they seem to have been able to get their entire org to recognize that AI is an existential threat so at least that's a good sign.
A counter point would be OAI and Anthropic can pay with more equity that they can promise will go to the moon. But all the equity base compensation eventually dilutes earnings per share so it's not free once you go public and people start caring about that.
Then notice how many of the advantages you enumerated are directly related to their ad business. But the ad business itself is under threat. I just cannot see how they can stuff as many ads in an agentic interface as their SERPs. E.g. what's the net outcome of shoving AI everywhere if it's not going to be monetized nearly as well as their ads?
My point is, Google has finetuned their ad business and surrounding ecosystem to an extreme level to sustain this absolute firehose of cash (including antitrust and rig-bidding shenanigans they were literally found guilty of) but almost all of that is disrupted by the shift to chat interfaces.
I think Google will do extremely well as an AI chatbot / agent and cloud AI provider, but it will not be nearly as lucrative as the ad business they have to cannibalize to get there.
You’re painting a dichotomy that doesn’t exist and hasn’t for a good while.
Edit: I missed your final paragraph. I still disagree with this argument though, people still want to go to websites.
https://www.niemanlab.org/2026/07/search-traffic-has-decline...
https://www.techspot.com/news/113199-reddit-major-publishers...
Email is a good window but lots of people talk in more depth with a chat bot.
Search is a good “purchase intent moment” but people are telling a chat bot exactly what they care about in a purchase, rather than making indirect search terms and reading review sites.
> Search is a good “purchase intent moment” but people are telling a chat bot exactly what they care about in a purchase, rather than making indirect search terms and reading review sites.
When I look at my history with AI chatbots, 90 % of the conversations revolve around research-level problems in math and computer science. :-D
But 90% (99%?) of user sessions don’t require a huge amount of reasoning tokens or output; keep in mind LLMs are replacing search for single-turn answers.
Google is going to be able to push cost down further and for longer than anyone else. They also rightfully believe this is a company ending gamble if they fail they could be a second tier player for decades.
So better inference margins, far better operation experience serving cheap AI at massive scale, a massive warchest, and the fear of complete company irrelevence.
That's going to be one hell of a company to beat for free tier LLMs. Also chatgpt is impressive but it's notthing like google search quite yet. On top of that AI labs must use google's product for their AI.
They also own more data by an exponential margin. Anything but dominance of free AI on google's side would mean they are just so incompetent they deserve to fail.
There's no reason why Google's public stuff is this stale, overpriced and underwhelming. But at least until their next round of models drops, even calling them a "frontier lab" is starting to feel like a stretch. Which is weird!
Especially considering that in the 2010s they were The Big AI company, especially after buying out boutique shops like DeepMind.
Because everyone now outsources much of their thinking and researching to LLM's, our collective culture + brain is shaped in a cyclical manner by using them.
It's the mechanical homogenization of culture and groupthink.
TBH I don’t find it useful at all for personal use. It’s totally soulless for creative ventures and absolute dogshit at researching the things I want it to be good at (I.e. planning a vacation or finding new music).
All that combined with the social stigma makes me feel pretty skeptical that it’s some kind of pop culture shaping mechanism, at least not for a few more years.
They took away the button a few months ago and are now putting it back.
Perhaps I should have said "proper access".
That would explain a lot of the terrible AI/LLM takes online.
This could be the Opus 4.5 moment for normal humans.
The Aug 6 update has forced the entry box to auto-format Markdown in an attempt to imitate Claude. The implementation is buggy and even simple copy-and-paste has gone entirely haywire. They also forgot to leave a switch to turn the confounded autoformatting thing off.
Chat mode in general is currently crawling with more UX bugs than a porch screen in summer.
It made me wonder how many paid subscribers realize they are using the same 5.5 instant model as free users by default. A dark pattern or oversight?
I’ve literally seen people ask ChatGPT for a link to Gmail.
Is it a dark pattern, or design decision that makes users happier?
I expect a few things to happen in the next year:
1) Exclusive MCP server deals/API integrations
2) Significant switch to B2B marketing, even moreso than we've seen before, with API interfaces being paid and chat-client interfaces becoming more and more free, perhaps just with limits more on integrations or data visualization/analysis
3) US restrictions on B2B contracts with non-US hosted models that do any sort of contracting with the government
Obviously there's a bunch of stuff I'm not foreseeing. But it really does feel like the bottom of the market is collapsing into free. I assume OpenAI and Anthropic think their next generation of models will restore their halo tier status and that the cash burn is justified to just get there, but this has to really mess up IPO plans.
1) Back then, even as a free user you'd be able to use the strongest model (even if with tight limits). Now, you need to pay to use Sol, and you need to pay to use Opus or Fable. It does seem fairly premium in that sense. Idk about 5.6 Luna, but the previous Instant was really bad, even for very casual users. It would hallucinate non stop.
2) When $100 and $200 per month plans launched, they were received as outrageous even here. Nowadays they are pretty common among power users.
This is kind of what I'm saying though. Bottom has fallen out, differentiation is just can you be much more premium than the competition. Currently that remains unanswered.
EDIT: I'm basing this off the assumption that for chat, premium is not a point of differentiation at all. For coding/analysis, it is.
Back when coding for me still meant copy-paste from the web version, it was only worth the $20/month for me.
They only added the $100 Pro plan in April during GPT 5.4 times.
Today I happily pay $400/month for Codex and Claude Code.
So same name, and no indication of which "version" you're using except whether you're in "Chat" or "Work"? Why make it so confusing?
Maybe Luna efficiency gain was actually significant enough that putting all the free users and giving them super generous limits makes sense.
They might be doing this to improve the messaging of AI among causal users since right now there is a huge amount of datacenter backlash in the US due to AI grievances.
Maybe they have too much excess capacity or they really want to juice token numbers and market share on their dashboards for marketing.
I also wonder if being given access to an actually a decent model like luna with actual thinking budget instead of brainless "instant" modes will start to make causal users understand the real capabilities of these models.
All of this is of pretty minor importance though. You can't read as many tokens as a subcription can produce so more chat is not the value add nor super important.
I mean there are literally so many providers for free chat if you are willing to use several seperate apps.
The real value in these subs is using codex cli, much like the real point of anthropic subs is using claude code. Because agentic work actually does require a lot of tokens.
Definitely this. The recent 80% discount was a reaction to Deepseek's update so that they still position near the frontier. My theory: Luna has always had a much higher efficiency. You do know that the model didn't get faster after the discount?
luna is very good
Both models are cheap enough that I can run 4 sessions at the same time without running out of the 20 USD codex and 10 USD Opencode plan. I've burned through almost a billion tokens this week and I've done some pretty big refactors as well.
I have a Claude Max subscription but I've barely been using it because of the many issues they've had this week.
Pricing absolutely matters.
I buy 2 hot dogs for $1 at Sheetz sometimes. For that price, I get them with mustard, onions, and sauerkraut already on them.
They are not excellent in any way. I will probably never love them. :)
But they're available 24/7/365 and they don't take long for the staff to throw together.
And most importantly: They sure are cheap.
(Costco's $1.50 dog+Coke is much higher quality and presents a better value, but it requires visiting a Costco and that has its own cost.)
(This is a subtle nudge at anyone from OpenAI who reads this to make sure they get updated.)
OpenAI have a model called "chat-latest" - I wonder if that's running this new model yet: https://developers.openai.com/api/docs/models/chat-latest
It's described as "points to the latest Instant model currently used in ChatGPT" - so presumably that's "GPT-5.6 Instant" in the app.
https://gist.github.com/simonw/aae4febd3c6f7bc5b7811857edb3c... has screenshots that still show "Instant" as an option for ChatGPT Chat... but not for ChatGPT Work.
It’s hard to understand this .. like sure we can select instant but is there an actual model called 5.6 instant? Like is it on LM Arena and OpenRouter or available via API etc
5.5 instant is definitely A Thing it’s even name checked in this OAI post
Every week, 1 billion people turn to ChatGPT for everything from quick questions and web searches to planning, research, advice, and complex decisions.
Guess, Google's AI Mode is chipping away at their consumers (I know I haven't used Chat in a long, long while for 'quick questions and web searches' after OpenAI did away with "think" which I always use). The money-minting office & coding market Anthropic has cornered is hyper-competitive at both the frontier & low-cost ends. OpenAI is reactive [0] and seems right up against it, despite the strength of its excellent models.[0] Won't put it past OpenAI (and/or Google) to open weight larger models!
My mission is world conquest. I'm writing a comment on an HN thread.
No, those two clauses have no relation whatsoever. I just felt like saying the first sentence because it sounded cool.
It seems like giving it a time limit or a budget in dollars would be clearer, though?
Or, keep searching until I come back to the computer and ask about progress.
very nontrivial problem knowing when to stop and making assumptions
ChatGPT already competes with Google...
Sadly it's not available anymore on the updated desktop app (formerly Codex) and the previous desktop app (formerly ChatGPT) is abandoned.
I feel that work is basically split into two tiers, hard (which requires a good model and lots of reasoning by definition) and relatively easy (which won't consume much of my limits despite a great model and reasoning, so I may just as well keep it on Sol High).
Luna: Can be useful if price sensitive, always use with AT LEAST high effort. But codex plans are very generous, so just ignore it.
Terra: Forget it exists
Sol: Just use this. Medium can work for specific edits. Otherwise just use high or xhigh.
OpenAI basically agrees with this, and the slider gives you those options.
TL;DR: Just use sol medium/high/xhigh
Ultra: Might be useful, the one time I tried it, it was extremely wasteful in terms of tokens. Definitely not a daily driver.
There's also max effort, I think it can be useful but the gains compared to xhigh are quite small.
I think max and ultra are only worth it when the others fail.
I don’t know if this is a good workflow.
This equation significantly changes if you're paying API prices vs on a subscription. (Such as if you're integrating it into a different product, where it becomes worth it to figure out what the cheapest you can leverage is.)
I hope to see a memory-less mode that is not incognito. Want fresh contexts sometimes, but also want to keep the chats saved in history. Memory can spoil some creative work, it dials the model in too tightly.
The value is shifting up the stack... make core intelligence free, then monetize the ecosystem built on top of it .. kinda similar to how the internet itself is free, but platforms and apps capture the value.
I think we'll see a huge push towards connectors for work and personal tools, along with much deeper os level integration. That's where the long-term moat is, not the base model itself.
Improving the free offerings is meant to increase visibility and therefore market share. It's not about goodwill, and it never will be. :)
The free stuff is primarily marketing and marketing always has costs.
In terms of compute, I have no way to really look behind the curtain and see what goes on back there. But I know with codex CLI, in terms of weekly quota: I can get a ton of work done with luna and usually get reasonable results. It feels very compute-light in this way.
I'm amazed by the work luna on xhigh can do for the price I pay (just $20, every month). It has the presentation of something that is very efficient to run, while also being something that can actually produce OK results. It's also fairly quick.
It differs from many previous smaller offerings of yore in this way. Like, I mean: I found stuff like the -mini models and 4o to be utterly useless wastes of my time. Luna isn't like that at all; it can get some stuff done.
So far for me, luna is the most impressive part of the 5.6 rollout. Not because it is best, but because it is useful and cheap.
So if luna is decent (it seems that it is), and if it is in fact light (which seems to be true from what I can observe), and it is offered for free, then it may very well be better, faster, and cheaper than the competition is.
And that's good for visibility. Marketing is all about buying eyeballs.
Hm, does this mean that 5.6 Pro in ChatGPT web is somehow different/not as good now? I found it really good for code review (upload your repo and patch and off it goes).
But seeing the graphic with the visual weather report: that makes me think that is not the goal at all. :)
Even after identifying the 0% chance of rain, it still drags the conversation on and on and on
I wonder if they actually do it to optimize inference. I maintain a corporate AI server and one of the tricks to reduce the load was to modify the system prompt to be as terse as possible so the average response completes faster and requests queue up less often.
Less verbose output = less context & less token gen.
Serving cheaply at that scale means routing, batching, and cache hits, not a better model :D
I was not impressed with 5.6 and this hits exactly why.
Also, this instant, medium, and high slider situation we now have everywhere is batshit crazy. It’s a major step back in technology.
I’ll put money on the table there will be a surprise in revenue because users have no clue what to choose, so they are constantly choosing high because they don’t want to risk getting inaccurate answers. If this was intentional by the dark patterns department, then brilliant. However, I’m guessing Anthropic and OpenAI are struggling to know how to deploy their models.
Also, the models were already amazing. They need to slow down and do a model release once a year and only do extremely minor iterations instead. Some amazing things are accomplished, but how people actually want to utilize the models gets screwed up every time in the process.
If you're going to force people to specify manually, at least make it 0.0 - 1.0 normalised such that 0.5 is the default.
There are a ton of use-cases where the models still struggle a lot and make bad decisions, I'd much prefer they continue their acceleration for a while more.
Yes, please give free users more access and leave paying codex users in the dust.
Wise plan, Sam.
I've paid for ChatGPT (and in the recent ~year, codex) for about as long as it was available to pay for. I'm not upset by the offer of giving luna to the masses for free -- not at all.
Should I be upset that people can cut-and-paste to the lessest of the new model variations for free? If so, then why?
Nothing was taken from me here. I'm still going to keep doing whatever it is that I do with the tools that I've been using.
I'm not in competition with anyone, and even if I were then it wouldn't be with those who are using ChatGPT for free on the web.
I dunno, I've honestly never been happier as a Codex customer. Sure, there's been less bonus usage resets recently, but the amount of productivity and value I've gotten out of my $20/month personal subscription is bonkers.
If you're price insensitive, Fable 5 is definitely still the best, but a lot of people aren't.