GPT-4.5: "Not a frontier model"?
interconnects.ai
interconnects.ai
https://www.reddit.com/r/singularity/comments/1izpb8t/gpt45_...
I'm guessing that this model was finished pre-training at least a year ago (it's been 2 years since GPT 4.0 was released) and they just didn't see the hoped-for performance gains to think it warranted releasing at the time, and so put all their effort into the Q-star/strawberry = eventual O1 reasoning effort instead.
It seems that OpenAI's reasoning model lead isn't perhaps what they thought it was, and the recent slew of strong non-reasoning models (Gemini 2.0 Flash, Grok 3, Sonnet 3.7) made them feel the need to release something themselves for appearances sake, so they dusted off this model, perhaps did a bit of post-training on it for EQ, and here we are.
The price is a bit of a mystery - perhaps just a reflection of an older model without all the latest efficiency tricks to make it cheaper. Maybe it's dense rather than MoE - who knows.
And the successor model was called "Orion", not "Omni".
And the successor model was called "Orion", not "Omni".
I think it at least is somewhat analogous to what happened with pricing on previous models. GPT 4, despite being less capable than 4o, is an order of magnitude more expensive, and comparably expensive to o1. It seems like once the model is out, the price is the price, and the performance gains emerge but they emerge attached to new minified variations of previous models.
One theory is that they're worried about the increasing tide of LLM-generated slop that's been posted online since that date. I don't know if I buy that or not - other model providers (such as Anthropic, Gemini) don't seem worried about that.
The general public will naturally expect it to be the next big thing. Wasn’t that the point of releasing it? To seem like progress is being made? To try to make that point with a model that doesn’t deliver is a misstep.
If I were Sam Altman, I’d be pulling this back before it goes on general release, saying something like it was experimental and after user feedback the costs weren’t worth it and they’re working on something else as a replacement. Then o3 or whatever they actually are working on instead can be the “replacement” even if it’s much later.
Sonnet 3.7 is actually reasoning model.
I might be wrong but I couldn't find a source that indicates that the "base" model also implements reasoning.
It seems, based again on limited experimentation doing sort-of-real work, that analysis works quite well and extended thinking is so-so. Whereas DeepSeek R1 seems to be willing and perhaps even encouraged to second-guess itself (maybe this is a superpower of the “wait” token”), Sonnet 3.7 doesn’t seem to second-guess itself as much as it should. It will happily extended-think, generate a wrong answer, and then give a better answer after being asked a question that it really should have thought of itself.
(I’m not complaining. I’ve been a happy user of 3.7 for a whole day! But I think there’s plenty of room for improvement.)
I've heard the theory that's intentional, because after roughly that point the internet training data became full of AI slop
The leap from GPT-4o to 4.5 isn't a leap—it's an expensive tiptoe toward incremental improvements, priced like a luxury item without the luxury payoff.
With pricing at 15x GPT-4o, they're practically daring us not to use it. Given this, I wouldn't be surprised if GPT-4.5 quietly disappears from the API once OpenAI finishes squeezing insights (and cash) out of this experiment.
I don't at all blame OpenAI for going down this path (indeed, I laud them for making expensive bets), but I do blame all the quote-un-quote "thought leaders" who were writing breathless posts about how AGI was just around the corner because things would just scale linearly forever. It was classic "based on historical data, this 10 year old will be 20 feet tall by the time he's 30" thinking, and lots of people called them out on this, and they either just ignored it or responded with "oh, simple not-in-the-know peons" dismissiveness.
To your point, though, I would add not only who has seen any grand conspiracy actually be accomplished, who has seen one even attempted and kept under wraps? Such that the absence of corroborating sources was more consistent with an effectively executed conspiracy theory than the simple absence of such a plan.
Of course, that's my point. Again, I think it's great that OpenAI swung for the fences. My beef is again with these "thought leaders" who would write this blather about AGI being just around the corner in the most uncritical manner possible (e.g. https://news.ycombinator.com/item?id=40576324). These folks tended to be in one of two buckets:
1. "AGI cultists" as I called them, the "we're entering a new phase of human evolution"-type people.
2. People who had a motive to try and sell something.
And it's not about one side or the other being "right" or "wrong" after the fact, it's that so much of this just sounded like magical thinking and unwarranted extrapolations from the get go. The actual experts in the area, if they were free to be honest, were much, much more cautious in their pronouncements.
Either way, AGI or not, LLMs are pretty magical.
And yeah, LLMs are awesome. But you can't predict scientific discovery, and all future AI capabilities are literally still a research project.
I've had this on my HN user page since 2017, and it's just as true as ever: In the real world, exponentials are actually early stage sigmoids, or even gaussians.
I ask chatgpt to give me a map highlighting all spanish speaking countries, gives me stable diffusion trash.
Just gotta do the grunt work, add a tool with a map api. Integrate with google maps for transit stuff.
It's a good LLM model already it doesn't need to be einstein and solve aerospatial equations. We just need to wait until they realize their limits and find the humility to build yet another useful product that won't conquer the world.
Works great when the output is small enough to unit test or immediately try in situations with no possible negative outcomes.
Anything larger? Skip the LLM slop and go to the source. You have to go to the source, anyway.
Yeah, and then check that. I don't get this argument at all.
People who uncritically swallow the first answer or two they get from Google have a name... but that would just derail the thread into politics.
It turns out that going from not knowing what you don't know to knowing what you don't know adds an order of magnitude improvement to people's experience.
Are you referring to the search tool? Like when you ask the LLM something beyond its cutoff date it searches and gives you what it searched?
While it's cool that that feature shows the sources, the core of the LLM still does not provide sources, again by design it forgets what the sources are, it cannot link to the common crawl datapoints.
4o can link to sources to some extent, but it sounds like it doesn't do "research-y" things at all unless you have a paid account. It works for me -- note how it cites the manufacturer's literature here: https://i.imgur.com/lRkH948.png -- but when I ask it the same thing in a private window it doesn't cite any sources: https://i.imgur.com/Kc6Y0IR.png .
So that does seem kind of lame. They will offer source citations in the free model the minute they realize they can deliver ads that way, I'm sure...
For example, have you ever personally verified that humans went to the moon? Have you ever done the experiments to prove the Earth is round?
I have, actually! Thanks, astronomy class!
I've even estimated the earth's diameter, and I was only like 30% off (iirc). Pretty good for the simplistic method and rough measurements we used.
Sometimes authorities are actually authoritative, though, particularly for technical, factual material. If I'm reading a published release date for a video game, directly from the publisher -- what is there to contest? Meanwhile, ask an LLM and you may have... mixed results, even if the date is within its knowledge cutoff.
Was it the stick on the earth and measuring the right triangle casted by the shadow?
IIRC you also had to do the same thing on another spot far apart and measure the difference between both triangles. At the same time.
I specifically remember that we measured how far the shadow moved over time using chalk. I think this was in lieu of having a second triangle somewhere else, although, for all I know we might have used a second reference distance.
Our campus had this elaborate outdoor "observatory" that had all sorts of interesting features. The gnomon was one of them, but there were other cool things, like IIRC there was a metal sculpture where Polaris would be seen through a central hole (and a guide inscribed in the concrete to figure out where to stand for that to happen -- I think it was based on viewing height).
For example, if I'm looking for some medical finding and I get to a source that's a clinical study from a reputable publication, I may be satisfied and stop there since this is not my area of expertise. However, a person with knowledge of the field may be able to parse the study and pick it apart better than I could. Hence, their search would not end there since they would be unsatisfied with just the source I was satisfied with.
On the other hand, having no verifiable sources should leave everyone unsatisfied.
I'm getting this. https://claude.ai/share/7a8ecdb0-a28c-4d48-ad81-2d9e95fab538
However they’ll need to replace vast swathes of the economy to justify these AI companies’ market caps.
This is kind of the crux though. The only way to make LLMs more useful is to basically make them traditional AI. So it's not really a leap forward nevermind path to AGI.
As for API price, when it matters businesses and people are willing to pay much more for just a bit better results. OpenAI doesn't take the other options away. So we don't lose anything.
Disclaimer: have not tried 4.5 yet, just skimmed through the announcement, using 4o regularly.
More info from ChatGPT here: https://chatgpt.com/share/67c5ad36-2f84-8001-a61f-11d4e17135...
And unfortunately one not exclusive to OpenAI. Anthropic credits also expire after 1 year.
Very quick mvp comparison for the show me what you mean crew: https://chatgpt.com/share/67c48fcc-db24-800f-865b-c0485efd7f... & https://chatgpt.com/share/67c48fe2-0830-800f-a370-7a18586e8b... (~30 seconds vs ~3 minutes)
> Mission is the operationalized version of vision; it translates aspiration into clear, achievable action.
The "Mission is the operationalized version of vision" is not in the corpus that I am find and is obviously a confabulated mixture of classic Taylorist like "strategic planning"
SOPs and metrics, which will be tied to compensation and the unfortunate ubiquitous nature of Taylorism would not result in shared purpose, but a bunch of Gantt charts past the planning horizon.
IMHO I would consider "complex nuanced thought" as understanding the historical issues and at least respect the divide between classical and neo-classical org theory. Or at least avoid pollution of more modern theories with classical baggage that is a significant barrier to delivering value.
Mission statements need to share strategic intent in an actionable way, strategy is not operationalization.
The quality of writing can be much better than Claude 3.5/3.7 at times but struggling with similar confabulation of information that is not in the original text but "sounds good/flows well". Which isn't ideal for a personal journal... I am still playing around with the system prompt but given the astronomical cost (even with me as the only user) with marginal benefits I am probably going to end up sticking with Claude for now.
Unless others have a recommendation for a less robot-y sounding model (that will, however, follow instructions precisely) with API access other than the mainstream Claude/OpenAI/Gemini models?
(also: the person you are responding to is doing exactly what you're saying you don't want done, take something unrelated to the original text (Taylorism) but could sound good, and jam it in)
Concise, clear, and memorable statement that outlines a company's core purpose, values, and target audience.
> "made operational/actionable at a strategic level".
Taken the common definition from the first part of this plan, what do you think the average manager would do given that in the social sciences, operationalization is explicitly about measuring abstract qualities. [1]
"operationalization" is a compromise, trying to quantify qualitative properties, it is not typically subject to methods like MECE principal, because there are too many unknown unknowns.
You are correct that "operationalization" and "strategic intent" are not mutually exclusive in all aspects, but they are for mission statements that need to be durable across changes that no CEO can envision.
The "made operational/actionable at a strategic level" is the exact claim of pseudo scientific management theory (Greater Taylorism) that Japan directly targeted to destroy the US manufacturing sector. You can look at the former CEO of Komatsu if you want direct evidence.
GM:s failure to learn form Toyota at NUMII (sp?) is another.
The planning process needs to be informed by stratagy, but planning is not strategic, it has a limited horizon.
But you are correct that it is more nuanced and neither Taylor nor Tolstoy allowed for that.
Neo-classical org theory is when bounded rationality was first acknowledged, although the Prussian military figured that out long before Taylor grabbed his stopwatch to time people loading pig iron into train cars.
I encourage you to read:
Strategy: A History By sir Lawrence Freedman
For a more in depth discussion.
[1] https://socialsci.libretexts.org/Bookshelves/Sociology/Intro...
He clearly communicated in 1977 that his ideas were never formally validated and that he cautioned about their use in other contexts.
I think that the concepts can be useful, if you took them as anything more than a guiding framework that may or may not be appropriate for a particular need.
https://core.ac.uk/download/pdf/36725856.pdf
I personally find value in team and org mission statements, especially for building a shared purpose, but to be honest, any of the studies on that are more about manager satisfaction then anything else.
There is far more data on the failure of strategy execution, and linking strategy with purpose as well as providing runways and goals is one place I find vision and mission statements useful.
As up to 90% of companies fail on strategy execution, and because employee engagement is in free fall, the fact that companies are still in business means little.
Context is king, and this is horses for courses, but I would caution against ignoring more recent, Nobel winning theories like Holmström's theorem.
Most teams don't experience the literal steps Tuckman suggested, rarely all at once, and never as one time singular events. As the above link demonstrated, some portions like the storming can be problematic.
Make them operationalize their mission statement, and they will and it will be in concrete.
Remember von MoltKe "No plan of operations extends with certainty beyond the first encounter with the enemy's main strength."
There is a balance between C2 and mission command styles, the risk is trying to force or worse intentionally causing people to resort to c2 when almost always you need a shifting balance between command and intent based solutions.
The Feudal Mode of Production was sufficient for centuries, but far from optimal.
The NUMMI reference was exactly related to the same reason Amazon profits historically raised higher despite head count increases that should have allowed.
Small cross functional teams, with clearly communicated tasks, and enough freedom to accomplish those tasks efficiently.
You can look at Trist's study about the challenges with incentivizing teams to game the system. Same problem happened under Balmer at MS, and DEC failed the opposite way, trying to do everything at once and please everyone.
https://www.uv.es/=gonzalev/PSI%20ORG%2006-07/ARTICULOS%20RR...
The reality is that the popularity of frameworks rarely relates to their effectiveness, building teams is hard, making teams work as teams across teams is even harder.
Tuckerman may be useful in that...but this claim is wrong:
> "Modern mission frameworks prioritize adaptability within durable purpose, avoiding the pitfalls you’ve rightly flagged"
Modern _ frameworks prioritize adoption and depending on the framework to solve your companies needs will always fail. You need to choose a framework that fits your strategy and objectives, and adapt it to fit your needs.
Learn from others, but don't ignore the reality on the ground.
I feel we're talking past each other here. My original point was about which AI model is better for MY WORK. (I run a starup accelerator for first time founders) 4.5, in 30 seconds over minutes, provided more practical value to founders building actual businesses, and saved me time. While I appreciate your historical references and academic perspectives, they don't address my central argument about GPT-4.5's response being more pragmatically useful. The distinction between academic precision and practical utility is exactly what I'm highlighting. Founders don't need perfect theoretical models - they need frameworks that help them bridge vision and execution in the real world. When you bring up feudal production modes and von Moltke, we're moving further from the practical question of which AI response would better guide someone trying to align teams around a meaningful mission that drives business results. It's exactly why I formed the 2 prompts in the manner I did, I wanted to see if it was an academic or an expert.
My assessment stands that GPT-4.5's 30 seconds of thinking reflects well how mission operationalizes vision reflects how successful businesses actually work, not how academics might describe them in theoretical papers. I've read the papers, I've studied the theory deeply, but I also have NYSE and NASDAQ ticker symbols under my belt, from seed. That, is the whole point here.
If I were say in middle management and you asked me to "operationalize" the impact of mission statements, I would try to associate the existence of a mission statement on a team to some metric like financial performance.
If I was on a small development team and you asked me to "operationalize" our mission statement, I would probably make the same mistake the software industry always does, like trying it to tickets closed, lines of code, or even the Dora metrics.
Under my understanding of "operationalize" and the only way I can find it referenced related to mission statements themselves, I would actually de-emphasize deliverables, quality, stakeholders changing demands etc...
Even if I try to "operationalize" in a more abstract way, like define an impact score, which may not directly map to business objectives or even team building.
Almost every LLM offers a similar definition I offered above E.G.
> "operationalization" refers to the process of defining an abstract concept in a way that allows it to be measured and observed through specific, concrete indicators
Impact scores, which are subjective can lead to Google's shelfware problems, and even scrum rituals often leads to hard but high value tasks being ignored because of the incentives don't allow for it.
In both of your cites, they were situations where existing cultures were enhanced, not fully replaced.
Both were also short term, and wouldn't capture the long tail problems I am referencing.
Heck even Taylorism worked well for the auto industry until outside competition killed it. Well at least for the companies, consumers suffered.
The point is that "operationalization" specifically is counterproductive under a model, where infighting during that phase is bad.
If you care about delivering on execution, it would seems to be important to you. But I realize that you may not be targeting repeat work...I just don't know.
But I am sure some McKinsey BA probably has put that consern in a PDF someplace by now because the GoA Agile assessment guide is being incorporated and even ITIL and TOGAF reference that coal face paper I cited.
The BCGs and McKinseys of the world are absolutely going to shift to detection of common confabulations to show value.
While I do take any tools possible to make me more productive, correctness of content concerns me more than exact verbage.
But yes, different needs, I am in the nitch of rescuing failed initiatives, which admittedly is far from the typical engagement style.
To be honest the lack of scratch space on 4.5 compared to CoT models is the main blocker for me.
The irony of a company that has distilled the word's information complaining about another company distilling their model...
- Not writing things that go off in weird directions / staying grounded in "reality"
- Responding very well to tone preferences and catching nuance in what I say
It seems like it's less that it has a great "personality" like Claude, but that it's capable of adapting towards being the "personality" I want and "understanding" what I'm saying in ways that other models haven't been able to do for me.
GPT picked up on unspecified requirements almost instantly. It is subtle (and may be undesirable in some contexts). For example in my songs, I have to bracket the section headings, it picked up on that from my original input. All the other frontier models generally have to be reminded. Additionally, I separately asked for an edit to a music style description. When I asked GPT-4.5 to write a song all by itself, it included a music style description. No other model I have worked with has done this.
These are subtle differences, but in aggregate the model just generally needs less nudging to create what is required.
Other times it locks itself into a dull style and ignores what I ask of it and just produces boring generic garbage, and I have to wrangle it hard to get some of the spark back.
I have no idea what's going on inside, but just like with Stable Diffusion, it's fairly easy to make something that has the spark of genius, and is very close to being perfect, but getting the last 10% there, and maintaining the quality seems almost impossible.
It's a very weird feeling, it's hard to put into words what is exactly going on, and probably even harder to make it into a benchmark, but it makes me constantly flip-flop between scared of being how good the AI is, and questioning why I ever bothered with using it in the first place, as I would've progress much faster without it.
1) For coding (API) most probably will stick to Claude 3.5 / 3.7 - big market but still small comparing to all world wide problems
2) For non-coding API IMHO gemini 2.0 flash is the winner - dirty cheap (cheaper than 4o-mini), good enough and even better than gpt-4o, cheap audio and image input.
3) For subscription app ChatGPT is probably still the best but only slightly - they have the best advanced voice audio conversation but Grok will be probably eating their lunch here
Means nothing until they do
Claude is still stuck at 10 messages per day and gemini is less accurate/useful.
It strikes me as a form of fraud that steals my most precious resources: time and attention. I read creative writing to feel a human connection to the author. If the author is a machine and this is not disclosed, that's a huge lie.
It should be required that publishers label AI generated content.
Creativity has lost its meaning. Should it be illegal? The courts will take a long time to settle the matter. Reselling people's work against their will as creative machine output seems unethical, to say the least.
> It should be required that publishers label AI-generated content.
Strongly agree.
I much would prefer a super lightning fast model that is cheaper but the same quality as these frontier models.
Let me query these things to death.
This is how modern capitalism operates. We keep receiving new version numbers without any meaningful change — consumers keep updating their subscriptions out of habit — xyzAI PR managers, HR managers, corporate lawyers, and countless other bureaucrats keep collecting their paychecks, secretly dreaming of retirement. Meanwhile, xyzAI's top management burns money on endless acquisitions just to fight boredom, gradually morphing into xyz(ai)MEGAcorp, a monolith dabbling in everything from crude oil processing and styrofoam cups to web services and AI models.
No modern megacorporation is capable of creating anything fundamentally new beyond what worked for them once. Universal welfare and prosperity could have been achieved 60 years ago, but that would have disrupted the cycle. Instead, we got planned obsolescence, endless subscription models, and a world where everything “new” is just a slightly repackaged version of last year’s product.
>"Scaling to this size of model did NOT make a clear jump in capabilities we are measuring."
> "The jump from GPT-4o (where we are now) to GPT-4.5 made the models go from great to really great."
I think the main takeaway of GPT4.5 (Orion) is that it basically gives a perspective to all the "hit a wall" talk from the end of last year. Here we have a model that has been trained on by many accounts 10-100X the compute of GPT4, is likely several times larger in parameter count, but is only... subtly better, certainly not super-intelligent. I've been playing around w/ it a lot the past few days, both with several million tokens worth of non-standard benchmarks and talking to it and it is better than previous GPTs (in particular, it makes a big jump in humor), but I think it's clear that the "easy" gains in the near future are going to be figuring out how as many domains as possible can be approximately verified/RL'd.
As for the release? I suppose they could just have kept it internally for distillation/knowledge transfer, so I'm actually happy that they released it, even if it ends up not being a really "useful" model.
Aren't we making mistake of assuming benchmarks are purely 100% correct?
https://learn.microsoft.com/en-us/azure/ai-services/openai/c...
So I'm still on the Aug 24 release, which, with your reminding me, is not to be deprecated till Aug 2025, but that's less than 5 months from now, and we're skipping the Nov 2024 release just as OpenAI themselves have chosen to do.
Still, being early with a product and still often ”good enough” still takes them a long way. I think GPT-5 and where their competition will be then will be quite important for OpenAI though. I think the signs on the horizon is that everyone will close up on each other as we hit the diminishing returns, so the underlying business model, integrations, enterprise reach, marketing and market share will probably be king rather than the underlying LLM in 2026.
Since GPT-5 is meant to select the best model behind the scenes, one issue might be that users won’t have the same confidence in the model, feeling like it’s deciding for them or OpenAI tuning it to err on the side of being cheap.
I'm not even sure what is being alleged there—o1's reasoning tokens are kept secret precisely to avoid the kind of distillation that's being alleged. How can you distill a reasoning process given only the final output?
The first model has a much harder problem. It must figure out a way to iron out the differences between all the different models (i.e. multiple humans, or noise in the dataset) it is trained on, to come up with its single viewpoint of relationships. It must then be tuned to optimally fit its purpose, instead of simply operating like a (composite) rando human.
The second model then has these advantages:
1. Trained on a single model, whose behavior comes out of consistent learned relationships, as apposed to many inconsistent models with noise (lots of different humans operating differently at any given time for reasons beyond what the data can show).
2. Trained on post-tuned model behavior directly. One-step training is a considerably simpler problem than two-step. It doesn't learning anything it will need to unlearn or modify.
3. These simplifications both reduce the necessary size of the second model.
And with these combined advantages (a simpler more consistent single-system to model, smaller size, and fewer training steps), it can be trained much faster.
There, of course, can be disadvantages. The distilled model may not be as versatile, if the data used to prompt the original model isn't as diverse as the first models data. So prompt data for training still matters.
OpenAI was just first out of the gates, there'll always be some company that's first, essence is how they handle their leadership, and they've sadly been absolutely terrible and scummy.
Actually i think Google was a pretty good example of the exact opposite, decades of "actually not being evil", while openAI switched up 1 second after launch.
Google's overwhelming victory in search had ~ nothing to do with marketing.
There was a time when Google was thought of as a respectable, high-quality, smart and nimble company. That has faded as the marketing grew.
What is your point? OpenAI wasn’t the first out of that gate as your own argument cites Google prior. All these companies are predatory, who is arguing against that? OP said OpenAi was irrelevant. That’s just dumb. They are not. Feel free to advance an argument in favor of that narrative if you wish as I was just trying to provide a single example that shows that some of these lightweight models are building directly off the backs of giants spending the big money. I find nothing wrong with distillation and am excited about companies like DeepSeek.
Companies are 100% using these big players to generate synthetic data. Distillation is extremely powerful. How is this even in question?
The vast majority of users, which are over 300 million weekly, will mainly use 4o and whatever is the default. In the future they’ll use 4.5 and think it’s most human like and less robotic.
OpenAI has done a good job of making the model less important and the domain gptGPT.com more important.
Most of the time the model rarely matters. When you find something incorrect you may switch models but that rarely fixes the problem. Rewording a prompt has more value than changing a model.
Thousands of small, specific models are infinitely more efficient than a general one.
The more narrowed the task - the better algorithms work.
That's obvious.
Why are general models pushed so hard by its creators?
Their enormous valuations are based on total control over user experience.
This total control is justified by computational requirements.
Users can't run general models locally.
Giant data centers for billions are the moat for Model creators and corporations behind.