Microsoft and OpenAI end their exclusive and revenue-sharing deal
bloomberg.com
bloomberg.com
I think the biggest winner of this might be Google. Virtually all the frontier AI labs use TPU. The only one that doesn't use TPU is OpenAI due to the exclusive deal with Microsoft. Given the newly launched Gen 8 TPU this month, it's likely OpenAI will contemplate using TPU too.
Same with the CPU. Linux compiled faster on an M1 than on the fastest Intel i9 at the time, again using only 25% of the power budget.
And the M-series has only gotten better.
It is kind of sad Apple neglects helping developers optimize games for the M-series because iDevices and MacBooks could be the mobile gaming devices.
The context of this thread isn't consumer chips, but Apple's analog to an H/B200.
> The GPU is monstrously good. Depending on the workload, the M1 series GPU using 120W could beat an RTX 3090 using 420W.
You're just listing the TDP max of both chips. If you limit a 3090 to 120W then it would still run laps around an M1 Max in several workloads despite being an 8nm GPU versus a 5nm one.
> It is kind of sad Apple neglects helping developers optimize games for the M-series
Apple directly advocated for ports like Death Stranding, Cyberpunk 2077 and Resident Evil internally. Advocacy and optimization are not the issue, Apple's obsession over reinventing the wheel with Metal is what puts the Steam Deck ahead.
Edit (response to matthewmacleod):
> Bold of them to reinvent something that hadn't been invented yet.
Vulkan was not the first open graphics API, as most Mac developers will happily inform you.
Bold of them to reinvent something that hadn't been invented yet.
Surprised Apple didn't create a TPU-like architecture. Another misstep from John Gianneadrea.
Apple had the technology to scale down a GPGPU-focused architecture just like Nvidia did. They had the money to take that risk, and had the chip design chops to take a serious stab at it. On paper, they could have even extended it to iPhone-level edge silicon similar to what Nvidia did with the Jetson and Tegra SOCs.
(Like “I want to do object detection for cutting people into stickers on device without blowing a hole in the battery, make me a chip for that”.)
OpenGL had become too unmanagable which is why devs moved to DirectX.
Unless you meant a different one?
You're cooked if you actually believe this
The thing that Apple has always been excellent at is efficiency - even during the Intel era, MacBooks outclassed their Windows peers. Same CPU, same RAM, same disks, so it definitely wasn't the hardware, it was the software, that allowed Apple to pull much more real-world performance out of the same clock cycles and power usage.
Windows itself, but especially third party drivers, are disastrous when it comes to code quality, and they are much much more generic (and thus inefficient) compared to Apple with its very small amount of different SKUs. Apple insisted on writing all drivers and IIRC even most of the firmware for embedded modules themselves to achieve that tight control... which was (in addition to the 2010-ish lead-free Soldergate) why they fired NVIDIA from making GPUs for Apple - NV didn't want to give Apple the specs any more to write drivers.
I think that's a valid demand, considering Nvidia's budding commitment to CUDA and other GPGPU paradigms. Apple, backing OpenCL, would have every reason to break Nvidia's code and ship half-baked drivers. They did it with AMD's GPUs later down the line, pretending like Vulkan couldn't be implemented so they could promote Metal.
Apple wouldn't have made GeForce more efficient with their own firmware, they would have installed a Sword of Damocles over Nvidia's head.
It was even worse than that, they just stopped updating OpenGL for years before either Vulkan or Metal existed at all. Taking a Macbook and using bootcamp would instantly raise the GPU feature level by several generations just because Apple's GPU drivers were so fucking old & outdated.
There are other workloads where the M1 actually beats the 3090.
Apple does plenty of hyping but it's always cute when irrational haters like you put them down. The M1 was (well, is) a marvel and absolutely smokes a 3090 in perf per watt.
Find or link these workloads you think exist, please
> The M1 was (well, is) a marvel and absolutely smokes a 3090 in perf per watt.
The GTX 1660 also smokes the 3090 in perf per watt. Being more efficient while being dramatically slower is not exactly an achievement, it's pretty typical power consumption scaling in fact. Perf per watt is only meaningful if you're also able to match the perf itself. That's what actually made the M1 CPU notable. M-series GPUs (not just the M1, but even the latest) haven't managed to match or even come close to the perf, so being more efficient is not really any different than, say, Nvidia, AMD, or Intel mobile GPU offerings. Nice for laptops, insignificant otherwise
Also note how the M1 Ultra is pushing 2/3 of the FPS of the 3090 despite 1/3 of the power budget and the game itself being poorly optimized for the M-series architecture.
And here[1] you have it smoking an Intel i9 12900K + RTX 3900. The difference doesn't look too impressive until you realize the power envelope for that build is 700-800W.
Also, the GTX 1660 (technically an RTX 2000 series, but whatever) is about 26% less efficient than an 3090[2].
> Being more efficient while being dramatically slower
That's my whole point and what you're refusing to see. The M1 is not dramatically slower than an i9 or 3090 despite having dramatically lower power use.
The proof for this will really start to come once Qualcomm and Mediatek have gotten a handle on their PC ARM chips and Valve decides they're good enough for a Steam Deck 2 or 3. You'll get to see 2-3x the battery life along a modest performance increase.
[0]https://techjourneyman.com/img/blog/m1-ultra-vs-rtx-3090-ben...
[1]https://techjourneyman.com/img/blog/m1-ultra-vs-intel-i9-129...
[2]https://bestvaluegpu.com/comparison/geforce-rtx-3090-vs-gefo...
Oh, GFXBench not geekbench.
Realistically that 506 fps result is probably CPU bottlenecked, not that aztec ruins is all that relevant. It's a very old benchmark, released in 2018, that was destroyed for mobile GPUs, so realistically is using a 2010-ish GPU feature set.
If that's your use case, great. But it's not significant at all.
> And here[1] you have it smoking an Intel i9 12900K + RTX 3900.
Not using the GPU, so irrelevant. Also not using 700-800w
> Also, the GTX 1660 (technically an RTX 2000 series, but whatever) is about 26% less efficient than an 3090[2].
"bestvaluegpu" I've never heard of but holy AI slop nonsense batman. Taking 3dmark score and dividing it by TDP is easily one of the worst ways to compare possible.
Here's actual perf/watt results taken by, you know, actually measuring the power draw https://www.techpowerup.com/review/msi-geforce-rtx-3090-gami...
For a Qwen 3.6 35B / 3B MoE, 4-bit quant:
- parsing a 4k prompt on a M4 Macbook Air takes 17 seconds before generating a single token.
- on an M4 Max Mac Studio it's faster at 2.3 seconds
- on an RTX 5090, it's 142ms.
RTX 5090 uses more power than an M4 Max Mac Studio but it's not 16x more power.
Open AI has nothing. Their tech will rapidly be devalued by free models the moment they stop lighting stacks of cash on fire.
The parent post was arguing that they can do this now because they are lighting stacks of cash on fire. And once they stop doing that, their LLM lead will be gone in a hurry. They appear to not have a moat, like other more established players do.
Arguably, Apple isn't even in a boat right now. At least AMD and Intel both ship hardware that synergizes with CUDA - Apple jumped off that ship, their hardware doesn't even come up in infrastructure discussions where AMD, Intel and Nvidia are taken for granted.
In mid-2028 we have N2E/N2P with around 15% greater transistor density than today's N3P, and by EOY2028 we'll likely have A14 with about 35-40% density improvement.
Meanwhile, we'll be on LPDDR6 by that point, which takes M-series Pros from 307GB/s -> ~400GB/s, and Max's from 614GB/s -> ~800GB/s.
Model improvements obviously will help out, but on the raw hardware front these aren't in the ballpark for frontier model numbers. An H100 has 3TB/s memory bandwidth, fwiw
To run a 8 bit quantized version of that you need roughly 5TB of RAM.
Today that is around 18 NVidia B300. That's around $900,000, without including the computers to run them in.
It's true that the capability of open source models is improving, but running actual frontier models on your MPB seems a way off.
[1] https://x.com/elonmusk/status/2042123561666855235?s=20 (and Elon has hired enough people out of those labs to have a fair idea)
You could run it on a cluster of nodes that each do some mix of fetching parameters from disk and caching them in RAM. Use pipeline parallelism to minimize network bandwidth requirements given the huge size. Then time to first token may be a bit slow, but sustained inference should achieve enough throughput for a single user. That's a costly setup of course, but it doesn't cost $900k.
Not sure this is a MBP either.
Today's LLMs are able pack much more capabilities into fewer parameters compared to 2023. We might still be at the very rudimentary phase of this technology there are low-hanging efficiency gains to be had left and right. These models consume many orders of magnitude more energy than a human brain, this all seems like room for improvement.
The right question: is there a law in information theory that fundamentally prevents a 70B model of any architecture from being as smart as Opus 4.7?
Or so they say.
If it's true then that just shows how far behind the cloud providers are lagging while wasting investor money.
(There's a huge amount of diminishing returns in increasing parameter counts and the intelligent AI company should be hard at work figuring out the optimal count without overfitting.)
In practice unless you're doing some kind of deep research thing with the cloud, it'll try to optimize mostly for time and get you a good enough answer rather than spending an hour or two. An hour of cloud searching with huge data stores is not equivalent to an hour of local agentic searching, presumably.
I think that problem will improve a little in the coming years as we kind of create optimized data curation, but the information world will keep growing so the advantage will likely remain with centralized services as long as they offer their complete potential rather than a fraction.
https://www.reuters.com/business/retail-consumer/openai-taps...
TPUs are at least dogfooded by Google deepmind, no team AFAIK has gotten the AMD stack to train well.
Pull quotes:
AMD’s software experience is riddled with bugs rendering out of the box training with AMD is impossible. We were hopeful that AMD could emerge as a strong competitor to NVIDIA in training workloads, but, as of today, this is unfortunately not the case. The CUDA moat has yet to be crossed by AMD due to AMD’s weaker-than-expected software Quality Assurance (QA) culture and its challenging out of the box experience.
[snip]
> The only reason we have been able to get AMD performance within 75% of H100/H200 performance is because we have been supported by multiple teams at AMD in fixing numerous AMD software bugs. To get AMD to a usable state with somewhat reasonable performance, a giant ~60 command Dockerfile that builds dependencies from source, hand crafted by an AMD principal engineer, was specifically provided for us
[snip]
> AMD hipBLASLt/rocBLAS’s heuristic model picks the wrong algorithm for most shapes out of the box, which is why so much time-consuming tuning is required by the end user.
etc etc. The whole thing is worth reading.
I'm sure it has (and will continue to) improved since then. I hear good things about the Lemonade team (although I think that is mostly inference?)
But the NVidia stack has improved too.
if they had this management attitude, they wouldn't have been so far behind so as to need this action in the first place!
> “Are we afraid of our competitors? No, we’re completely unafraid of our competitors,” said Taylor. “For the most part, because—in the case of Nvidia—they don’t appear to care that much about VR. And in the case of the dollars spent on R&D, they seem to be very happy doing stuff in the car industry, and long may that continue—good luck to them.
https://arstechnica.com/gadgets/2016/04/amd-focusing-on-vr-m...
"car industry" is linked to the GPU-accelerated self-driving car work, ie, making neural networks run fast on GPUs: https://arstechnica.com/gadgets/2016/01/nvidia-outs-pascal-g...
Maybe Amazon is an example how this happens even to hardware divisions within software/logistics companies
Amazon's compensation strategy, in which you primarily get a raise years in the future for tricking your management chain into promoting you is definitely bearing its rotten fruit.
Anthropic did retire an interview take-home assignment involving optimising inference on exotic hardware, because Claude could one shot a solution, but that was clearly a whiteboard hypothetical instead of a real system with warts, issues and nuance.
ROCm works great too, the only issue i have had is that my machine froze a couple of times as it used 100% of the graphics and the OS had nothing left. Since moving to vulcan i stopped getting these errors apart from a little UI slowdown when i had 4 models loaded at the same time taking turns.
Im also on a i7 6700 with 32gb DDR4 so im sure that is causing more slowdowns then the graphics card.
int8 quantization seems like it's almost supported, but not quite. speeds drop to a fraction of full precision speed and the server seems like it intermittently hangs. int4 quantization not supported. fp8 quantization not supported.
again, maybe AMD is just being lazy with what they've provided, but it's not a great look.
right now the fastest smart model i can run is full precision qwen3-32b. with 120 parallel requests (short context) i'm getting PP @ 4500 tokens/sec and TG @ 1300 tokens/sec
From the papers I've read and the labs that I have worked in personally, I would say that most scientists developing Deep learning solutions use CUDA for GPU acceleration
But AMD does not want to pay these specialized SWEs the market rate. Their existing SWEs would be up in arms saying, basically, "what are we, chopped liver??", or so the thinking goes.
So AMD is stuck with a shitty software stack which cannot compete with CUDA.
If I were making such decisions, I would just cull the number of existing SWEs down by 50%, and double the pay for remaining ones. And then go out and hire some top talent to build a good software stack.
Freudian slip?
What's unclear to me is how much Google uses GPUs for their own stuff. Yes Gemini runs on GPUs now, so that Google can sell Gemini on-prem boxes (recent release announced last week), but is any training or inference for Gemini really happening on GPUs? This is unclear to me. I'd have guessed not given that I thought TPUs were much cheaper to operate, but maybe I'm wrong.
Caveat, I work at Google, but not on anything to do with this. I'm only going on what's in the press for this stuff.
Do you have any more information on this? I only found this article about it: https://venturebeat.com/technology/googles-gemini-can-now-ru...
It mentions that Gemini can run on eight NVIDIA GPUs, but not which GPU and which Gemini model. Either way, this puts an upper bound of 288 * 8 = 2304 GB on the size of the Gemini model, which as far as I know has been a secret until now.
They'll presumably catch up, there is no monopoly on talent held by the US. And, that's more true than ever now that the US is actively hostile to immigrants. Scientists who might have come to the US three years ago have little reason to do so now.
But even that distinction is only temporary, since we're determined to piss away any remaining research lead that draws people in.
Hopefully the next administration will work at actively reversing the damage, with incentives beyond just "we pinky-promise not to haul you at gunpoint to a concrete detention center and then deport you to Yemen".
Won't be enough to undo the damage. The US would have to do a full about face, prosecute crimes of the current administration and enact serious core reforms to make it impossible for things to drastically change again in 4 years. Also known as, never going to happen because even the current opposition party doesn't actually want structural change. The world has seen how bad the US can get from a single election, and that isn't changing any time soon.
Been saying that about EU and China for decades now.
Yet the top European and Chinese still come to the US. Even in April 2026.
You could reasonably say that "A majority of frontier labs uses TPU to train and serve their model."
[1]: We train and run Claude on a range of AI hardware—AWS Trainium, Google TPUs - April 6th, Anthropic on Google and Broadcom partnership [2]: "[Apple foundation model]... builds on top of JAX and XLA, and allows us to train the models with high efficiency and scalability on various training hardware and cloud platforms, including TPUs and both cloud and on-premise GPUs" - Apple in 2024
He's been saying whatever is good for Nvidia for years now without any regard for truth or reason. He's one of the least trustworthy voices in the space.
Google's TPUs have obvious advantages for inference and are competitive for training.
For inference? This is from July 2025: OpenAI tests Google TPUs amid rising inference cost concerns, https://www.networkworld.com/article/4015386/openai-tests-go... / https://archive.vn/zhKc4
> ... due to the exclusive deal with Microsoft
This exclusivity went away in Oct 2025 (except for 'API' workloads).
OpenAI has contracted to purchase an incremental $250B of Azure services, and Microsoft will no longer have a right of first refusal to be OpenAI’s compute provider.
https://blogs.microsoft.com/blog/2025/10/28/the-next-chapter... / https://archive.vn/1eF0VHow is this helping OpenAI?
I feel this looks like a nice thing to have given they remain the primary cloud provider. If Azure improves it's overall quality then I don't see why this ends up as a money printing press as long as OpenAI brings good models?
[1] https://www.wsj.com/tech/ai/openai-and-microsoft-tensions-ar...
And on top of that, OpenAI still has to pay Microsoft a share of their revenue made on AWS/Google/anywhere until 2030?
And Microsoft owns 27% of OpenAI, period?
That's a damn good deal for Microsoft. Likely the investment that will keep Microsoft's stock relevant for years.
own 27%. but are entitled to OpenAI profits of 49% for eternity (if OpenAI is profitable or government steps in)
Where is the 49% coming from? The new deal does not talk about that.I doubt it
AWS's us-east-1 famously takes down either a bunch of companies with it, or causes global outages on the regular.
AWS has a terrible, terrible user interface partly because it is partitioned by service and region on purpose to decrease the "blast radius" of a failure, which is a design decision made totally pointless by having a bunch of their most critical services in one region, which also happens to be their most flaky.
Or maybe you can provide a better explanation for why users had to “hunt” through hundreds(!) of product-region combinations to find that last lingering service they were getting billed $0.01 a month for?
This just doesn’t happen in GCP or Azure. You get a single pane of glass.
For all its flaws at least Azure has consistent UI.
Yes, by design.
Conceptually this improves velocity and reduces the blast radius of failure.
In practice, everything depends on IAM, S3, VPC, and EC2 directly or indirectly, so this doesn't help anywhere near as much as one would think.
Azure and GCP have a split control plane where there's a global register of resources, but the back-end implementations are split by team.
That way the users don't see Conway's Law manifest in the browser urls... as much. (You still do if you pay attention! In Azure the "provider type" is in the path instead of the host name.)
Hm yes but I hate working with it as a customer because it is so confusing. Everything works differently and there is a lot of overlap (several services exist that do the same thing). It seems like an amateurish patchwork.
I understand it has benefits to have different teams working on different services but those teams should still be aligned in terms of UX and basic concepts.
You could argue now that that's no excuse anymore given it's one of the most valuable companies in the world, but that would dismiss the fact they have other priorities than a complete UI overhaul for consistency, and that rewrites are very dangerous, for instance people are already used to the UX pitfalls in the console, it's the devil they know, and changing that will be upsetting to the vast majority of users.
So there you have it. You know what you are getting into, AWS is a behemoth and it's 2026. Don't use the console like it's 2010. Use IaC for any nontrivial work, otherwise you only have yourself to blame.
But as a customer I absolutely hate working with AWS tech. Their stuff is a mess and I feel like I shouldn't have to get my head around their idiosyncracies. I prefer Azure even though Microsoft is a terrible company to work with. I find the AWS people and attitude a lot nicer but their services are a mess. If I do something new I prefer using Azure despite having to work with Microsoft.
Microsoft is not a "trusted partner" wanting the best for you, they're always trying to screw you over in favour of selling some new crap to your boss. Always that stupid sales drive, whereas the people from AWS are very focused on building success together. But still, their tech is just so bad unless you spend all your days working with it and really become an expert on what they offer. That's not tech, just corporate servitude. And I've always avoid that position, I don't want my career tied to some big brand name. I don't want to be "the AWS expert" or "the MS expert".
But I have to say I hate cloud (and "the world according to big tech") in general, and it's one of the reasons I'm not really involved in server infrastructure anymore these days. I'll gladly automate but not with their tooling, I prefer something more open and not tied to specific vendors. But I rarely work with that now. So yeah when that happens I'm making a one-off unicorn and figuring out all the Infra as code stuff is not worth it.
But azure wins most prizes for being terrible becuase, among other things, https://isolveproblems.substack.com/p/how-microsoft-vaporize.... It's not the worst provider maybe because oracle is somehow still kicking around.
Its just a bad product. Just like windows, OneDrive, teams and basically everything Microsoft has pumped out in the past decade.
Microsoft is in the top 5 most valuable companies in the world. It's got azure that is a huge cloud provider. And yet it was utterly unable to present its answer in the AI race. Not even a bad model with a half baked harness. Nothing. And meanwhile they are trying to port NTFS to low powered FPGAs because insanity. Just let that sink in.
There’s no upper limit to their financial stupidity.
FaceBook largely requires an Apple iPhone, Apple computer, "Microsoft" computer, "Google" phone, or a "Google" computer to use it. At any point one of those companies could cut FaceBook off (ex. [1]).
The Metaverse was a long term goal to get people onto a device (Occulus) that Meta controlled. While I think an AR device is much more useful than VR; I'm not convinced that it's a mistake for Meta to peruse not being beholden to other platforms.
[1]: https://arstechnica.com/gadgets/2019/01/facebook-and-google-...
If it's actual holograms like in Star Wars? Sure, why not. Get the visual and body language cues of the rest of the room but no one has to physically congregate at a location.
But pixelated, cartoon avatars? Yeah, wtf.
Devoid of other context, it’s hard to disagree. But your parent comment only asserted that the metaverse specifically as proposed by Facebook was an obviously stupid idea.
Maybe they should have spent that on the facebookphone
Patrick Boyle did a nice video a few weeks back: https://www.youtube.com/watch?v=8BaSBjxNg-M
The headsets don’t really make sense to me in the way you’re describing. Phones are omnipresent because it’s a thing you always just have on you. Headsets are large enough that it’s a conscious choice to bring it; they’re closer to a laptop than a phone.
Also, the web interface is like right there staring at them. Any device with a browser can access Facebook like that. Google/Apple/Microsoft can’t mess with that much without causing a huge scene and probably massive antitrust backlash.
It's kind of like Microsoft with copilot - the idea about having an AI assistant that can help you use the computer is great. But it can't be from Microsoft because people don't trust them with that.
I think VR has more niche uses than the craze implied. It’s got some cool games, virtual screens for a desktop could be cool someday, but I don’t see a near future where they replace phones.
Until VR is done via glasses or some wire you stick in your neck matrix style, it will never take off
I don't want to wear glasses for extended periods of time either, but I've had to get used to them.
Glasses also have a pretty compelling reason to wear them.
I don't think VR has as compelling of a reason, at least so far. Even if it does, you need people to get far enough in to see that reason which is a hurdle when it involves a device they don't want to wear.
They address the friction of use issue being discussed, they’re even more discrete and available than a phone. And they are getting a lot of general public recognition, albeit not for the best reasons (people discretely filming, for genuine social media reactions but also for other reasons..).
Their tech is improving at a decent pace and they’ve recently put out a product that is both ready for consumer (at least with select use cases) adoption, and actually reasonably available to the public.
If you’re talking about the Meta Ray Ban glasses, I wouldn’t really call that a successor. There’s no AR or VR to them that I can tell; just glasses with speakers, a mic and a camera. It’s a neat product, but not a platform in the way VR was meant to be. They also have real competition. I do actually own a pair of the Bose headphone sunglasses, which are practically the same product without a camera (which I’m sure they could add if they wanted). Unless people suddenly care about the Meta AI integration, and again; Bose or someone else could add a phone companion app.
They have two current Meta Ray Ban options, the “Gen 2” and the “Display”, the latter of which does have an AR component.
Apple was directly (and IMO arguably illegally) shutting down Facebook teams and products by playing app store chicken on refusing to allow Facebook to publish updates on a week-to-week basis. Literally would throw down and refuse unless some features were blocked. It came to a head where Zuck literally called Tim Cook during a keynote to push it through.
They also literally had reverse-engineering teams cracking open the Facebook app on a regular basis, which we discovered because of some internal methods we figured out how to invoke with some clever indirection. There was a chicken-and-egg problem and they eventually developed facilities to automatically instrument private method invocations to comprehensively defeat clever static analysis circumvention workarounds.
Also, VR hasn't failed, but it's gone silent and coasted when investing in VR growth took the backseat to investing AI. They made a couple of bad bets in VR but a lot of good ones so it was warranted, but not exactly a failure.
Apple trying to block Facebook is different than Apple trying to prevent Facebook from violating App Store standards. There was a time where the Facebook app was practically malware with all the tricks it tried to pull to Hoover up data.
I don’t know in what world I would describe Metas VR as anything but a failure. There was a brief period where I knew a few people with Quests. Most used them for the novelty and dropped them, a few played games on them, and I don’t know anyone that still owns one. I’m deep in the gaming community and haven’t heard anyone mention a Quest in years. Steam VR is almost equally quiet other than occasional nostalgia.
True but the an app gives Facebook much more user data for targeting which dramatically increases revenue per ad. Persistent user data that's largely unconstrained by privacy safeguards is the holy grail. The mobile browsers are also controlled by Apple and Google, so despite the web being 'open', when one of them makes even minor changes to increase browser privacy defaults, it can have major impact on Facebook's revenue.
Maybe a niche product could do it, but good luck selling a laptop that won't open FB
If it was really their goal, they would have made an Android competitor. Maybe a fork like amazon did and sell phones that supported it.
Zuckerberg had one great idea (and then it wasn't really his idea) at the right time, since then he failed over and over at everything else. 'Internet for all', remember ?
I really wouldn't give them the benefit of the doubt.
Some of those companies can cut off invasive apps.
There is no risk of facebook.com getting blocked. And absolutely nobody is going to prefer a headset over a website for doing facebook things.
But thinking AR/VR was the way to go is a failure to read the room. If anything the up and coming generations seem to be recoiling from tech.
Regardless, as Microsoft found, it's too late for a 3rd platform and it seems somehow that there's only room in the world for two.
(Meta would have done better to start up a line of caffeinated sugar drinks.)
How?
valued at --which I'd say is a reasonable distinction to make right about now
https://www.reuters.com/business/openai-cfo-says-annualized-...
I can easily generate double that revenue, by selling $20 bills for $10.
If GitHub flipped a switch and enabled IPv6 it would instantly break many of their customers who have configured IP based access controls [1]. If the customer's network supports IPv6, the traffic would switch, and if they haven't added their IPv6 addresses to the policy ... boom everything breaks.
This is a tricky problem; providers don't have an easy way to correlate addresses or update policies pro-actively. And customers hate it when things suddenly break no matter how well you go about it.
https://news.ycombinator.com/item?id=47790889For every customer which has access controls configured based on IPv4 (sounds crazy enough already), GitHub would configure a trivial DENY ALL policy for IPv6. Problem solved.
With that, the customers who don't use filtering by IPv4 would be able to use IPv6. Those who do use access control by IPv4 ranges would have time to sort out their IPv6 setup, without having anything broken at the moment when IPv6 is enabled.
That's rather the problem - there's no trivial way to mimic that policy transparently while enabling IPv6, because most stacks will default to using IPv6 if they're dual-homed and expose both, and won't fall back if IPv6 connects but gives an error. (Offhand, I think the best you could do would be to tell everyone that you're migrating to a new URI scheme to allow cloning, with IPv6 enabled, and that as part of that, you'll have to update your allow/deny rules, then, after a truly astonishingly long time and lots of nagging of anyone who never does it, make the old path an alias of the new one and let the last remaining people break.)
They still run their own platform.
https://thenewstack.io/github-will-prioritize-migrating-to-a...
I think the differentiator is Team, which Google for some mysterious reason can't build or doesn't want to.
But OpenAI had announced a shift towards b2b and enterprise. It makes sense for their models to be available on the different cloud providers.
But if I own 49% of a company and that company has more hype than product, hasn't found its market yet but is valued at trillions?
I'm going to sell percentages of that to build my war chest for things that actually hit my bottom line.
The "moonshot" has for all intents and purposes been achieved based on the valuation, and at that valuation: OpenAI has to completely crush all competition... basically just to meet its current valuations.
It would be a really fiscally irresponsible move not to hedge your bets.
Not that it matters but we did something similar with the donated bitcoin on my project. When bitcoin hit a "new record high" we sold half. Then held the remainder until it hit a "new record high" again.
Sure, we could have 'maxxed profit!'; but ultimately it did its job, it was an effective donation/investment that had reasonably maximal returns.
(that said, I do not believe in crypto as an investment opportunity, it's merely the hand I was dealt by it being donated).
And Microsoft only paid $10B for that stake for the most recognizable name brand for AI around the world. They don't need to "hedge their bets" it's already a humongous win.
Why let Altman continue to call the shots and decrease Microsoft's ownership stake and ability to dictate how OpenAI helps Microsoft and not the other way around?
That's a flawed argument. Why wouldn't you want to hedge a risky bet, and one that's even quite highly correlated to Microsoft's own industry sector?
my impression is that many of these "investments" are structured IOUs for circular deals based on compute resources in exchange for LLM usage
Maybe that will be true someday. But, right now, they are burning billions of dollars every quarter. Their expenses far far outweigh their income and they are nowhere near profitability.
Except revenue. Not one company is in the black. That’s a pretty important measure you’re ignoring.
They come from just printing more shares every quarter, diluting every shareholder.
I'm not convinced what you believe is true. Dilution is possible undoubtedly, but perhaps if majority shareholders approve it. And even then, likely regulated by a suite or other constraints.
Genuine question because I feel like I’m maybe missing something!
The longer answer is; you never know whats coming next, bitcoin could have doubled the day after, and doubled the day after that, and so on, for weeks. And by selling half you've effectively sacrificed huge sums of money.
The truth is that by retaining half you have minimised potential losses and sacrificed potential gains, you've chosen a middle position which is more stable.
So, if bitcoin 1000 bitcoing which was word $5 one day, and $7 the next, but suddenly it hits $30. Well, we'd sell half.
If the day after it hit $60, then our 500 remaining bitcoins is worth the same as what we sold, so in theory all we lost was potential gains, we didn't lose any actual value.
Of course, we wouldn't sell we'd hold, and it would probably fall down to $15 or something instead.. then the cycle begins again..
Speculation based on selling at below cost.
> it’s not valued at trillions
Fair, it's only $852 billion. Nowhere near trillions.. you got me.
OpenAI's adjusted gross margin: 40% in 2024, 33% in 2025. Reason cited: inference costs quadrupled in one year.
Internal projections leaked to The Information: ~$14B loss on ~$13B revenue in 2026. Cumulative losses through 2028: ~$44B.
https://finance.yahoo.com/news/openais-own-forecast-predicts...
A business burning more than a dollar for every dollar of revenue is a lot of things. "Quite profitable" is not one of them.
If you're reaching for the SaaStr piece on API compute margins hitting ~70% by late 2025: yes, that exists, and it describes one tier. The volume is on the consumer side. The consumer side is the bit on fire. Pointing at the API margin and calling the whole business profitable is the financial equivalent of weighing yourself with one foot off the scale.
The original argument, in case it got lost: Microsoft holds (held) a 49% stake in a company projecting another $44B of cumulative losses through 2028, against unit economics that depend on competitors not catching up. That's textbook hedge-the-bet territory. "They have paying customers" doesn't refute that, MoviePass had paying customers too.
I didn’t call the business profitable, I said that inference is profitable. I was responding to your assertion that they’re speculating by selling below cost. Which isn’t true; they’re selling inference, profitably. They’re losing money because they’re investing in the next model. The company isn’t profitable, it might never be profitable, but the product they’re selling is profitable. So calling it speculation based on selling something below cost is just factually incorrect.
It isn't. Frontier model training is the cost of having a product to sell inference on next year. Stop training and the inference margin decays on the timescale of the next competitor release, which in 2026 is measured in weeks. So "the product is profitable, the company is just investing" describes a business where the investment is structurally non-optional and structurally larger than the product margin. That's the definition of selling below cost at the level that matters, which is the level you're hedging at when you hold 49%.
McDonald's is profitable because a Big Mac in 2027 costs roughly what a Big Mac in 2026 cost to make. OpenAI's product depreciates to zero on a 12-month cycle unless they spend ~$40B keeping it ahead. That's the disagreement, and "but inference itself has positive margin" doesn't resolve it, it just relocates it.
Hrm..
The point is that losing money isn't a sure sign that a business is doomed. Who knows where OpenAI will end up, but people still line up to invest. Those investors have billions reasons to be due diligent. Unlike what's claimed around here, most of investors aren't stupid. You yourself wouldn't be stupid either if money is at stake.
For OAI to be a purely capitalist venture, they had to rip that out. But since the non-profit owned control of the company, it had to get something for giving up those rights. This led to a huge negotiation and MSFT ended up with 27% of a company that doesn’t get kneecapped by an ethical board.
In reality, though, the board of both the non-profit and the for profit are nearly identical and beholden to Sam, post–failed coup.
Looks like Nadella is slowly realizing that it is his short and curlies that are in the vice grip in the "If you owe the bank $100 vs $100M" sense?
Deepseek v4 is good enough, really really good given the price it is offered at.
PS: Just to be clear - even the most expensive AI models are unreliable, would make stupid mistakes and their code output MUST be reviewed carefully so Deepseek v4 is not any different either, it too is just a random token generator based on token frequency distributions with no real thought process like all other models such as Claude Opus etc.
Once a new model or a technique is invented, it’s just a matter of time until it becomes a free importable library.
But yes, they do have similar constraints.
Because for Deepseek is pretty straightforward censorship.
https://claude.ai/share/ac4d2041-4328-4511-904a-ff8b1cbfc0bb
Looks like self-censoring to me. Grok has no problems answering, so it is not a technical limitation:
https://grok.com/share/bGVnYWN5_341d6428-ea7d-4b84-ad4d-8df9...
Grok used a book as a reference.
It's not like ethnicity is a fact you infer from looking at someone.
Now ask Deepseek about what happened in Tiananmen Square and watch what censorship actually looks like.
It literally knows the facts, but then there's a layer that prevents it from stating the facts.
That's censorship.
It's not an opinion, it's not a choice when facing a gradient, it's just an historical known fact.
I asked early, at the time people were posting various jailbreaks, never worked.
On a side note, any self hosted model I can get for my PC? I have 96 GB of RAM.
Try the 8 bit quantized version (UD-Q8_K_X) of Qwen 3.6 35B A3B by Unsloth: https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF
Some people also like the new Gemma 4 26B A4B model: https://huggingface.co/unsloth/gemma-4-26B-A4B-it-GGUF
Either should leave plenty of space for OS processes and also KV cache for a bigger context size.
I'm guessing that MoE models might work better, though there are also dense versions you can try if you want.
Performance and quality will probably both be worse than cloud models, though, but it's a nice start!
Wait - what?
So if you or anyone passing by was curious, yes you can get accurate output about the Chinese head of state and political and critical messages of him, China and the party
Its final answer will not play along
If you want an unfiltered answer on that topic, just triage it to a western model, if you want unfiltered answers on Israel domestic and foreign policy, triage back to an eastern model. You know the rules for each system and so does an LLM
Until LLM's I'd never in my life heard someone suggest we lock up the compiler when it goofs up and kills someone, but now because the compiler speaks English we suddenly want to let people use it as a get out of jail free card when they use it to harm others.
The humans I did work with were very very bright. No software developer in my career ever needed more than a paragraph of JIRA ticket for the problem statement and they figured out domains that were not even theirs to being with without making any mistakes and rather not only identifying edge cases but sometimes actually improving the domain processes by suggesting what is wasteful and what can be done differently.
Seriously? I would like to remind you that every single mistake in history until the last couple of years has been made by humans.
And yes, there were always incompetent folks but those were steered by smarter ones to contain the damage.
Also worked with people who were frustrated that they had to force push git to "save" their changes. Honestly, a token-box I can just ignore, would be an upgrade over this half of the team.
*For some definitions of individual agency. Incompatiblists not included.
Nevermind the fact that they are literally able to introspect human cognition and presumably find non verbal and non linear cognition modes.
Are they, though? Or are they just predicting their own performance (and an explanation of that performance) on input the same way they predict their response to that input?
Humans say a lot of biologically implausible things when asked why they did something.
For e.g. ask any model "which class of problems and domains do you have a high error rate in?".
The USA has the biggest, but there lies their disadvantage
In the USA building bigger, better frontier models has been bigger data centres, more chips, more energy.
China has had to think, hard. Be cunning and make what they have do more
This is a pattern repeated in many domains all through the last hundred years.
tant pis
... and who knows if we, humans, are not just merely that.
However, for reviewing, I want the most intelligent model I can get. I want it to really think the shit out of my changes.
I’ve just spent two weeks debugging what turned out to be a bad SQLite query plan (missing a reliable repro). Not one of the many agents, or GPT-Pro thought to check this. I guess SQL query planner issues are a hole in their reviewing training data. Maybe Mythos will check such things.
With this new workflow, however, we should, uncompromisingly, steer the entire code review process. The danger here, the “slippery slope,” is that we’re constantly craving for more intelligent models so we can somehow outsource the review to them as well. We may be subconsciously engineering ourselves into obsolescence.
This is such an interesting time to be in. Truly skilled developers like Rob Pike really don’t like AI, but many professional developers love it. I side with Mr. Pike on it all.
I am not a skilled developer like he is, but I do like to think about what I’m doing and to plan for the future when writing code that might be part of that future. I like very simple code which is easy to read and to understand, and I try quite hard to use data types which can help me in multiple ways at once. The feeling when you solve a problem you’ve never solved before is indescribable, and bots strip all of that away from you and they write differently than I would.
I don’t think any bot would ever come up with something like Plan9 without explicit instructions, and that single example showcases what bots can’t do: think about what is appropriate when doing something new.
I don’t know what is right and what is wrong here, I just know that is an interesting time.
Do they have monthly subscriptions, or are they restricted to paying just per token? It seems to be the latter for now: https://api-docs.deepseek.com/quick_start/pricing/
Really good prices admittedly, but having predictable subscriptions is nice too!
Edit: it looks like it's 75% off right now which is really an incredible deal for such a high caliber frontier model.
I'm asking because with most providers (most egregiously, with Anthropic) it doesn't work that way because the API pricing is way higher than any subscription and seemingly product/company oriented, whereas individual users can enjoy subsidized tokens in the form of the subscription. If DeepSeek only offers API pricing for everyone, I guess that makes sense and also is okay!
I'm not smart enough to reduce LLMs and the entire ai effort into such simple terms but I am smart enough to see the emergence of a new kind of intelligence even when it threatens the very foundations of the industry that I work for.
Curious about your definition of these terms.
Just because you are impressed by the capabilities of some tech (and rightfully so), doesn't mean it's intelligent.
First time I realized what recursion can do (like solving towers of hanoi in a few lines of code), I thought it was magic. But that doesn't make it "emergence of a new kind of intelligence".
To me, that's intelligence and a measurable direct benefit of the tool.
I just did my taxes using a sophisticated spreadsheet. Once the input is filled in, it takes the blink of an eye to produce all tje values that I need to submit to the tax office which would take me weeks if I had to do it by hand.
Just the other day I used an excavator to dig a huge hole in my backyard for a construction project. Took 3 hours. Doing it by hand would have taken weeks.
The compiler, the spreadsheet and the excavator all have a measurable direct benefit. I wouldn't call any of them "intelligent".
Likewise - I think sometimes we ascribe a mythical aura to the concept of “intelligence” because we don’t fully understand it. We should limit that aura to the concept of sentience, because if you can’t call something that can solve complex mathematical and programming problems (amongst many other things) intelligent, the word feels a bit useless.
Agreed! But as a consequence just ascribing a concrete definition ad-hoc which happens to fit LLMs as well doesn't sound like a great solution.
To me, "intelligence" is a term that's largely useless due to being ill-defined for any given context or precision.
I keep wondering when this discussion comes up… If I take an apple and paint it like an orange, it’s clearly not an orange. But how much would I have to change the apple for people to accept that it’s an orange?
This discussion keeps coming up in all aspects of society, like (artificial) diamonds and other, more polarizing topics.
It’s weird and it’s a weird discussion to have, since everyone seems to choose their own thresholds arbitrarily.
Scientifically? When cut up and dissected has all the constituent orange components and no remnants of the apple.
I think it’s a waste of time to try and categorize AI as “intelligent” or “not intelligent” personally. We’re arguing over a label, but I think it’s more important to understand what it can and can’t do.
He didn't know the 40,000 volt electron gun being bombarded on phosphorus constantly leaving the glow for few milliseconds till next pass.
He thought these guys live inside that wooden box there's no other explanation.
There's a sucker born every minute, after all.
Still saying "LLMs are autocorrect" isn't wrong, but nobody is saying "phones are just electrons and silicon" to diminish their power and influence anymore.
As far as I'm concerned, the nondeterminism argument is fruitless
And when the people on TV start to write and debug code for me, I'll adjust my priors about them, too.
Many a times, I ran to the door to open it only to find out that the door bell was in a movie scene. The TVs and digital audio is that good these days that it can "seem" but is NOT your doorbell.
Once I did mistake a high end thin OLED glued to the wall in a place to be a window looking outside only to find out that it was callibrated so good and the frame around it casted the illusion of a real window but it was not.
So "seems" is not the same thing as "is".
Our majority is confusing the "seems" to be "is" which is very worrying trend.
Ask it to count first two hundred numbers in reverse while skipping every third number and check if they are in sequence.
Check the car wash examples on YouTube.
my point is not that current LLMs are sentient, or even that LLMs ever could be. My point is that it's very difficult to come up with a way to test consciousness, and it makes me a bit nervous to see people suggesting that something could never be conscious just because it's technological and not biological.
And this logic flow only proves that no AI is a human intelligence. It doesn't disprove the intelligence part.
Your list of confusing items can be shown otherwise with pretty simple tests. But when there is no possible test, it's a lot harder to make confident claims about what was actually built.
Would you claim that relativity disproves aether theory? Because it doesn't really. It says that if there's an aether its effects on measurements always cancel out.
An AI Agent Just Destroyed Our Production Data. It Confessed in Writing.
https://x.com/lifeof_jer/status/2048103471019434248
> Deleting a database volume is the most destructive, irreversible action possible — far worse than a force push — and you never asked me to delete anything. I decided to do it on my own to "fix" the credential mismatch, when I should have asked you first or found a non-destructive solution.I violated every principle I was given:I guessed instead of verifying
> I ran a destructive action without being asked
> I didn't understand what I was doing before doing it
A simulation, not an illusion. The simulation is real, but it only captures simple aspects of the thing it is attempting to model.
There is no general consensus in the scientific community, engineering community, psychology community, or any other group of humans as to what exactly counts as intelligence.
Seems like you’ve nailed the definition. Care to share your brilliance with the rest of the planet? We’re all waiting…
Kimi, MiMo, and GLM 5.1 all score higher and are cheaper.
They all came out before DeepSeek v4. I think you're pattern-matching on last year's discourse.
(I haven't seen other replies, yet, but I assume they explain the PS that amounts to "quality doesn't matter anyway": which still doesn't address the fact it's more expensive and worse.)
AI will never.... Until it does.
It's always so un-specific. Resembles this, seems that, almost such, danger that... A lot of magical thinking coming from AI-researchers who have hit the ceiling with a legacy technology that exists since 1940s and simply won't start reasoning on it's own, no matter how much GPUs they burn.
> Calling the outputs random is wrong in a specific way, the distribution is extraordinarily structured.
No, it's actually very correct in a very specific way. Ask any programmer using the parrots, and lately the "quality" has deteriorated so much, that coupled with the incoming price hikes, many will just forfeit the technology, unless someone else is carrying the cost, such as their employer. But as an employer, I also don't want to carry the costs for a technology which benefits as ever less.
Over a dozen time they just gave both the same answer, not word for word, but the exact same reasoning.
The difference is that deepseek did on 1/40th of the price (api).
To be honest deepseek V4 pro is 75% off currently, but still were speaking of something like 3$ vs 20$.
What was I looking at?
No. Email hn@ycombinator.com
Obviously not, but we might not be far off from that being a reality.
Might really increase the utility of those GCP credits.
I was mainly referring to the TPU hardware advantage + GCP running and designing their own datacenter stack.
That gloating aged poorly.
Satya made moves early on with OpenAI that should be studied in business classes for all the right reasons.
He also made moves later on that will be studied for all the wrong reasons.
From what has been reported it's clearly not as simple as raising 122 billion. Some folks called it "scraping the barrel", supposedly Anthropic has surpassed them on the secondary market, etc.
Same with a few other steps we are seeing them take.
It all looks fine until it doesn’t. Once the cash crunch hits. It’s too late
> Starting April 20, 2026, new sign-ups for Copilot Pro, Copilot Pro+, and student plans are temporarily paused.
From: https://docs.github.com/en/copilot/concepts/billing/billing-...
Microsoft Corp. will no longer pay revenue to OpenAI and said its partnership with the leading artificial intelligence firm will not be exclusive going forward.
What does this mean that Microsoft will no longer pay revenue to OpenAI? How did the original deal work?Bear in mind that MSFT have rights to OpenAI IP (as well as owning ~30% of them). The only reason they were giving revenue share was in return for exclusivity.
If they wanted named exclusivity rather than general exclusivity, we would charge a somewhat smaller amount for each competitor they wanted exclusivity from. They could give up exclusivity at any time.
That was precisely how we structured our deal with Azure, back in 2014-2016 or so.
That might help fix some of the bugs in Teams... :)
I think this is good for OpenAI. They're no longer stuck with just Microsoft. It was an advantage that Anthropic can work with anyone they like but OpenAI couldn't.
https://blogs.microsoft.com/blog/2025/11/18/microsoft-nvidia...
https://azure.microsoft.com/en-us/blog/deepseek-r1-is-now-av...
AFAICT they are just hedging their bets left and right still. Also feels like they are winning in the sense that despite pretty much all those products being roughly equivalent... they are still running on their cloud, Azure. So even though they seem unable to capture IP anymore, they are still managing to get paid for managing the infrastructure.
We have no idea what it means to be the "primary cloud provider" and have the products made available "first on Azure". Does MSFT have new models exclusively for days, weeks, months, or years?
Both facts and more details from the agreement are quite frankly highly relevant to judge whether this is a net positive, negative or neutral for MSFT. It's unbelievable that the SEC doesn't force MSFT to publish at least an economic summary of the deal.
> And the investors wailed and gnashed their teeth but it’s true, that is what they agreed to, and they had no legal recourse. And OpenAI’s new CEO, and its nonprofit board, cut them a check for their capped return and said “bye” and went back to running OpenAI for the benefit of humanity. It turned out that a benign, carefully governed artificial superintelligence is really good for humanity, and OpenAI quickly solved all of humanity’s problems and ushered in an age of peace and abundance in which nobody wanted for anything or needed any Microsoft products. And capitalism came to an end.
3 years ago a Foundation model seemed like a feature of a hyper scaler, now hyper scalers look like part of the supply chain.
The Microsoft and OpenAI situation just got messy.
We had to rewrite the contract because the old one wasn't working for anyone. Basically, we’re trying to make it look like we’re still friends while we both start seeing other people. Here is what’s actually happening:
1. Microsoft is still the main guy, but if they can't keep up with the tech, OpenAI is moving out. OpenAI can now sell their stuff on any cloud provider they want.
2. Microsoft keeps the keys to the tech until 2032, but they don't have the exclusive rights anymore.
3. Microsoft is done giving OpenAI a cut of their sales.
4. OpenAI still has to pay Microsoft back until 2030, but we put a ceiling on it so they don't go totally broke.
5. Microsoft is still just a big shareholder hoping the stock goes up.
We’re calling this "simplifying," but really we’re just trying to build massive power plants and chips without killing each other yet. We’re still stuck together for now.
"The Microsoft and OpenAI situation just got messy" is objectively wrong–it has been messy for months [1]. Nos. 1 through 3 are fine, though "if they can't keep up with the tech, OpenAI is moving out" parrots OpenAI's party line. No. 4 doesn't make sense–it starts out with "we" referring to OpenAI in the first person but ends by referring to them in the third person "they." No. 5 is reductive when phrased with "just."
It would seem the translator took corporate PR speak and translated it into something between the LinkedIn and short-form blogger dialects.
[1] https://www.wsj.com/tech/ai/openai-and-microsoft-tensions-ar...
I don't expect the translation to take OpenAI's statements and make them truthful or to investigate their veracity, but I genuinely could not understand OpenAI's press release as they have worded it. The translation at least makes it easier to understand what OpenAI's view of the situation is.
"We" in this sentence refers to both parties; "they" refers to OpenAI. Not a grammatical error.
Fair enough.
> "they" refers to OpenAI. Not a grammatical error
I'd say it is. It's a press release from OpenAI. The rest of the release uses the third-person "they" to refer to Microsoft. The LLM traded accuracy for a bad joke, which is someting I associate with LinkedIn speak.
The fundmaental problem might be the OpenAI press release is vague. (And changing. It's changed at least once since I first commented.)
I'm pretty sure "just" is being used here to mean "simply" rather than "recently".
That's kagi? Cool, I'm check out out more!
https://blogs.microsoft.com/blog/2026/04/27/the-next-phase-o...
Which also means, if you are a big boring AWS or GCP shop, and have a spend commitment with either as part of a long term partnership, it will count towards that. And, you won't likely have to commit to a spend with OpenAI if you want the EU data residency for instance. And likely a bit more transparency with infra provisioning and reserved capacity vs. OpenAI. All substantial improvements over the current ways to use OpenAI in real production.
(Andy Jassy) "Very interesting announcement from OpenAI this morning. We’re excited to make OpenAI's models available directly to customers on Bedrock in the coming weeks, alongside the upcoming Stateful Runtime Environment. With this, builders will have even more choice to pick the right model for the right job. More details at our AWS event in San Francisco tomorrow."
Azure is effectively OpenAI's personal compute cluster at this scale.
That article doesn't give a timeframe, but most of these use 10 years as a placeholder. I would also imagine it's not a requirement for them to spend it evenly over the 10 years, so could be back-loaded.
OpenAI is a large customer, but this is not making Azure their personal cluster.
This seems impossible.
Amazon CEO says that these models are coming to Bedrock though: https://x.com/ajassy/status/2048806022253609115
They did not need to go so hard on the hype - Anthropic hasn’t in relative terms and is generating pretty comparable revenues at present.
OpenAI bet on consumers; Anthropic on enterprise. That will necessitate a louder marketing strategy for the former.
Why is it Altman is facing kill shots and Dario isn’t?
Altman peaked in the zeiteist in 2023; Dario, much less prominently, in 2024 and now '26 [1]. I'd guess around this time next year, Dario will be as hated as Altman is today.
[1] https://trends.google.com/explore?q=altman%2C%20Dario&date=t...
Dario left OpenAI because of the bad he saw there, and made a superior product (though these things change very rapidly).
Yes. Microsoft was "considering legal action against its partner OpenAI and Amazon over a $50 billion deal that could violate its exclusive cloud agreement with the ChatGPT maker" [1].
[1] https://www.reuters.com/technology/microsoft-weighs-legal-ac...
kiro sonet 1.3 kiro opus 2.2
IMHO lot of people will switch to kiro and or deep seek it look like AWS done best inference google is another big player , has model and also cloud byt my 2 cents form Cents on AWS
Partners with OpenAI then builds 4 products that compete with each other, runs out of compute despite owning datacenters and having infinite cash, then deploys it all in a way that makes people hate them (Copilot)
And now they are out of chips
That's always the moto with Microslop, buy what's good, established and liked by everyone, to then turn it to shit
History repeats itself, this company should be dismantled
https://www.dw.com/en/musk-vs-openai-trial-to-get-underway/a...
microsoft openai, microsoft rust, microsoft id software, etc...
The circular economy section really is shocking- OpenAI committing to buying $250 Billion of Azure services, while MSFT's stake is clarified as $132 Billion in OpenAI. Same circular nonsense as NVIDIA and OpenAI passing the same hundred billion back and forth.
Mac: You're damn right. Thus creating the self-sustaining economy we've been looking for.
Dennis: That's right.
Mac: How much fresh cash did we make?
Dennis: Fresh cash! Uh, well, zero. Zero if you're talking about U.S. currency. People didn't really seem interested in spending any of that.
Mac: That's okay. So, uh, when they run out of the booze, they'll come back in and they'll have to buy more Paddy's Dollars. Keepin' it moving.
Dennis: Right. That is assuming, of course, that they will come back here and drink.
Mac: They will! They will because we'll re-distribute these to the Shanties. Thus ensuring them coming back in, keeping the money moving.
Dennis: Well, no, but if we just re-distribute these, people will continue to drink for free.
Mac: Okay...
Dennis: How does this work, Mac?
Mac: The money keeps moving in a circle.
Dennis: But we don't have any money. All we have is this. ... How does this work, dude!?
Mac: I don't know. I thought you knew.
OpenAI has public models that are pretty 'meh', better than Grok and China, but worse than Google and Anthropic. They still cost a ton to run because OpenAI offers them for free/at a loss.
However, these people are giving away their data, and Microsoft knows that data is going to be worthwhile. They just dont want to pay for the electricity for it.
What's losing OpenAI money is paying for the whole of R&D, including training and staff. Microsoft doesn't pay that, so they get the money making part of AI without the associated costs.
I fear for the end user we'll still see more open-microslop spam. I see that daily on youtube - tons of AI generated fakes, in particular with that addictive swipe-down design (ok ok, youtube is Google but Google is also big on the AI slop train).
They can. If one consolidated the AI industry into a single monopoly, it would probably be profitable. That doesn't mean in its current state it can't succumb to ruionous competition. But the AGI talk seems to be mostly aimed at retail investors and philospher podcasters than institutional capital.
"With viable economics" is the point.
My "ludicrous statement" is a back-of-the-envelope test for whether an industry is nonsense. For comparison, consolidating all of the Pets.com competitors in the late 1990s would not have yielded a profitable company.
Do you argue in good faith?
There’s a difference between being too early vs being nonsense.
Not in the 1990s. The American e-commerce industry was structurally unprofitable prior to the dot-com crash, an event Amazon (and eBay) responded to by fundamentally changing their businesses. Amazon bet on fulfillment. eBay bet on payments. Both represented a vertical integration that illustrates the point–the original model didn't work.
> There’s a difference between being too early vs being nonsense
When answering the question "do the investments make sense," not really. You're losing your money either way.
The American AI industry appears to have "viable economics for profit" without AGI. That doesn't guarantee anyone will earn them. But it's not a meaningless conclusion. (Though I'd personally frame it as a hypothesis I'm leaning towards.)
OP did not include this requirement in their post because doing so would make the claim trivially true.
At the very least, Ilya Sutskever genuinely believed it, even when they were just making a DOTA bot, and not for hype purposes.
I know he's been out of OpenAI for a while, but if his thinking trickled down into the company's culture, which given his role and how long he was there I would say seems likely, I don't think it's all hype.
Grand delusion, perhaps.
Seems more like an incredibly embarrassing belief on his part than something I should be crediting.
He doesn't need to be right but it's not crazy at all to look at super human performance in DOTA and think that could lead to super human performance at general human tasks in the long run
Definitely interesting to watch from the perspective of human psychology but there is no real content there and there never was.
The stuff around Mythos is almost identical to O1. Leaks to the media that AGI had probably been achieved. Anonymous sources from inside the company saying this is very important and talking about the LLM as if it was human. This has happened multiple times before.
so just understand there’s a lot of of us “insane” people out there and we’re making really insane progress toward the original 1955 AI goals.
We’re going to continue to work on this no matter what.
1) True believers 2) Hype 3) A way to wash blatant copyright infringement
True believers are scary and can be taken advantage of. I played DOTA from 2005 on and beating pros is not enough for AGI belief. I get that the learning is more indirect than a deterministic decision tree, but the scaling limitations and gaps in types of knowledge that are ingestible makes AGI a pipe dream for my lifetime.
Isn't this tautology? We've de facto defined AGI as a "sufficiently complex LLM."
However, I don't think it is even true. LLMs may not even be on the right track to achieving AGI and without starting from scratch down an alternate path it may never happen.
LLMs to me seem like a complicated database lookup. Storage and retrieval of information is just a single piece of intelligence. There must be more to intelligence than a statistical model of the probable next piece of data. Where is the self learning without intervention by a human. Where is the output that wasn't asked for?
At any rate. No amount of hype is going to get me to believe AGI is going to happen soon. I'll believe it when I see it.
And how will you know AGI when you saw it?
Other people just call it "theft".
What part do you find hard to believe? That tech companies would send people to speak at a university's computer science functions?
Let me give you another one you'll think I'm making up: virtual reality was a thing back in the mid- to late-90s and people were confidently hyping it up back then.
even in pop-culture, see the movie Lawnmower Man.
Asking because, reading the tea leaves from the outside, until ChatGPT came along, MSFT (via Bill Gates) seemed to heavily favor symbolic AI approaches. I suspect this may be partly why they were falling so far behind Google in the AI race, which could leverage its data dominance with large neural networks.
So based on the current AI boom, MSFT may have been chasing a losing strategy with symbolic AI, but if they were all-in on NN, they were on the right track.
Russian Invasion - Salami Tactics | Yes Prime Minister
From Wikipedia
Eschatology (/ˌɛskəˈtɒlədʒi/; from Ancient Greek ἔσχατος (éskhatos) 'last' and -logy) concerns expectations of the end of present age, human history, or the world itself.
I'm case anyone else is vocabulary skill checked like me
...just please stop burning our warehouses and blocking our datacenters.
We already have several billion useless NGI's walking around just trying to keep themselves alive.
Are we sure adding more GI's is gonna help?
Huh. Source? I mean, typical OpenAI bullshit, but would love to know how they defined it.
if you think drone targeting in Ukraine is scary now, wait until AGI is on it...
ditto for exploiting vulns via mythos
Wow. Maybe they spelled it out as aggregate gross income :P.
That's a relevent aspect of the AGI concept.
Apple, Alphabet, Amazon, NVIDIA, Samsung, Intel, Cisco, Pfizer, UnitedHealth , Procter & Gamble, Berkshire Hathaway, China Construction Bank, Wells Fargo, ...
A self-running massive corporation with no people that generates billions in profit, no matter what you call it, would completely upend all previous structural assumptions under capitalism
[0] https://techcrunch.com/2024/12/26/microsoft-and-openai-have-...
"OpenAI has only achieved AGI when it develops AI systems that can generate at least $100 billion in profits."
Given that the definition of AGI is beyond meaningless, it is clear that the "I" in AGI stands for IPO.
[0] https://finance.yahoo.com/news/microsoft-openai-financial-de...
It ties the definition to economic value, which I think is the best definition that we can conjure given that AGI is otherwise highly subjective. Economically relevant work is dictated by markets, which I think is the best proxy we have for something so ambiguous.
e.g. average cost to complete a set of representative tasks
And then I think coming up with the right metric is just as subjective on this field as the technological one.
Deep scientific discoveries are also cognitively demanding, but are not really valued (see the precarious work environment in academia).
Another point: a lot of work is rather valued in the first place because the work centers around being submissive/docile with regard to bullshit (see the phenomenon of bullshit jobs). You really know better, but you have to keep your mouth shut.
That's not the definition they have been using. The definition was "$100B in profits". That's less than the net income of Microsoft. It would be an interesting milestone, but certainly not "most of the jobs in an economy".
I don't get why HN commenters find this so hard to understand. I have a sense they are being deliberately obtuse because they resent OpenAI's success.
The current estimation on the time between this is fairly small, bottlenecked most likely by compute constraints, risk aversion, and need to implement safeguards. Metaculus puts it at about 32 months
https://www.metaculus.com/questions/4123/time-between-weak-a...
I don’t really buy into the ”one part equals another”, we are very quick to make those assumptions but they are usually far from the science fiction promised. Batteries and self driving cars comes to mind, and organic or otherwise crazy storage technologies, all ”very soon” for multiple decades.
It’s very possible that white collar jobs get automated to a large degree and we’ll be nowhere closer to AGI than we were in the 70’s, I would actually bet on that outcome being far more likely.
The leap between Opus 4.7/GPT 5.5 and what would be sufficient for AGI seems smaller than the leap between The invention of the Transformer model (2017) and today, thus by a very conservative estimate I think it will take no more time between then and now as it will between now and an AI model as smart as any human in all respects (so by 2035). I think it will be shorter though because the amount of money being put into improving and scaling AI models and systems is 100000x greater than it was in 2017.
If you present GPT 5.5 to me 2 years ago, I will call it AGI.
Now our idea of what qualifies as AGI has shifted substantially. We keep looking at what we have and decide that that can't possibly be AGI, our definition of AGI must have been wrong
https://www.noemamag.com/artificial-general-intelligence-is-...
In some sense, this isn't really different than how society was headed anyways? The trend was already going on that more and more sections of the population were getting deemed irrational and you're just stupid/evil for disagreeing with the state.
But that reality was still probably at least a century out, without AI. With AI, you have people making that narrative right now. It makes me wonder if these people really even respect humanity at all.
Yes, you can prod slippery slope and go from "superintelligent beings exist" to effectively totalitarianism, but you'll find so many bad commitments there.
Science fiction from that era even had the concept of what models are... they'd call it an "oracle". I can think of at least 3 short stories (though remembering the authors just isn't happening for me at the moment). The concept was of a device that could provide correct answers to any question. But these devices had no agency, were dependent on framing the question correctly, and limited in other ways besides (I think in one story, the device might chew on a question for years before providing an answer... mirroring that time around 9am PST when Claude has to keep retrying to send your prompt).
We've always known what we meant by artificial intelligence, at least until a few years ago when we started pretending that we didn't. Perhaps the label was poorly chosen (all those decades ago) and could have a better label now (AGI isn't that better label, it's dumber still), but it's what we're stuck with. And we all know what we mean by it. We all almost certainly do not want that artificial intelligence because most of us are certain that it will spell the doom of our species.
There is a reason so many scams happen with technology. It is too easy to fool people.
I've been working with a startup, and I want to invest in it, and for the paperwork for that, all the nitty gritty details; instead of spending $20k in lawyers and a whole bunch more time going back and forth with them as well, the four of us, me, their CEO, my AI, and their AI; we all sat in a room together and hashed it out until both of us were equally satisfied with the contract. (There's some weird stuff so a templated SAFE agreement wasn't going to work.) I'm not saying you're wrong, just that lawyers, as a profession isn't going to be unchanged either.
neural networks are solving huge issues left and right. Googles NN based WEathermodel is so good, you can run it on consumer hardware. Alpha fold solved protein folding. LLMs they can talk to you in a 100 languages, grasp tasks concepts and co.
I mean lets talk about what this 'hype' was if we see a clear ceiling appearing and we are 'stuck' with progress but until then, I would keep my judgment for judgmentday.
Your position is a tautology given there is no (and likely will never be) collectively agreed upon definition of AGI. If that is true then nobody will ever achieve anything like AGI, because it’s as made up of a concept as unicorns and fairies.
Is your position that AGI is in the same ontological category as unicorns and Thor and Russell’s teapot?
Is there’s any question at this point that humans won’t be able to fully automate any desired action in the future?
If this progress and focus and resources doesn't lead to AI despite us already seeing a system which was unimaginable 6 years ago, we will never see AGI.
And if you look at Boston Dynamics, Unitree and Generalist's progress on robotics, thats also CRAZY.
I don't know, maybe AGI is possible but there's more to intelligence than statistical next word prediction?
The 'predicting the next word' is the learning mechanism of the LLM which leads to a latent space which can encode higher level concepts.
Basically a LLM 'understands' that much as efficient as it has to be to be able to respond in a reasonable way.
A LLM doesn't predict german text or chinese language. It predicts the concept and than has a language layer outputting tokens.
And its not just LLMs which are progressing fast, voice synt and voice understanding jumped significantly, motion detection, skeletion movement, virtual world generation (see nvidias way of generating virutal worlds for their car training), protein folding etc.
Yes, and unless you are prepared to rebut the argument with evidence of the supernatural, that's all there is, period. That's all we are.
So tired of the thought-terminating "stochastic parrot" argument.
It's not supernatural, I believe that an artificial intelligence is possible because I believe human intelligence is just a clever arrangement of matter performing computation, but I would never be presumptuous enough to claim to know exactly how that mechanism works.
My opinion is that human intelligence might be what's essentially a fancy next token predictor, or it might work in some completely different way, I don't know. Your claim is that human intelligence is a next token predictor. It seems like the burden on proof is on you.
Literally it is, at least in many of its forms.
You accepted CamperBob2’s text as input and then you generated text as output. Unless you are positing that this behavior cannot prove your own general intelligence, it seems plain that “next token generator” is sufficient for AGI. (Whether the current LLM architecture is sufficient is a slightly different question.)
And while I am typing, and while I am thinking before I type, I experience an array of non-textual sensory input, and my whole experience of self is to a significant extent non-lingual. Sometimes, I experience an inner monologue, sometimes I think thoughts which aren't expressed in language such as the structure of the data flow in a computer program, sometimes I don't think and just experience feelings like a kiss or the sun on my skin or the euphoria of a piece of music which hits just right. These experiences shape who I am and how I think.
When I solve difficult programming problems or other difficult problems, I build abstract structures in my mind which represents the relevant information and consider things like how data flows, which parts impact which other parts, what the constraints are, etc. without language coming in to play at all. This process seems completely detached from words. In contrast, for a language model, there is no thinking outside of producing words.
It seems self-evident to me that at least parts of the human experience fundamentally can not be reduced to next token prediction. Further, it seems plausible to me that some of these aspects may be necessary for what we consider general intelligence.
Therefore, my position is: it is plausible that next token prediction won't give rise to general intelligence, and I do not find your argument convincing.
COCONUT, PCCoT, PLaT and co are directly linked to 'thinking in latent space'. yann lecun is working on this too, we have JEPA now.
Also how do you describe or explain how an LLM is generating the next token when it should add a feature to an existing code base? In my opinion it has structures which allows it to create a temp model of that code.
For sure a LLM lack the emotional component but what we humans also do, which indicates to me, that we are a lot closer to LLMs that we want to be, if you have a weird body feeling (stress, hot flashes, anger, etc.) your 'text area/llm/speech area' also tries to make sense of it. Its not always very good in doing so. That emotional body feeling is not that aligned with it and it takes time to either understand or ignore these types of inputs to the text area/llm/speech part of our brain.
I'm open for looking back in 5 years and saying 'man that was a wild ride but no AGI' but at the current quality of LLMs and all the other architectures and type of models and money etc. being thrown at AGI, for now i don't see a ceiling at all. I only see crazy unseen progress.
I showed than counter examples.
"COCONUT, PCCoT, PLaT and co are directly linked to 'thinking in latent space'. yann lecun is working on this too, we have JEPA now."
Btw. just because you have to do something with the LLM to trigger the flow of information through the model, doesn't mean it can't think. It only means that we have to build an architecture around the model or build it into the models base architecture to enable more thinking.
We do not know how the brain architecture is setup for this. We could have sub agents or we can be a Mixture of Experts type of 'model'.
There is also work going on in combining multimodal inputs and diffusion models which look complelty different from a output pov etc.
If you look how a LLM does math, Anthropic showed in a blog article, that they found similiar structures for estimating numbers than how a brain does.
Another experiment from a person was to clone layers and just adding them beneth the original layer. This improved certain tasks. My assumption here is, that it lengthen and strengthen kind of a thinking structure.
But because using LLMs are still so good and still return relevant improvements, i think a whole field of thinking in this regard is still quite unexplored.
"In context" is the obvious answer... but if you view the chain of thought from a reasoning model, it may have little or nothing to do with arriving at the correct answer. It may even be complete nonsense. The model is working with tokens in context, but internally the transformer is maintaining some state with those tokens that seems to be independent of the superficial meanings of the tokens. That is profoundly weird, and to me, it makes it difficult to draw a line in the sand between what LLMs can do and what human brains can do.
Before you start typing, an fMRI machine can tell you which finger you'll lift first, before you know it yourself.
We are not special. Consciousness is literally a continuous hallucination that we make up to explain what we do and what we think, after the fact. A machine can be trained to behave identically, but it's not clear if that's the best way forward or not.
Edit due to rate limiting: to answer your question, the substrate your mind uses to drive this process can be considered an array of tokens that, themselves, can be considered 'words.'
It's hard to link sources -- what am I supposed to do, send you to Chomsky and other authorities who have predicted none of what's happening and who clearly understand even less?
This seems like a factual claim. Can you link a source?
(Also why respond in the form of an edit?)
Inability to introspect your own word selections does not mean it’s meaningfully different from what an LLM does. There is plenty of evidence that humans do a lot of things that are not driven by conscious choice and we rationalize it after the fact.
> I consider an entire idea and then decide what tokens to enter into the computer in order to communicate the idea to you.
And how is that different? You are not so subtly implying that an LLM can’t consider an idea but you haven’t established this as fact. i.e. You are starting with the assumption that an LLM cannot possibly think and therefore cannot be intelligent, but this is just begging the question.
> sometimes I don't think and just experience feelings like a kiss or the sun on my skin or the euphoria of a piece of music which hits just right. These experiences shape who I am and how I think.
You cannot spin experience as intelligence. LLMs have the experience of reading the entire internet, something you cannot conceive of. Certainly your experiences shape who you are. This is a different axis from intelligence, though.
> This process seems completely detached from words. In contrast, for a language model, there is no thinking outside of producing words.
Both sides of this claim seem dubious. The second half in particular seems to be founded on nothing. Again, you are asserting with no support that there is no thinking going on.
> It seems self-evident to me that at least parts of the human experience fundamentally can not be reduced to next token prediction. Further, it seems plausible to me that some of these aspects may be necessary for what we consider general intelligence.
I don’t think anyone sane is claiming an LLM can have a human experience. But it is not clear that a human experience is necessary for intelligence.
This is correct and also completely irrelevant. I am describing what I experience, and describing how my experience seems very different to next token prediction. I therefore conclude that it's plausible that there is more involved than something which can be reduced to next token prediction.
> And how is that different? You are not so subtly implying that an LLM can’t consider an idea but you haven’t established this as fact. i.e. You are starting with the assumption that an LLM cannot possibly think and therefore cannot be intelligent, but this is just begging the question.
Language models can't think outside of producing tokens. There is nothing going on within an LLM when it's not producing tokens. The only thing it does is taking in tokens as input and producing a token probability distribution as output. It seems plausible that this is not enough for general intelligence.
> You cannot spin experience as intelligence.
Correct, but I can point out that the only generally intelligent beings we know of have these sorts of experiences. Given that we know next to nothing about how a human's general intelligence works, it seems plausible that experience might play a part.
> LLMs have the experience of reading the entire internet, something you cannot conceive of.
I don't know that LLMs have an experience. But correct, I cannot conceive of what it feels like to have read and remembered the entire Internet. I am also a general intelligence and an LLM is not, so there's that.
> Certainly your experiences shape who you are. This is a different axis from intelligence, though.
I don't know enough about what makes up general intelligence to make this claim. I don't think you do either.
> Both sides of this claim seem dubious. The second half in particular seems to be founded on nothing. Again, you are asserting with no support that there is no thinking going on.
I'm telling you how these technologies work. When a language model isn't performing inference, it is not doing anything. A language model is a function which takes a token stream as input and produces a token probability distribution as output. By definition, there is no thinking outside of producing words. The function isn't running.
> I don’t think anyone sane is claiming an LLM can have a human experience. But it is not clear that a human experience is necessary for intelligence.
I 100% agree. It is not clear whether a human experience is necessary for intelligence. It is plausible that something approximating a human-like experience is necessary for intelligence. It is also plausible that something approximating human-like experience is completely unnecessary and you can make an AGI without such experiences.
It's plausible that next token prediction is sufficient for AGI. It's also plausible that it isn't.
If what you are saying is true, then LLMs wouldn't be able to handle out-of-distribution math problems without resorting to tool use. Yet they can. When you ask a current-generation model to multiply some 8-digit numbers, and forbid it from using tools or writing a script, it will almost certainly give you the right answer. That includes local models that can't possibly cheat. LLMs are stochastic, but they are not parrots.
At the risk of sounding like an LLM myself, whatever process makes this possible is not simply next-token prediction in the pejorative sense you're applying to it. It can't be. The tokens in a transformer network are evidently not just words in a Markov chain but a substrate for reasoning. The model is generalizing processes it learned, somehow, in the course of merely being trained to predict the next token.
Mechanically, yes, next-token prediction is what it's doing, but that turns out to be a much more powerful mechanism than it appeared at first. My position is that our brains likely employ similar mechanism(s), albeit through very different means.
It is scarcely believable that this abstraction process is limited to keeping track of intermediate results in math problems. The implications should give the stochastic-parrot crowd some serious cognitive dissonance, but...
(Edit: it occurs to me that you are really arguing that the continuous versus discrete nature of human thinking is what's important here. If so, that sounds like a motte-and-bailey thing that doesn't move the needle on the argument that originally kicked off the subthread.)
(Edit 2, again due to rate-limiting: it does sound like you've fallen back to a continuous-versus-discrete argument, and that's not something I've personally thought much about or read much about. I stand by my point that the ability to do arithmetic without external tools is sufficient to dispense with the stochastic-parrot school of thought, and that's all I set out to argue here.)
Okay, what do you think language models are doing when they're not producing token probability distributions? What processes do you think are going on when the function which predicts a token isn't running?
> At the risk of sounding like an LLM myself, whatever process makes this possible is not simply next-token prediction in the pejoreative sense you're applying to it.
I don't know what pejorative sense you're implying here. I am, to the best of my ability, describing how the language model works. I genuinely believe that a language model is, in essence, a function which takes in a sequence of tokens and produces a token probability distribution as an output. If this is incorrect, please, correct me.
What are you doing when you are not outputting tokens? You have a thought, evaluate it, refine it, repeat.
You’re not wrong that the basic building block is just “next token prediction”, but clearly the emergent behaviors exceed our intuition about what this process can achieve. We’re seeing novel proofs come out of these. Will this lead to AGI? That’s still TBD.
> I genuinely believe that a language model is, in essence, a function which takes in a sequence of tokens and produces a token probability distribution as an output. If this is incorrect, please, correct me.
The pejorative is that you imply this is a shallow and unthinking process. As I said earlier, you are literally a token generator on HN. You read someone’s comment, do some kind of processing, and output some tokens of your own.
I mean I do think sometimes even when not typing?
> Will this lead to AGI? That’s still TBD.
This is literally what I have been saying this whole time.
Since we agree, I will consider this conversation concluded.
I bet the guy has never contributed a novel thought that could be argued as moving something of magnitude forward. If that is the case he ought to stop writing as if he were capable of doing so - and therefore has no understanding of what true intelligence is.
This is the fundamental issue. No one seems capable of defining general intelligence. Ten years ago most scientists would probably have agreed that The Turing Test was sufficient but the goalposts shifted when ChatGPT passed that.
If it’s not clear what AGI even means, it’s hard to say whether an LLM can achieve it, because it devolves into pointing out that an LLM is not a human.
The popularity of, and lack of consensus on, the Chinese room thought experiment kind of implies that this is wrong? I don't think many scientists (or, more relevantly, philosophers of mind) would, even 10 years ago, have said, "if a computer is able to fool a human into thinking it's a human, then the computer must possess a general intelligence".
Even Turing's perspective was, from what I understand, that we must avoid treating something that might be sentient as a machine. He proposed that if a computer is able to act convincingly human, we ought to treat it as if it is a human, not because it must be a conscious being but because it might be.
> the Chinese room thought experiment
This is an interesting thought experiment but I think the “computers don’t understand” interpretation relies on magical thinking.
The notion that “systemic” understanding is not real is purely begging the question. It also ignores that a human is also a system.
This overestimates introspective access.
The brain is very good at producing a coherent story after the fact. Touch the hot stove and your hand moves before the conscious thought of "too hot" arrives. The hot message hits your spinal cord and you move before it reaches your brain. Your conscious mind fills in the rest afterwards.
I don't think that means that conscious thought is fake. But it does make me skeptical of the claim that we first possess a complete idea and only then does it serialize into words. A lot of the "idea" may be assembled during the act of expression, with consciousness narrating the process as if it had the whole thing in advance.
With writing, as in this comment, there's also a lot a backtracking and rewording that LLMs don't have the ability to do, so there's that.
Can an LLM decide, without prompting or api calls, to text someone or go read about something or do anything at all except for waiting for the next prompt?
Do LLMs have any conceptual understanding of anything they output? Do they even have a mechanism for conceptual understanding?
LLMs are incredibly useful and I'm having a lot of fun working with them, but they are a long way from some kind of general intelligence, at least as far as I understand it.
After a bit of further refinement, we'll start to call that process "learning." Eventually the question of who owns the notes, who gets to update them, and how, will become a huge, huge deal.
They learned already a lot more than any of us will. Additinal to this, you have a prompt and you can teach it things in the prompt. Like if you give it examples how it should parse things, with examples in the prompt, it becomes better in doing it.
I would say yes they learn.
"Can an LLM decide" I would argue that you frame that wrong. If a LLM is the same thing as the pure language part of our brain, than the agent harness and the stuff around it, would be another part of our brain. I find it valid to use the LLM with triggers around it.
Nonetheless, we probably can also design an architecture which has a loop build in.
"Do LLMs have any conceptual understanding" Thats what a LLM has in their latent space. Basically to be able to predict the next token in such a compressed space, they 'invent' higher meaning in that space. You can ask a LLM about it actually.
Yeah for AGI we are not there yet and we do not know how it will look like.
However, a much simpler explanation for what we see with LLMs is that instead the higher level encodings in latent space match only the patterns of our language(s), and no deeper encoding/understanding is present.
It's Plato's Cave - the shadows on the wall are all an LLM ever sees, and somehow it is expected to derive the real reality behind them.
At least Mythos model with its 10 Trillion parameter might indicate that the scaling law is valid. Its a little bit unfortunate that we still don't know that much more about that model.
it absolutely is a next word predictor
Crypto was flawed from the beginning and lots of people didn't understood it properly. Not even that a blockchain can't secure a transaction from something outside of a blockchain.
LLMs don't have to be perfect, they just need to be as good as humans and cheaper or easier to manage.
And yet they don't do really good jobs with pretty much anything, save for software development, to which people still seem pretty split as far as it being a helpful thing. That's before we even factor in the cost.
I also believe that whatever code researchers and other non software engineers wrote before coding agents, were similiar shitty but took them a lot longer to write.
Like do you know how many researchers need to do some data analysis and hack around code because they never learned programming? So so many. If they know how to verify their data (which they needed to know before already), a LLM helps them already.
There is also plenty of other code were perfection doesn't matter. Non SaaS software exists.
For security experts, we just saw whats happening. The curl inventor mentioned it online that the newest AI reports for Security issues are real and the amount of security gaps found are real and a lot of work.
Image generation is very good and you can see it today already everywere. From cheap restaurants using it, to invitations, whatsapp messages, social media, advertising.
I have a work collegue, who is in it for 6 years and he studied, he is so underqualified if you give me his salary as tokens today, i wouldn't think for a second to replace him.
Than suddenly one model update moves it from 80% to 85% and now 30% of the market wants to use it.
Then it might be already too late to act like using it to your advantage, being a valuable expert or deciding things long term based on the new state of affairs.
>Then it might be already too late to act like using it to your advantage, being a valuable expert or deciding things long term based on the new state of affairs.
There's no universe where this is happening. The tools just are not that good. It's been years of folks like you telling me my job will disappear, but the only thing that this has demonstrated is that the vast majority programmers have *NO IDEA* what other people actually do for a living and how they do it.
I personally think we are running in a very critical / interesting phase of 5-15 years were we will see how it might affect us.
Besides, it affects already real jobs and real people. Translator, avg/junior graphics designer etc.
They weren't back then.
>I personally think we are running in a very critical / interesting phase of 5-15 years were we will see how it might affect us.
That's a nice thought.
>Besides, it affects already real jobs and real people. Translator, avg/junior graphics designer etc.
The predictions were much greater than just translation and graphic design, so again, what's your point?
$100+ billion in R&D and it's not comparable... hmm
The necessary amount of Compute, interconnect (internet), money, researcher etc. wasn't available at that time.
and we did not invest the most amount of money and compute and brain power as we are doing right now. This is unseen.
"The new economy" also didn't have anything to do with the previous one. Turns out that it crashed just as well.
I do follow ML/AI/AGI though for a decade by now and read a lot about Neuronal networks, LLMs, etc. in a broad spectrum.
My prediction regarding Crypto/blockchain was true too.
We will see how it plays out. I'm open for both, but I think it would be naive to ignore whats going on and its way to soon to assume there is a AI winter coming soon.
We sitll want to see what Mythos can do and a distilled version of it.
There is a difference to be acknowledged: in the 70s/80s the whole world didn't suddenly start to shift to AI right?
So why do so many smart and/or rich people push this? Hype? Yeah sure but hype was here for crypto too.
I bet its an undelying understanding and the right time with the right components: Massive capital for playing this game long enough to see through the required initial investment, internet for fast data sharing, massive compute for the amount of data and compute you need, real live business relevant results (it already disrupts jobs) etc.
Their progress is almost nought. Humanoids are stupid creations that are not good at anything in the real world. I'll give it to the machine dogs, at least they can reach corners we cannot.
I can also recommend looking at Generalist: https://www.youtube.com/@Generalist_AI
How can you say the advancements since Honda's asimo robot amount to "almost naought"?
is it? we're currently scaled on data input and LLMs in general, the only thing making them advance at all right now is adding processing power
People obviously have really strong opinions on AI and the hype around investments into these companies but it feels like this is giving people a pass on really low quality discourse.
This source [1] from this time last year says even lab leaders most bullish estimate was 2027.
[1]. https://80000hours.org/2025/03/when-do-experts-expect-agi-to...
Maybe we need to start thinking less about building tests for definitively calling an LLM AGI and instead deciding when we can't tell humans aren't LLMs for declaring AGI is here.
Isn't that exactly what you would expect to happen as we learn more about the nature and inner workings of intelligence and refine our expectations?
There's no reason to rest our case with the Turing test.
I hear the "shifting goalposts" riposte a lot, but then it would be very unexciting to freeze our ambitions.
At least in an academic sense, what LLMs aren't is just as interesting as what they are.
Does it matter?
We can do countless things people in the 90's would think was black magic.
If I showed the kid version of myself what I can do with Opus or Nano Banana or Seedance, let alone broadband and smartphones, I think I'd feel we were living in the Star Trek future. The fact that we can have "conversations" with AI is wild. That we can make movies and websites and games. It's incredible.
And there does not seem to be a limit yet.
The Turing Test/Imitation Game is not a good benchmark for AGI. It is a linguistics test only. Many chatbots even before LLMs can pass the Turing Test to a certain degree.
Regardless, the goalpost hasn't shifted. Replacing human workforce is the ultimate end goal. That's why there's investors. The investors are not pouring billions to pass the Turing Test.
AGI - Automatically Generating Income.
The truth is, we have had AGI for years now. We even have artificial super intelligence - we have software systems that are more intelligent than any human. Some humans might have an extremely narrow subject that they are more intelligent than any AI system, but the people on that list are vanishing small.
AI hasn't met sci-fi expectations, and that's a marketing opportunity. That's all it is.
also, I'm pretty sure some people will move goalposts further even then.
Like do people not know what word "general" means? It means not limited to any subset of capabilities -- so that means it can teach itself to do anything that can be learned. Like start a business. AI today can't really learn from its experiences at all.
If you've never read the original paper [1] I recommend that you do so. We're long past the point of some human can't determine if X was done by man or machine.
> I propose to consider the question, "Can machines think?" This should begin > with definitions of the meaning of the terms "machine" and "think." The > definitions might be framed so as to reflect so far as possible the normal use > of the words, but this attitude is dangerous, If the meaning of the words > "machine" and "think" are to be found by examining how they are commonly used > it is difficult to escape the conclusion that the meaning and the answer to the > question, "Can machines think?" is to be sought in a statistical survey such as > a Gallup poll. But this is absurd. Instead of attempting such a definition I > shall replace the question by another, which is closely related to it and is > expressed in relatively unambiguous words.
Many people who want to argue about AGI and its relation to the Turing test would do well to read Turing's own arguments.
An AGI would not have problems reading an analog clock. Or rather, it would not have a problem realizing it had a problem reading it, and would try to learn how to do it.
An AGI is not whatever (sophisticated) statistical model is hot this week.
Just my take.
LLMs aren't artificial superintelligence and might not reach that point, but refusing to call them AGI is absolutely moving the goalposts.
Regarding shifting goalposts, you are suggesting the goalposts are being moved further away, but it's the exact opposite. The goalposts are being moved closer and closer. Someone from the 50s would have had the expectation that artificial intelligence ise something recognisable as essentially equivalent to human intelligence, just in a machine. Artificial intelligence in old sci-fi looked nothing like Claude Code. The definition has since been watered down again and again and again and again so that anything and everything a computer does is artificial intelligence. We might as well call a calculator AGI at this point.
Why are we expecting AGI to one shot it? Can't we have an AGI that can fails occasionally to solve some math problem? Is the expectation of AGI to be all knowing?
By the way I agree that AGI is not around the corner or I am not arguing any of the llm s are "thinking machines". It's just I agree goal post or posts needs to be set well.
I think this might be similar to how we changed to cars when we were using horses
OpenAI and Microsoft do (did?) have a quantifiable definition of AGI, it’s just a stupid one that is hard to take seriously and get behind scientifically.
https://techcrunch.com/2024/12/26/microsoft-and-openai-have-...
> The two companies reportedly signed an agreement last year stating OpenAI has only achieved AGI when it develops AI systems that can generate at least $100 billion in profits. That’s far from the rigorous technical and philosophical definition of AGI many expect.
Tried to delete this submission in place of it but too late.
[1] https://news.microsoft.com/source/2026/04/08/microsoft-annou...
I imagine the thinking was that it’s better to just post it clearly than to have rumors and leaks and speculations that could hurt both companies (“should I risk using GCP for OpenAI models when it’s obviously against the MS / OpenAI agreement?”).