Keeping financiers of their IPO from understanding what they might be buying seems to swerve into fraud territory. Keeping their investors from knowing/understanding what the problems are also seems fraudulent.
7,510 karma · joined October 4, 2013
Keeping financiers of their IPO from understanding what they might be buying seems to swerve into fraud territory. Keeping their investors from knowing/understanding what the problems are also seems fraudulent.
Have you ever been in poverty? I grew up that way and the then-equivalent of $230 would have meant not skipping Christmas when I was 6. It would have meant a full refrigerator, gas for the car for a couple of months, or so many other things.
Google, Amazon, Oracle, Microsoft, Meta, SpaceX, and tons of other startups are spending trillions on data centers, but everyone is renting them out. They aren't renting them to each other (and paying the extra overhead when they have their own servers).
That leaves basically just OpenAI and Anthropic on the hook to pay for everything BEFORE the GPU half of it depreciates away, but they are busy cutting prices to compete with Chinese models which runs counter to their need to increase prices to fulfill their obligations.
How do you add them? Inside the model has tons of problems. From user context makes a bit more sense, but how do you decide and how do you ensure the context is actually beneficial for the company paying for the ads (advertisers are very sensitive about what content gets subconsciously associated with their brand).
How do you make sure a human sees them? This is a very hard problem even in normal ads and seems even more problematic with chatbots.
How do you attract advertiser dollars? Ad spending is zero-sum. You must somehow convince advertisers that your chatbot is a better ad platform than Youtube, Facebook, etc.
The only AI solution that seems reasonable is something like Google's AI search because the RAG backend ensures Google can control where ads show effectively through deterministic means (based on the websites it pulls). Google also has the ad network, consumer base, and provable humans to drive up ad value (this is without mentioning that Gemini falls behind on coding/math benchmarks, but seems to be better at "normie" interactions).
None of this helps any of the AI startups and running all this extra stuff for the same ad revenue cuts Google's profit margins too.
Even if you eliminate 100% of the people involved in creating software, that's still not enough money (and of course, that's not going to happen because someone has to know what to build).
Everyone is adding agents to enterprise stuff, but an overwhelming majority of the general population now hate AI for most things -- especially the "AI support" these companies are using. I think most people would rather suffer through overseas call centers with absolutely terrible representatives than deal with AI support (studies seem to indicate 80+% prefer a human to AI for support in general).
AI as a search engine is useful, but a very different animal. Proficient users want a way to check the AI like Google's AI search does (though the links sometimes don't agree with the summary), but these run very small models (8b or so) with the RAG backend doing the real work.
How many competitors does this space need? Companies can build their own proprietary RAG search engines, but users would almost always prefer the company make that public data available to Google and just use one well-optimized search engine instead of dozens of bad copies.
What about profitability? Google enshittified their search to increase retention and ad time, but AI search should reduce retention/ad time AND costs a lot more money to run too meaning it should lower their bottom line. If that weren't enough, their RAG system is almost certainly more replaceable by users with alternatives than their traditional search system. This seems like all downside for Google.
Solving a merge conflict for $60/hr is not so bad, but what if it's $600/hr?
Why are high level managed languages still being used?
For that matter, why aren’t LLMs generating everything in wasm or native assembly?
“Slopping the hogs” is probably where the modern negative interpretation comes from of forcing garbage onto someone.
A20 Pro scores around 4725 at 8.9w. Geekerwan's review showed that the iPhone 18 pro could dissipate as much as 6.4w during long gaming tests. Cutting power by 50% likely still keeps around 80% of the clockspeed.
As 9950x scores around 3400-3450 in geekbench, or around 28% slower.
There's a very good chance that the iPhone single thread performance is genuinely faster even during sustained loads.
In the end, everyone lost and there are millions of bikes in landfills.
If you're interested in the bikeshare bubble, Asianometry did a video on it a while ago.
Left to its own devices, that system will result in AI autophagy (model collapse) and iterative degradation.
The only area with serious uncertainty is how humans interact, but we now have research showing humans suffer cognitive issues very quickly using AI (some studies indicate effects happen in as little as 10 minutes) with cognitive surrender being a particularly big issue.
In a lot of systems, the only new data seems to be a few brainstorming sentences (you can read slop as entropy decaying things). The AI slops that into requirements. That slop feeds into an agent which generates a bunch of “reasoning” slop, maybe compacts everything (more slop), and spins up agents that get handed slop. They then open files with who knows how many generations of slop (maybe never even touched by a human) and write out a bunch more slop (it’s ironic that humans get better the more they edit a file, but AI gets worse). That slop gets “tested” by another agent reading all the other slop and maybe all this recurses a few generations.
At the end of this AI equivalent to “the human centipede”, you get a developer who’s handed 10x or maybe even 100x more code than their brain would possible process. They are suffering complete cognitive surrender (not to mention often reaching mental and maybe physical collapse from the workload and stress). They don’t understand the system and rubber stamp it so they can move on to the next 50 PRs of the day.
From start to finish, it’s 100% entropy outside a handful of lines worth of human input.
Many people predicted bugs and even discussed entropy issues before AI coding was popular. The buggy mess timing aligns not only with AI adoption, but happens to each company ramping up as they ramp up AI usage.
This is like seeing Einstein’s predictions happen, but arguing he can’t prove correlation/causation. What evidence would you actually accept that is feasible to study?
Every study I’ve seen correlates the use of AI with large increases in the number of bugs. Look at Amazon dialing back AI after massive outages. Microsoft patch Tuesday releases are bricking computers (they even managed to break notepad somehow). The rash of Facebook bugs also coincided with their move to AI. Leaks from Google have engineers saying AI either doesn’t save any time because it takes so to remote stuff or it causes breakages if they speed up.
These companies can afford to get the best devs. They have access to essentially unlimited token budgets. They have STILL fallen off a cliff in quality.
What more proof could there be that this isn’t sustainable?
Just one collision in space carries an ever-increasing risk of cascading into a major event with long-term effects (there's no a good way to call in a space wrecker to haul billions of tiny bits out of orbit).
Bit manipulation offers up to almost 10% advantage.
Zicond allows branchless code which represents significant speedups.
There’s also serious gains to be had from crypto support.
I’d guess the rest aren’t as important to Python, but those are quite important.
Qualcomm recently bought Ventana and are looking at paying ARM billions (in addition to their current billions) for the privilege of designing their own cores. They sit on the RISC-V consortium and made proposals like Znew.
We could be seeing a high-performance RISC-V release from Qualcomm quite soon (especially if they can simply swap out the ARM decode for RISC-V).
Qualcomm beat ARM in court and reportedly pays 2-3% royalties where ARM was demanding 5-10% royalties.
Qualcomm's ARM license expires around 2028 with an option to extend to 2033 (for some amount of money). If Qualcomm is locked into ARM when renewal comes up, ARM is going to not only name the higher price, but likely charge even more to recover their lost revenue.
Qualcomm is already designing their own cores and ARM is charging them billions for the privilege. Increasing net profits 2-3% for simply doing the thing you are doing seems like a very easy choice. Increasing net profits 5-10% (maybe more) after the price hike seems like a fiduciary responsibility.
This isn't just talk either. Qualcomm already proposed a RISC-V Znew extension (nearly 300 pages of changes to make the ISA more like ARMv8). They bought Ventana (a company already done with their second very wide RISC-V design). They partnered with other companies to make Quintauris for promoting RISC-V too.
There's a non-zero chance we see a RISC-V design from Qualcomm before 2028 and I believe a near 100% certainty of a RISC-V design before 2033.
APX proposes exactly this along with 3-register syntax and some other things.
There are two big issues IMO.
1. It will take at least 15 years before most software ships with this because unlike something like AVX where you typically just rewrite a small part of your code that needs AVX, APX requires a 100% rewrite to take advantage.
2. APX instructions require an additional byte each time you use them compared to current instructions. I think there are still savings to be found, but they won't be as big as it might seem at first.
Lack of serious incentive combined with decades-long rollout seems a recipe for non-adoption at a time when RISC chips offer these features now.
That's a sales tactic -- not a logical position.
RISC-V does a 20-bit LUI (load upper immediate) then a 12-bit addi to the same register for the lower bits. Having access to 40-60 bit immediates makes 32-bit immediates a lot easier (with 64-bit immediates being multi-step, but quite uncommon).
2-register to 3-register also just involves different wiring and costs nothing. I think you'd see 15-bit stick with 2-register. 20-bit would more interesting. You could choose to spend 3 bits on a third register or you could widen 2-register instructions to access the 32 core registers (or something between where you do 3-register, but only on 16 registers). 20-bit also reduces some of the need for very large 15-bit immediates (especially jump which is upward of 10% of the total space on 32-bit designs) which could allow more 15-bit instructions further improving effective density.
Easy access to 40/60-bit instructions mean stuff like vsetvli could simply go away and very useful instructions like FMA4 (instead of FMA3) could be added. Vector masking is another big one. They don't have enough bytes for a full vector mask set resulting in some hacks.
The big question is about jumping and predicting inside packets. You can add 2 bits for what externally looks like 16-bit addressing (where the 2 bits indicate packet position to jump to) or have faster jumps that always hit the beginning of the packet (at the expense of code density due to nops). There might even be a hybrid approach where short jumps can jump within a packed, but long jumps must jump to packet boundaries (which makes sense as most compilers make functions align on cache line boundaries anyway). There is a point for eliminating 20-bit (and all that compression goodness) for 45+15-bit pairs instead) as branches inside packets are immediately calculable.
64-bit instructions with 4 bits indicating instruction formats (60-bit, two 40+20-bit variants, 30+30-bit, 20+20+20-bit, three 30+15+15-bit variants, and 15+15+15+15-bit). Have each larger instruction type be a strict superset of the smaller instructions, but with larger immediates, more registers, and maybe additional instruction formats (eg, for SIMD).
Something like that would be even easier to decode (converting short instructions to long is simply a bit of wiring). Instruction density should increase due to 20-bit instruction type. Having properly-aligned instructions would help with fetching performance. Larger instructions means you can jump 4x further with the same immediate and 16-bit offsets. No need to have some of the V extension workarounds (from not wanting to add 48-bit instructions).
SSE has inconsistencies like SSE4.x vs SSE4a. AVX is an even more mixed bag. There are some 19 AVX-512 extensions and ZERO chips support all of them.
The situation is so bad that AMD and Intel got together to make AVX10 to unify everything. That seemed great, but Intel now has AVX 10.1 and 10.2 in addition to the base set, so there we go again...
x86 is a massive battleground with tons of competing extensions like FMA3 vs FMA4 (why did FMA3 win???) and in cases where one of the competing variants didn't win, we get something like virtualization extensions being completely different between Intel and AMD. There's also the rash of security extensions that have gone through various support and dropped support (not to mention using some of this stuff for market segmentation and further fragmenting the ecosystem).
x86 is anything but consistent if you look into its history (or even it's present).