I seek all kinds of answers, including ones about fundamental logic, mathematical physics, etc
I seek all kinds of answers, including ones about fundamental logic, mathematical physics, etc
But once we understand a little bit about the problem we can model 80-90% of its behavior with a handful of parameters. Add in some bias and noise parameters and you have an accurate trainable machine learning model that’s orders of magnitude more efficient.
Take for example a spring which can be modeled by 1 or 2 parameters. But its impulse response looks like a sin curve multiple by exponential decay.
If you just train neurons to match input/outputs from a spring you need a ridiculous number of model parameters to describe that shape.
CNNS have seen an enormous amount of success due to this fact: a lot of processes can be modeled by convolution.
Nvidia is ceding the low-end GPU market to anyone who wants it. Not only does it allow a competitor to establish a reliable source of revenue for their R&D department, but it could cut off the sale of the binned chips that are inevitably produced on the expensive, tiny processes that Nvidia uses - which would hurt their margins to some degree.
bro there is like $10 of margin in your idea for a $200 gpu lol, nobody is "ceding" anything (actually 4060 is a more advanced card than 7600 on literally every front, for ~10% more money) but the cost floor has climbed to the point where $200-300 gpus just don't progress that much anymore.
There's very good reasons for this - shrinks are the least effective on low-tier cards (because memory controllers don't shrink), and you simply don't gain much actual savings from shrinking a 200mm2 die - congrats it's 150mm2 now, on a more expensive node, meaning your $10 chip is now $9. And meanwhile gamers want more VRAM every year, manufacturing and testing and shipping costs have gone up (and cost the same for a 4090 as a 4060), etc. The economics of low-end cards is literally terrible and they are simply falling off the edge of profitability.
Intel is willing to lose money hand-over-fist just to get into the market, but AMD and NVIDIA are pretty much charging fair-ish prices, and gamers just are too emotionally immature to accept that moore's law really really actually is dead for realsies and things aren't going to progress 40% perf/$ per gen anymore.
It's so weird, nobody cries about the CPU market like this. A 1600AF went for $85, a 3600 went for $160, nobody said "boo" when the 5600X increased that to $330 or whatever. Nowadays you are spending at least 50% more on your CPU than you did 5 years ago, sometimes closer to 2x. The enthusiast market is buying $250-400 cpus now, not $85-160. And obviously everyone understands that upgrading your CPU every gen is terrible value too, especially when prices have drifted upwards. But they don't have a half-decade of negging from reviewers telling them that this is a market in crisis, and that they should feel bad about buying a CPU, etc.
The literal half-decade of warfare from reviewers against the GPU market is so tired at this point. Bro, things are going to slow down, it just is how it is. GPUs are the processor that's most dependent on moore's law providing growth in transistors at the same cost, and wafer price increases hit them the hardest. Go complain to TSMC instead, or ASML, or the brick wall - it's ultimately a physics problem. But there's a hell of a lot of clicks and youtube ad money to be made whining about it in the meantime.
At least reviewers are finally coming to jesus on DLSS - mostly because they know AMD will finally have a decent upscaler within a year tops, and that RDNA4/5 will be pushing forward on tensor etc. The writing was on the wall as soon the specs leaked for PS5 Pro, which is basically adopting RTX features wholesale. https://www.youtube.com/watch?v=CbJYtixMUgI https://www.youtube.com/watch?v=BG-7vyw2YRg&t=1625s
Nvidia GPUs are pretty flexible in terms of computation and extremely power hungry. It may be that a next generation of more specialized hardware, such as TPUs or something, outperforms Nvidia GPUs on machine learning tasks to such an extent that those GPUs are obsolete for those tasks. This next generation could come to market sooner than Nvidia anticipates.
Another possibility is that ML researchers figure out some ways to radically reduce the amount of compute required for good training and inference on _less_ specialized hardware. It's really impressive what you can do with llama.cpp. If open source models running on consumer grade hardware ever get to 90% as good as ChatGPT (which, to be clear, is absolutely not the case currently), then those top end GPUs are overkill for most use cases.
I don't think either of those scenarios is particularly likely, but they're at least plausible.
just like the creation of radically simpler internal combustion engines led to us spending a lot less on internal combustion engines, right? /s
The most impressive thing about modern computing is that we've had exponential increase in compute speed, yet everything runs as slow as it did 30 years ago
Also bearish on programmers keeping up on their fundamental algorithms, rather than trying to throw NNs at every problem.
Besides interviews, at least 90% of developers have no need for 'fundamental algorithms'. The library being used uses them in some way sure, but the vast majority of devs simply need to know how to use the tool, not how the tool itself is developed.
Thinking back to a recent-ish discussion here about the (very elegant and efficient) algorithm behind Shazam, and some of the comments here being along the lines of "haha, that's so quaint, nowadays you could just use a neural net". Nevermind how that would even work as well.
Or programmers excited about neural networks forgoing the much simpler, and in many cases completely sufficient, computer vision algorithms built into OpenCV, in favor of trying to train their own model from scratch.
And the other smaller problem was using a hash array instead of a hash map. You don't really need a mental model of multiple algorithms to know o(1) is faster than o(n).
The whole thing is more of a sign of how poor coding practices must have been at Rockstar for such a bug to not only make it in, but persist for years.
1) AGI won't happen because we are on the wrong path
2) AI being a big part of our lives is still a theory. Aswath Damodaran has some brief thoughts on this.
But the biggest bear case has to be that the technology won't get better. Essentially, everyone assumes that it will without reservations.
"That AI isn't all that that and won't make much money" seems to be by far the biggest one. So far the applications are impressive and a little scary, but not actually something that anyone is going to pay for. Apple makes a zillion dollars because people want its phones. Google makes a zillion dollars because people want to sell junk to folks on the internet.
You need to posit a product built out of compute that does more. Maybe replaces a bunch of existing workers in an existing industry, something like that. So far the market is still looking.
Although, even if that is the ultimate result, there's still going to be plenty of money changing hands on the way to that conclusion. See also all the blockchain/web3 companies, when there was (to me at least) clearly a lot less substance/potential there.
For example, being able to feed a potential customer's invoice into GPT and ask it to see what kind of services we can offer to beat its price. Our sales people had to spend hours doing this before. Now it's done in 2 minutes through an engineered prompt. And it's incredibly accurate.
The problem with GPT4 API is context size and price. That's it. Both are bottlenecked by faster and cheaper compute.
That's my bull case for more compute, not bear case like OP asked.
But fundamentally technology gets better when we can do more with less.
When it has to draw that much current it limits the contexts in which it can be beneficial.
Nvidia went all-in on infiniband serDES while AMD chose pcie/CXL. But since Pcie signaling requirements are tighter, you need bigger stronger PHYs, which means you get less actual area per beachfront. The penalty is latency/power, but who cares when gpus are latency-hiding machines anyway?
https://www.semianalysis.com/p/cxl-is-dead-in-the-ai-era
https://www.semianalysis.com/nvidia-b100-b200-gb200-cogs-pri...
this in turn means that nvidia can implement more links or bigger links in their nvswitch networks, which means they can construct bigger systems and push the TCO down.
Two 7900X is still functionally a 7900X, but two 3090s is functionally a 48GB card. Nvidia has got the interconnect bandwidth to a point where it’s a significant enough fraction of the local bandwidth to be functionally one single gpu - this is the same argument as MI300X etc. Doesn’t matter whether the link is on-package or off-package, what matters is that it’s a significant fraction of the speed of your local memory or cache ports. Nvidia did that, with large numbers of gpus, not just a pair of chiplets.
Nvidia has been thinking about this one for a long time - nvswitch is on its third generation, and can switch literal terabytes of data per switch, times several switches. The Mellanox purchase too, but it goes back way longer.
And unlike AMD they actually have a driver that works and just trivially exposes these capabilities and gets out of the way. If you want to tinker and build the open alternative that’s fine, other people want to work.
This is shocking to many AMD fanboys but actually Jensen is a good engineer too, nvidia is mostly on top because they sell products that people want (to such a relentless degree they get furious if they don’t get faster every year etc) and cannot be trivially displaced by “just as good” Radeon drivers etc - just see the latest installment of the geohot saga. Nobody is trapped by nvidia, it is a golden cage - getting actual work done or just going and playing a game instead of spending hours playing with regedit hacks to disable dxnavi to fix DX11 shader compilation stutter is what you’re buying.
https://twitter.com/__tinygrad__/status/1770160392389771305
https://old.reddit.com/search/?q=Dxnavi+stutter+&include_ove...
Nvidia is on top because of relentlessly competent engineering and savant-level business direction, and as much as people scoff at the idea… that’s literally the reason you hate him lol. He is a Jobs-like visionary figure that can see what the tech can be and drive the engineering and business factors to align along the long-term to get him where he wants to go, while also providing the funding and profit in the short term.
https://m.youtube.com/watch?v=Xn1EsFe7snQ&t=1034
The only company with comparable parasocial negative attachment is apple and it’s for the exact same underlying reason . People are also systematically unable to understand that apple users are not “trapped” or in need of rescuing either. People buy apple because it does what they want it to really well, and they don’t care about installing Linux on their phones. And nerds resent that deeply. It’s not a coincidence there’s this axis of warfare around both Nvidia and the App Store with the EU etc. Nerds cannot abide someone choosing the “wrong” hardware. They are right and you will buy the same thing as them or they will get the EU to outlaw your product, or change the symbol licensing to prevent you running on Linux, etc. If you don't like the same filesystem as me, obviously that means I get to relicense some symbols that have been there for 20+ years and break your filesystem. Btrfs is better, the council has spoken.
It keeps happening for a reason, folks, lol. Nerds can’t tolerate others making different choices. And those users disproportionately self-select to “nerd” platforms like android and AMD.
There's still the secondary question of how compute heavy it will be, and I don't think anyone knows. But Sam Altman, in a recent speech he gave in Korea, expressed confidence that there isn't a limit in sight for returns from GPT scaling.
a) Humans at some point between here and eternity become more efficient easier, than scaling compute is hard. Seems unlikely.
b) Compute is overrated, now or will be in the near future. I will be happy to donate to the church of "compute is overrated" if that makes people get off of gpt-4+ and let me cook. Read that as: I doubt it.
I don't see a c)
Just with current AI models, the amount of value that is waiting to be created (take technology X, add AI to it) is incredible. Casual things that used to take years for a team to build, can now be solved by throwing a GPU at it with a generic model that is fine-tuned a bit. Basically, things that were unpractical 2y ago are now on the table.
The bear thesis is that compute will stop being scarce. Which is plausible, since in capitalism, the best cure for high prices tends to be high prices.
Something crazy to think about is that Accelerando by Charles Stross is starting to look like a prophecy being slowly fulfilled.