Maybe that’s all far enough afield to make the current state of things irrelevant?
Maybe that’s all far enough afield to make the current state of things irrelevant?
Currently rendering and local GPGPU compute is Nvidia dominated and I don’t see AMD competently going after the market segments.
Most will probably use something like Llama as base.
Believe it or not, we've actually been grappling with this scenario for almost a decade at this point. Originally the answer was to unite hardware manufacturers around a common featureset that could compete with (albeit not replace) CUDA. Khronos was prepared to elevate OpenCL to an industry standard, but Apple pulled their support for it and let the industry collapse into proprietary competition again. I bet they're kicking themselves over that one, if they still hold a stronger grudge against Nvidia than Khronos at least.
So - logically, there's actually a one-size-fits-all solution for this problem. It was even going to get managed by the same people handling Vulkan. The problem was corporate greed and shortsighted investment that let OpenCL languish while CUDA was under active heavy development.
> BTW, training is such a high cost that it seems like a major motive for the customer to reduce costs there.
Eh, that's kinda like saying "app development is so expensive that consumers will eventually care". Consumers just buy the end product; they are never exposed to building the software or concerned with the cost of the development. This is especially true with businesses like OpenAI that just give you free access to a decent LLM (or Apple and their "it's free for now" mentality).
Or, in the small business case (mind you, “long term” for tech reaching small businesses is looooong), these businesses again need much smaller models because a) they don’t need a model well versed in Shakespeare and multi variable calculus, and b) they want inference to be as low cost as possible.
These are just scenarios off the top of my head. The broader point is that a dramatic drop in training cost is a wildcard whose effects are really hard to predict.
But I don't know what "long term" is exactly, and have no idea how to time this thing. Besides, I'd bet the sibling evoking the Jevon's paradox is correct.
If one of those scenarios happens, maybe Nvidia can pivot, or if we see analog take over, we could see something really bizarre like a dark horse like Seagate taking over by pivoting from SSDs, just because their manufacturing pipeline is more compatible.
How do we get from here to there, cause I want to get there so bad.
However, I think we need AI beyond current LLMs to really take us there. I'm not saying LLMs can't get us there, we don't know, just beyond what we have. We need AI that we can trust with real tasks IRL.
It is kind of an interesting thought though. A big wall of SSD is a fabulous amount of storage. and maybe a clever read only architecture, would be cheaper than SSD. and a clever data structure for shared high order bits, maybe, maybe there is potential for some device to look up matrix multiply results, or close approximations that could be cheaply refined.
Right now, I doubt it. But big static cache it is a kind of interesting idea to kick around Saturday afternoon.
Shard that across the planet and you'd have a global cache for calculations. Or a lookup for every possible AI prompt and its results.
You are reading the GP the wrong way around.
You store partial results exactly because you can't store computation. Computation is perishable¹, you either use it or lose it. And one way to use it is to create partial results you can save for later.
1 - Well, partially so. Hardware utilization is perishable, but computation also consumes inputs (mostly energy) that aren't. How much it perishable depends on the ratio of those two costs, and your mobile phone has a completely different outlook from a supercomputer.