Can you elaborate on this? Would love to understand how that would play out.
The promise of the engine is the same as Lattner’s LLVM was for compilers: train on the back end of your choice, and run it on the front end of your choice. It is modular.
That breaks the hardware-software link.
https://www.modular.com/engine
Scroll down to the first flow chart. Note how they stuck Nvidia over to the bottom-right. That was intentional. ARM, the cheapest data center CPU hardware and the most prevalent at the edge, is in the middle. They launched with CPU support first, now doing GPUs.
The main dividing line in LLMs and other large models is becoming giant commercial models running on GPU cloud vs open source at the edge or CPU cloud.