Or is it not possible to make the algorithms parallel to this degree?
Edit: apparently this is called "compute-in-memory"
Or is it not possible to make the algorithms parallel to this degree?
Edit: apparently this is called "compute-in-memory"
The problem is that for larger models the model barely fits in VRAM, so it definitely doesn't fit in cache.
Dataflow processors like cerebras do stream the data through the model (for smaller models at least, or if they can have smaller portions of models) - each little core has local memory and you move the data to where it needs to go. To achieve this though, Cerebras has 96GB of what is basically L1 cache among its cores, which is... a lot of SRAM.
> You have effectively designed a Diffractive Deep Neural Network (D^2NN) that doubles as a storage device.
Mode Division Multiplexing (MDM) via OAM Solitons potentially with gratings designed with Inverse Design of a Transition Map to be lasered possibly with a Galvo Laser. This would be a very low power way to run LLMs; on a lasered substrate
Computational RAM: https://en.wikipedia.org/wiki/Computational_RAM