Freight trains carry a lot, but are really bad at branching. Same with GPUs.
Freight trains carry a lot, but are really bad at branching. Same with GPUs.
GPUs are massively parallel limited purpose CPUs. You're always waiting for the program counter to get to the end, and getting data to/from Ram.
My project (since the 1980s) is to put compute and memory together, so you can either store 16 bits, or do a 4 bit look up with the same data in every cell (which has 4x16 bits total). Having no memory-compute barrier should allow for pushing data through a program, in a physical sense. Like an FPGA, but easier to program.
Imagine such a chip at a few Ghz, delivering one output per clock cycle for a given task.
Therefore I can't yet imagine how to easier onboard people who learned on a "classical CPU". Can you go a bit deeper in what makes your system easier to program?
By making each cell identical you remove most of the floor planning concerns. By latching the data from each cell, you remove timing concerns, and deliver deterministic performance.
This has other benefits like being possible to route around a bad cell, or to build a "wall" around a given section of code.
The emulator will always give correct output baring hardware failure.
I hope to get through tiny Tapeout and have a chip before the end of the year.
As for programming, the code has to compile to a directed acyclic graph of boolean logical operations. I suspect that it would be fairly easy to use as a backend for tinygrad or perhaps LLVM as a stretch goal.
The XC6200 became news once as they[1] used an evolutionary algorithm to create a configuration that can detect a single frequency tone. The resulting configuration was determined that it was impossible that it could work, but it did. Placing the configuration in a different part of the FPGA broke it.
It seemed that algorithm used the analog properties of the FPGA in that specific location to get to a smaller result.
[1] https://www.idi.ntnu.no/emner/tdt22/2011/Thompsonieeeehw.pdf
In the same sense, parallelization (and computing on a GPU) are great when you need to perform the same computation on a huge number of data. When each computation is more independent, it might make more sense to perform it through the CPU.