Did startup Flow Computing just make CPUs 100x faster?
theverge.com
theverge.com
Historically, the chance of such research turning into a chip you can buy is zero.
Sounds like some of those 1970's miracle filters, that would let you fuel your car with microenergized water and hand-waves, instead of gasoline.
100X on selected workloads is easy to believe, I mean, it be a 100 wide vector unit and the workload they get 100X on could be adding up a big line of numbers, haha.
Exactly my thought, modern FPGAs have >1000 DSP elements; I can make 1000x paralle counters and theoretically “do more work”. It’s also extremely unimpressive. Haha
Not an entirely unconvincing idea, in fairness, but what strikes me as odd is that they’re trying to patent and sell the idea, rather than the technology.
The usual speed bumps with these SMP designs show up in the white paper. There's a section on how they can use recompilation to automatically accelerate existing code. There's a rich history of failed attempts at automatic parallelization at compile time. So this section is just an admission that you're going to have to write code specifically for this thing to get anything out of it.
This coprocessor concept looks like it needs to be tightly coupled with the host CPU to work. The upside needs to be massive to justify a vendor integrating with this IP and also overcome the software adoption costs. Especially since they're directly competing with GPGPU...
Also some strange red flags, like the comparison to quantum computing in the FAQ. (Did they feel the need to include this because an investor asked?)
My guess is they've done some interesting core research around memory bottlenecks/latencies in a massively SMP architecture, and picked up investment to attempt to productize it. But the way they're marketing themselves right now doesn't inspire much confidence.
Is the journalism quality these days so low that people don't understand multitudes vs percentages?
I'll believe it when I see 3rd party tests.
I slightly skimmed the white paper and FAQ, and it really sounds like they're being extremely hand-wavy about making a multi-core CPU.
"The blocks inside the CPU die are optimized and meant for different purposes - vector units for vector calculation, matrix units for matrix calculation - Parallel Processing Unit is optimized for parallel processing."
Parallel processing of what? GPUs are essentially simple CPUs but massively parallel. If you're claiming to be able to parallel processing of a general purpose CPU, then aren't you essentially just making a multi-core CPU? How are you supposedly scaling that up to 100X?
This reeks of being a scam to get money from VCs.
> Any headline that ends in a question mark can be answered by the word no.
As confirmed by the article, of course.
I remain cautiously optimistic. There are often large performance gains left unclaimed for the purpose of "generality". My favourite example is Postgres vs TimescaleDB; by exploiting the structure of certain tables (in the case of TimescaleDB, time-order), we can get better performance -- but of course that only works with time-series data.
Could it be that by focusing on parallel operations in a separate chiplet, these workflows can be made much faster? Maybe someone here has the background to tell me why I should be more pessimistic
This part makes me think particularly they are building off hype - they showed a tech demo that did well for some specific case that happens to be a buzzword.
The following statements:
> A. Nonexistent cache coherence issues. Unlike in current CPU systems, in Flow’s architecture there are no cache coherence issues in the memory systems due to the memory organization excluding caches in the front of the intercommunication network.
And
> E. Low-level parallelism for dependent operations. In Flow-enabled CPUs, it is possible to execute dependent operations with the full utilization within a step (with the help of chaining of functional units), whereas in current CPUs the operations executed in parallel need to be independent due to parallel organization of the functional units.
Raise my eyebrows, more than a little bit. If to compute e.g. (assume for a second you do not have optimized instructions for this) you have E = D * C and C = B + A, how exactly you can compute E without first knowing C? You can't, not without computing all possible values of E for C (inefficient).
Introduce locking into the situation (where cache coherence issues come into play), their "excluding caches in the front of the intercommunication network." makes no sense.
You need locks to make guarantees about memory safety, saying that you have none is fine if you have nothing to lock, but then you're likely not doing general processing, you're likely doing something specific.
Which I think is a clue as to what they are really doing:
> D. Flexible threading/fibering scheme. Flow-computing technology allows an unbounded number of fibers at the model level, which can also be supported in hardware (within certain bandwidth constraints). In current-generation CPUs, the number of threads is - in theory - not bounded, but if the number of hardware threads is exceeded in the case of interdependencies, the results can be very bad. In addition, the operating systems typically limit the number of threads to a few thousand at most. The mapping of fibers to backend units is a programmable function allowing further performance improvements in Flow.
To me, this sounds like someone CPU-ified a GPU. That's the primary purpose of a gpu, efficiently run a shit-ton of threads. Except of course GPUs aren't great at processing. But it fits the use case of AI algorithms well, which stand to gain a lot from general GPU improvements.
I think what the mean (according to their diagram) is each result has to be written back to the register file before another unit can use it. So conventionally you would compute C = A + B and update the C register and in the next step compute E = D*C.
What they seem to claim is they directly bypass the computed C result from the add unit the the multiply unit, hence it is pipelined. This is a bit disingenuous as any performant cpu worth it's salt will have operand bypassing.
Right, but that's an operation optimization. Nobody disputes that particular instructions (combined add/multiply) can be optimized and pipelined, but that's very different from claiming that arbitrary calculations no longer have dependent steps.
Also, the studies on that Wikipedia page disprove the law.
The answer is clearly, no, they didn't make CPUs 100x faster. Maybe they intend to, but that's not the same thing as having done it.
That’s… an interesting way to write an article. Maybe they will write an article about comments to bring it full circle.
Otherwise, they are completely different. The Flow processor is more like a SIMT GPU modified to run general-purpose code. The Mill is more like a DSP modified to run general-purpose code.
Any investors please join the queue.