Wafer Scale Compute: Setting Records in Computational Fluid Dynamics
cerebras.net
cerebras.net
Particularly the SpMV CS-1 code listing on page 5, and this comment (in section VI.B) about accuracy:
> The measured normwise relative residuals with mixed and 32-bit are shown in Figure 9. Up to iteration 7 the mixed precision implementation tracks the 32-bit, but then fails to reduce the residual further. With this precision, machine precision is about 10^−3. We have observed this accuracy for very well conditioned systems. Here, the growth of rounding errors during the iterative solve explains the loss of an additional factor of 10, leading to a plateau at a relative residual of 10^−2. We have not yet tested whether this accuracy is acceptable in the MFIX code.
> We expect that for some realistic situations, mixed precision solvers are usable as is; in others they may need to be coupled with a correction scheme such as an iterative refinement or an outer iteration that solves a nonlinear system; and in other situations one may need to use higher precision arithmetic
[MFIX is the NETL CFD described in the article]
It feels a bit odd for the authors to declare the 200x speedup over fp64 runs on Joule, followed by "we don't know if these results at a lower precision are acceptable."
very cool stuff regardless. 18 GB of SRAM on the same (very large) chip is insane.
This is mind-blowing
wow
Also: "the on-wafer interconnect can send out one 32-bit pair of 16-bit words to each of a given core’s four neighbors on every machine cycle and can take in one 32-bit word on the same cycle. [...] supercomputers require a ratio of perhaps one word communicated per 1,000 floating point operations."
I also don't buy that 50% of time is spent setting up the linear equations. They say that that is a conservative estimate; I would say that it's unrealistic - in the CFD code base I work on, it's much much less than that - about 30%, and that's when the linear solver is well parallelised and the equation set up is not.
Why did they choose not to support double precision floats? They are pretty needed for fluids/hpc work.
Looks like RAM with local compute. Another startup is sort of doing this (Arm cores etched in the usual DRAM silicon process [1]). I wonder if this is a long term trend in computing. It does brings a whole new meaning to "getting the code to the data".
Also, the way to program it seems like it is very rough and constraining. I wonder how our current cushy languages and runtimes would (could?) adapt to that architecture. Would OpenMP even fit the job?
Of course, assume we actually know how to simulate neurons in a meaningful way.
It is rectangular, not round. Bummer.
Normally, chips are measured in mm^2.
So that would be 462000000 mm^2 die area, if I did my math right.