Any questions, ask away, happy to explain or dodge as appropriate ;)
Any questions, ask away, happy to explain or dodge as appropriate ;)
I mean, it's nice to be assured the Mill is a general purpose CPU that handles most existing code really well, but what kind of code would it handle really well, if you know what I mean? And what kind of programming language would be well suited to making good use of its strengths?
Going through the most striking parts of the architecture: - They can pipeline outer loops, even with function calls in them. I expect that will perform well on event loop patterns like you find in non-blocking IO server frameworks. - They have cheap pipelining, vector operations, and (in the high-end models) lots of functional units. All the statistical languages like R and Julia should be able to get something out of that. - If you have a lot of repeated pure loops (ie: map() calls), your compiler could optimize two loops into one to try to saturate the mill's functional units. Using their pick and smear operations, I think you can even get this to work on different-size loops (with less benefit the more their sizes differ, of course). - They have novel memory protection stuff, but that's more OS than programming language.
I'm clearly not part of the Mill team though, so I could be wrong about any or all of those.
ILP is a bit of a misnomer on the Mill as each Mill instruction contains many operations and these execute in 'phases' over the subsequent cycles. In the extreme(ly common) case you have a single-instruction tight loop but the second phase of the first iteration is running as the first phase of the second iteration runs. Mind-bending.
We have more success at speculation than a mainstream machine as we have cheap null and `None` propagation; see http://millcomputing.com/topic/metadata/
And as the Mill is a belt machine, and as the hardware manages the call stack, the units can be in use by the caller even when the PC is in the callee. If you schedule, say, a `multiply` that takes 4 cycles and then next cycle call into a function (which usually takes just a cycle) then the machine can go run the function, however long it takes. When the function returns that `multiply` has long since completed but the spiller is going to use result replay to drop the results on the belt at the right time.
A lot of time is spent in loops, though, and the Mill does really well at those! On the Mill, all levels of loop can be pipelined because each level of loop is given its own call frame, allowing the result replay to work. See http://millcomputing.com/topic/pipelining/
Sorry for linking to just so many of the talks! ;)
( And... tails, so I have to submit this. Sorry! )
Interesting concept though!
Are the Mills creators afraid say Intel will use their ideas and create their own chips excluding them from business?
ARM itself is beaten by RISC-V's initial Rocket prototypes in power and area and the upcomming BOOMs are a good match for ARM Cortex A15+ (only 64 bit). And they only recently started just with OoO chips. And do you see ARM suddenly creating RISC-V chips? or borrowing their ideas?
I think it's safe to say that the future will have a lot of fat binaries.
We can however make generally less memory accesses overall, as we have backless memory
When you do a load you specify when it will retire, and it will get the value in memory at the time it retires. What happens is that the load-store unit goes off and gets the memory as soon as possible, swallowing latency, but snoops on the cache coherency traffic.
http://millcomputing.com/topic/memory/ is a good starting talk on this aspect of the Mill.
For example, the Mill is a belt machine and each stack frame has its own belt and the hardware takes care of spilling in-flights across calls to any depth. Also, the Mill loads are orthogonal to cache management and we offer immunity to aliasing and false sharing.