The Mill CPU Architecture – The Compiler [video]
youtube.com
youtube.com
Any questions, ask away, happy to explain or dodge as appropriate ;)
We can however make generally less memory accesses overall, as we have backless memory
When you do a load you specify when it will retire, and it will get the value in memory at the time it retires. What happens is that the load-store unit goes off and gets the memory as soon as possible, swallowing latency, but snoops on the cache coherency traffic.
http://millcomputing.com/topic/memory/ is a good starting talk on this aspect of the Mill.
For example, the Mill is a belt machine and each stack frame has its own belt and the hardware takes care of spilling in-flights across calls to any depth. Also, the Mill loads are orthogonal to cache management and we offer immunity to aliasing and false sharing.
ILP is a bit of a misnomer on the Mill as each Mill instruction contains many operations and these execute in 'phases' over the subsequent cycles. In the extreme(ly common) case you have a single-instruction tight loop but the second phase of the first iteration is running as the first phase of the second iteration runs. Mind-bending.
We have more success at speculation than a mainstream machine as we have cheap null and `None` propagation; see http://millcomputing.com/topic/metadata/
And as the Mill is a belt machine, and as the hardware manages the call stack, the units can be in use by the caller even when the PC is in the callee. If you schedule, say, a `multiply` that takes 4 cycles and then next cycle call into a function (which usually takes just a cycle) then the machine can go run the function, however long it takes. When the function returns that `multiply` has long since completed but the spiller is going to use result replay to drop the results on the belt at the right time.
A lot of time is spent in loops, though, and the Mill does really well at those! On the Mill, all levels of loop can be pipelined because each level of loop is given its own call frame, allowing the result replay to work. See http://millcomputing.com/topic/pipelining/
Sorry for linking to just so many of the talks! ;)
( And... tails, so I have to submit this. Sorry! )
Interesting concept though!
Are the Mills creators afraid say Intel will use their ideas and create their own chips excluding them from business?
ARM itself is beaten by RISC-V's initial Rocket prototypes in power and area and the upcomming BOOMs are a good match for ARM Cortex A15+ (only 64 bit). And they only recently started just with OoO chips. And do you see ARM suddenly creating RISC-V chips? or borrowing their ideas?
I mean, it's nice to be assured the Mill is a general purpose CPU that handles most existing code really well, but what kind of code would it handle really well, if you know what I mean? And what kind of programming language would be well suited to making good use of its strengths?
Going through the most striking parts of the architecture: - They can pipeline outer loops, even with function calls in them. I expect that will perform well on event loop patterns like you find in non-blocking IO server frameworks. - They have cheap pipelining, vector operations, and (in the high-end models) lots of functional units. All the statistical languages like R and Julia should be able to get something out of that. - If you have a lot of repeated pure loops (ie: map() calls), your compiler could optimize two loops into one to try to saturate the mill's functional units. Using their pick and smear operations, I think you can even get this to work on different-size loops (with less benefit the more their sizes differ, of course). - They have novel memory protection stuff, but that's more OS than programming language.
I'm clearly not part of the Mill team though, so I could be wrong about any or all of those.
I think it's safe to say that the future will have a lot of fat binaries.
moreover, since it is running kind of a cycle ahead, it seems that this whole network, seems more like a cross-bar switch, for renames :)
On a conventional register-based OoO interrupts require issue replay and this is a major source of complexity and budget.
On the Mill, which is a "belt machine", interrupts are as far as the hardware is concerned just an involuntary function call; the Mill does result replay.
Regards mispredicts the Mill has a shockingly short pipeline and mispredicts (where the target is in cache) is only around 5 cycles or so. The Mill has a novel prediction mechanism called transfer prediction as covered in this talk http://millcomputing.com/topic/prediction/
My off-the-cuff thought is that since the spiller has a mechanism for dumping its state in memory anyway, there is probably (or should be) a way for software to force it to dump this state, which could be used by the debugger to read all the required data from a specified (probably model-specific) structure in memory.
The OS can of course control who can debug what.
I wonder if they will use automatic synthesis of code for each instruction, and if so, how they are going to ensure good performance. Maybe superoptimization?
Another question I have is about their stance on free firmware. Will the mill family members require blobs?
[1] 13:50 in the linked video.
They just got their patents up last year, demanding working hardware seems unreasonable and they have working machines simulated in software.