Software Pipelining on the Mill CPU [video]
millcomputing.com
millcomputing.com
Clearly the Mill team likes the video format, but are there written expositions for people like me? I have to think I'm not the only one who is video-averse. The very sparse FAQ on the website (http://millcomputing.com/category/faq/general/) says that there are a few white papers in process. Are any of these available yet?
http://millcomputing.com/topic/introduction-to-the-mill-cpu-...
(Although it doesn't cover things in as much depth as the talks)
We are working on a public facing wiki, and do actually have some white papers drafted.
But these talks are not your average talks, and well worth trying even if you don't normally ;)
For me its not that I don't watch videos, its just that either I don't have time to watch a video (whereas reading I can do as fast or slow as I want and I can skim), or it happens that I'm unable to watch a video for whatever reason but would be ok with reading (this happens a lot while I'm travelling - also the most likely time where I would like to read something like this).
If it was purely marketing, they'd have a kickstarter up and be promising performance figures without a billion caveats.
At the end of the day, the problem with the Mill is there's no hardware implemented. No real world benchmarks to be had, and the world of hardware design is full of seemingly slamdunk ideas which have giant problems in actual implementation.
This isn't to disparage the Mill or its team - but the content and results are important. Not the style of the message. Anyone can style a message.
We are a DSP that can run general purpose code.
Traditionally, to run general purpose code fast you needed an out-of-order superscalar architecture, as all the x86 and RISC cores are these days.
DSPs have substantially better performance and substantially better efficiency, but have traditionally been ineffective executing general purpose code (such as the web browser you are using to read this).
The Mill is a synergy of lots of small breakthroughs that together deliver significant improvements to general purpose single threaded code.
Its been held that cores have stopped getting faster. We're faster.
And we have similar as yet not filed improvements for multicore too.
I've seen this Mill occdasionally pitching and always have been asking the same question: are your results from simulation or from FPGA? Now I know the answer.
The most suspicious thing in Mill is that belt thing. To produce operands for N operations you need N*3 (2 for reads, one for write) ports of RAM. For even two operations that means 6 ports. No FPGA allow that out-of-the-box. Given that, you have to implement that in registers and logic, wasting FPGA resources.
(AFAIK, silicon fabs also does not have such RAM blocks. you have to build them themselves, either from registers and logic (and make them slow) or using transistors (which make development process slow). this is THE source of relative slowness of Itanium and Elbrus thing from Russia.)
If you want an advice, go for Tabula. You'll need many R/W ports per block of RAM, they seem to have those (12 ports RAM blocks). Maybe your design won't be as slow as I think it will.
That's also nowhere near the source of slowness in the Itanium.
The reuse of output from functional units was done in TTA CPUs (Transport Triggered Architectures). Guess how they fare if you probable never heard of them?
Guess also how fast or slow they are compared to regular OOO CPUs.
You will be right if you guess that they are not that good in terms of raw performance and they are not that fast in terms of operating frequency.
They are not fast in either way precisely because they use crossbar as a operand delivery network. They also can have FIFOs as the switching network or as an another functional unit dedicated to spreading information, but most often it is not used.
What? This isn't the 1980s, that's not how you register file. Realistically a "register" these days is an abstract concept, a label attached to a value somewhere in the pipeline. It's the job of instruction decode to keep track of exactly where to load it from.
(disclaimer: I'm rebutting the general point about microarchitectures and register files/SRAMs, but haven't studied the Mill in any depth...)
Bypasses are nothing new; what is novel is how we are able to handle triple the number of data paths than other machines, the fact that the bypass is exposed to the program rather than being hidden behind the register metaphor, and that the program model is a single-assignment FIFO. See those talks for more.
There is no better way to speed up registers and SRAM than by having none at all.
They are not new and they can be used to create very efficient chips (in terms of operations/watt) for some fixed functions (precisely, FFT of 2^N).
But they are 1) not fast in terms of raw performance for general purpose tasks, 2) not fast in terms of operating frequency and most important 3) prone to stall when present with non-deterministic delays like access to RAM.
You can add whatever functional units you like to TTA design, including content-addressable memory in disguise as FIFO. TTA design with such device will be identical to what you've described above.
I won't think you will improve performance very much with this trick.
PS
"not fast for general purpose tasks" - in some benchmarks TTA architectures executed gcc 10+ times slower than general purpose CPU with same frequency.
"not fast in operating frequency" - TTA requires crossbar, which is slow in 2D. You cannot make it fast.
"prone to stall" - you have to stop complete pipeline for a cache miss, otherwise you'll have divergence in execution.
Are you being funded by any major silicon giants or are you being backed at all?
How many years do you think before you reach some sort of manufacturing or are you still in the "When it's done" phase?
In heavy semiconductor you don't really move out of "when its done" until the FPGA proof-of-principle is working. That's over a year plus "when it's done" :-)
One crucial question - are you compatible with X86?
> Ivan Godard has designed, implemented or led the teams for 11 compilers for a variety of languages and targets, an operating system, an object-oriented database, and four instruction set architectures. He participated in the revision of Algol68 and is mentioned in its Report, was on the Green team that won the Ada language competition, designed the Mary family of system implementation languages, and was founding editor of the Machine Oriented Languages Bulletin. He is a Member Emeritus of IFIPS Working Group 2.4 (Implementation languages) and was a member of the committee that produced the IEEE and ISO floating-point standard 754-2011.
), its actually designed to be easy to write a compiler for. It's he polar opposite of the "sufficiently smart compiler syndrome" :)
I keep suggesting we do a "sufficiently dumb compiler syndrome" talk, but it'd contain nothing novel; the art is well established by all the VLIW machines that have come before.
However, the scheduler must track lifetimes and make sure that nothing still live falls off the end. This is also (not quite so) easy to do, because the scheduler knows exactly what is the belt behavior of each operation it schedules (the Mill is exposed-pipeline), so it is sufficient to symbolically execute a candidate schedule to know if it is feasible.
If not, the scheduler inserts a spill/fill pair at appropriate places and reschedules. This is guaranteed to terminate, and in practice usually is immediately feasible and rarely takes more than one iteration.
The operation scheduling itself is the standard time-reversed tableau scheduler used in VLIWs, probably 40 years old at this point.
Goodbye mutexes?
It avoids raising exceptions wherever possible, which I like EXTREMELY. This saves space, allows for faster hardware and makes life of systems/compilator programmer easier.
It should be praised for that matter alone.
If you can answer this, what is the bussiness plan for the Mill? Who do you expect to buy it when it comes out. Has any company expressed interest is using it?
http://hackaday.com/2013/11/18/interview-new-mill-cpu-archit...
Hope this helps!
Any chance you have a solution to the cache dilemma? Unified cache's are brilliant for simplifying implementations but can often lead to stalling.
It's about using as much of the die area on the CPU chip for actual computation, rather than supporting an instruction set with outdated design.
The programmer's model of computation today has little correspondence to the realities of current semiconductor process technology in terms of what's fast, and what's easy to implement in hardware. The Mill is a bottom-up redesign that takes into account many of the design constraints with current technology, and attempts to design a good architecture that can maximize actual computational throughput.
Looks like this is Transmeta v2. :-)
The Mill on the other hand can just start with the segment of the market that is the least inconvenienced by having to recompile their code and expand from there. So basically it's easier for the Mill to gain a foothold (build the best cpu for one segment) but harder to climb up from there.
joking aside, seeing this talk makes me nostalgic for college and the long nights spent in the CprE labs. I ended up not in hardware/firmware because 4/5GL are so much more fun to work with and let me be so much more expressive. However, this shit is seriously cool. Can't wait to learn more.
edit: also wanted to say the presentation is a perfect example of awesome UX. The animations are not superfluous as in most presentations and are quite meaningful and help the audience better understand the material. It is also an example of something with great UX, but poor aesthetics. Green/yellow text on a blue background, eck. Could have been worse though :)