We are a DSP that can run general purpose code.
Traditionally, to run general purpose code fast you needed an out-of-order superscalar architecture, as all the x86 and RISC cores are these days.
DSPs have substantially better performance and substantially better efficiency, but have traditionally been ineffective executing general purpose code (such as the web browser you are using to read this).
The Mill is a synergy of lots of small breakthroughs that together deliver significant improvements to general purpose single threaded code.
Its been held that cores have stopped getting faster. We're faster.
And we have similar as yet not filed improvements for multicore too.
I've seen this Mill occdasionally pitching and always have been asking the same question: are your results from simulation or from FPGA? Now I know the answer.
The most suspicious thing in Mill is that belt thing. To produce operands for N operations you need N*3 (2 for reads, one for write) ports of RAM. For even two operations that means 6 ports. No FPGA allow that out-of-the-box. Given that, you have to implement that in registers and logic, wasting FPGA resources.
(AFAIK, silicon fabs also does not have such RAM blocks. you have to build them themselves, either from registers and logic (and make them slow) or using transistors (which make development process slow). this is THE source of relative slowness of Itanium and Elbrus thing from Russia.)
If you want an advice, go for Tabula. You'll need many R/W ports per block of RAM, they seem to have those (12 ports RAM blocks). Maybe your design won't be as slow as I think it will.
That's also nowhere near the source of slowness in the Itanium.
The reuse of output from functional units was done in TTA CPUs (Transport Triggered Architectures). Guess how they fare if you probable never heard of them?
Guess also how fast or slow they are compared to regular OOO CPUs.
You will be right if you guess that they are not that good in terms of raw performance and they are not that fast in terms of operating frequency.
They are not fast in either way precisely because they use crossbar as a operand delivery network. They also can have FIFOs as the switching network or as an another functional unit dedicated to spreading information, but most often it is not used.
What? This isn't the 1980s, that's not how you register file. Realistically a "register" these days is an abstract concept, a label attached to a value somewhere in the pipeline. It's the job of instruction decode to keep track of exactly where to load it from.
(disclaimer: I'm rebutting the general point about microarchitectures and register files/SRAMs, but haven't studied the Mill in any depth...)
Bypasses are nothing new; what is novel is how we are able to handle triple the number of data paths than other machines, the fact that the bypass is exposed to the program rather than being hidden behind the register metaphor, and that the program model is a single-assignment FIFO. See those talks for more.
There is no better way to speed up registers and SRAM than by having none at all.
They are not new and they can be used to create very efficient chips (in terms of operations/watt) for some fixed functions (precisely, FFT of 2^N).
But they are 1) not fast in terms of raw performance for general purpose tasks, 2) not fast in terms of operating frequency and most important 3) prone to stall when present with non-deterministic delays like access to RAM.
You can add whatever functional units you like to TTA design, including content-addressable memory in disguise as FIFO. TTA design with such device will be identical to what you've described above.
I won't think you will improve performance very much with this trick.
PS
"not fast for general purpose tasks" - in some benchmarks TTA architectures executed gcc 10+ times slower than general purpose CPU with same frequency.
"not fast in operating frequency" - TTA requires crossbar, which is slow in 2D. You cannot make it fast.
"prone to stall" - you have to stop complete pipeline for a cache miss, otherwise you'll have divergence in execution.
Are you being funded by any major silicon giants or are you being backed at all?
How many years do you think before you reach some sort of manufacturing or are you still in the "When it's done" phase?
In heavy semiconductor you don't really move out of "when its done" until the FPGA proof-of-principle is working. That's over a year plus "when it's done" :-)
One crucial question - are you compatible with X86?
> Ivan Godard has designed, implemented or led the teams for 11 compilers for a variety of languages and targets, an operating system, an object-oriented database, and four instruction set architectures. He participated in the revision of Algol68 and is mentioned in its Report, was on the Green team that won the Ada language competition, designed the Mary family of system implementation languages, and was founding editor of the Machine Oriented Languages Bulletin. He is a Member Emeritus of IFIPS Working Group 2.4 (Implementation languages) and was a member of the committee that produced the IEEE and ISO floating-point standard 754-2011.
), its actually designed to be easy to write a compiler for. It's he polar opposite of the "sufficiently smart compiler syndrome" :)
I keep suggesting we do a "sufficiently dumb compiler syndrome" talk, but it'd contain nothing novel; the art is well established by all the VLIW machines that have come before.
However, the scheduler must track lifetimes and make sure that nothing still live falls off the end. This is also (not quite so) easy to do, because the scheduler knows exactly what is the belt behavior of each operation it schedules (the Mill is exposed-pipeline), so it is sufficient to symbolically execute a candidate schedule to know if it is feasible.
If not, the scheduler inserts a spill/fill pair at appropriate places and reschedules. This is guaranteed to terminate, and in practice usually is immediately feasible and rarely takes more than one iteration.
The operation scheduling itself is the standard time-reversed tableau scheduler used in VLIWs, probably 40 years old at this point.
Goodbye mutexes?
It avoids raising exceptions wherever possible, which I like EXTREMELY. This saves space, allows for faster hardware and makes life of systems/compilator programmer easier.
It should be praised for that matter alone.
If you can answer this, what is the bussiness plan for the Mill? Who do you expect to buy it when it comes out. Has any company expressed interest is using it?
http://hackaday.com/2013/11/18/interview-new-mill-cpu-archit...
Hope this helps!
Any chance you have a solution to the cache dilemma? Unified cache's are brilliant for simplifying implementations but can often lead to stalling.
It's about using as much of the die area on the CPU chip for actual computation, rather than supporting an instruction set with outdated design.
The programmer's model of computation today has little correspondence to the realities of current semiconductor process technology in terms of what's fast, and what's easy to implement in hardware. The Mill is a bottom-up redesign that takes into account many of the design constraints with current technology, and attempts to design a good architecture that can maximize actual computational throughput.