The Mill CPU Architecture – Threading [video]
youtube.com
youtube.com
Although the same end result could be accomplished without going through a 4x16 "bit matrix" or "crossbar", the setup has some nice properties, particularly for a hobbyist TTL CPU:
* Generating 16 bit address lines and 8 or 16 bit memory-data IO lines could be done by grabbing 2 or 4 nibbles at a time.
* You could drive LED's or a (decoded) hex 7 segment display directly from the bit-matrix lines to see all N history values.
* If you have at least two mux units, the "A" and "B" inputs of the ALU chip could be pointed at different history values.
* The two mux units could be ganged to feed the wider RAM-address and RAM-data registers.
* When doing math or logic on 8, 16, etc. bit values, you wouldn't have to change the register selectors (mux addresses) to change nibbles: the act of pushing the result moves the next operands into position under the "tape head".
That last point means that a fairly simple TTL circuit could flexibly support 4,8,12,16,...,64 bit ALU ops (provided enough 74HCT595s were connected to provide 2x that many bits total). Just set up the initial data and operation, load a TTL counter with the desired number of nibbles, and let 'er rip at the max speed the 74181 can handle.
I call this "The Suspenders" CPU architecture.
I'd like to make a TTL CPU, but the more I look at it, the more it seems I'll have to make my ALU without the 74181 —I don't like relying on discontinued parts that may be hard or impossible to source in the near future.
Explaining success or failure is always fraught with peril and overly simplified beyond reason, but I can't help but think that Intel massively fucked up one thing and one thing only: they targeted the enterprise market. Not just the commodity data center market, but the enterprise enterprise market, the we-only-buy-IBM hyperconservative and slow to move market. The type that still run COBOL on mainframes (and still call them mainframes) because it works and recompiling isn't even close to an option. Nobody in this market wants a nominal increase in computing power if it comes at the expense of backwards incompatibility.
They couldn't have seen it beforehand because smartphones weren't a thing, but a few years after launch they had their ideal target market right in front of them. Smartphones manufacturers will do anything for an incremental improvement in power efficiency. They'll take that improvement and exploit every ounce of it both upmarket in high end phones and downmarket in android burners for 3rd world countries. And people go through phones like crazy...nobody is running software on their phones that is more than 1-2 years old, and most OSes and apps have been updated at least once in the last 3 months. A recompile and migration to a new architecture for this market isn't even 1% of the hurdle that enterprise software was.
I hope that if the Mill makes it to market, that they get the market right. I'd hate to see another innovation get the shaft because of something as dumb as some marketing decisions.
Itanium was supposed to take over everything, not "enterprise". Amusingly its performance projections were based on 36 hand coded instructions from a representative inner loop in Spec, and management went ahead based on that. Even though they would leapfrog x86 in theory, in practise x86 did a steady march in performance improvements (helped by Intel's fabs). As Itanium got late, rather than cancel the project, they decided x86 was for the masses and Itanium for enterprise.
I really like that they are trying Mill, but suspect Risc V is going to soak up the dollars and attention.
But then I noticed that Mill is not an ISA but a family of ISAs? (Small, Medium, Large or something like that) They are hacking LLVM so that it knows about all of the ISAs.
But doesn't this cause a problem for say JIT compilers? (JVM, every major JavaScript engine) Every single JIT compiler has to know about 4 ISAs? I get that they are similar, but that seems onerous. Debugging tools have to know about them too. Even strace is coupled to the ISA. I think the costs may have been underestimated.
Anything that changes the hardware/software boundary is already risky because you have to change two things at once. But if you're going to do that, I would think it should be a single stable interface?
In other words I think the coupling between their own compiler tech and the hardware is too close. Not everything is a portable C program. There's still people running Fortran, not to mention non-LLVM based compilers like Go.
But you're right that changing the hardware/software boundary is really complicated and I expect that to succeed the Mill will have to spend a lot of time incubating in some sort of high end embedded role before the ecosystem needed for a general purpose application processor is there. So things like network switches, cell towers, robots, etc. The sort of thing where you might already be running Linux on top of some sort of RTOS.
Actually it's way worse, but much better.
Better first: Typically, binary programs targeting the mill don't target a specific machine, but they target a hypothetical "general mill" machine code called GenAsm [0]. GenAsm makes some generous assumptions about the hardware, like infinite belt and presence of machine instructions. Included with each machine is the Specializer [1], a program that takes GenAsm and converts it down to the specialized binary encoding for this specific machine; think of it like a linker. This includes translating infinite belt semantics into finite belts, polyfilling any missing machine instructions with microcode, etc, etc. This process is very fast and the OS can cache the result. JITs can use an API to convert generated GenAsm into runnable machine code which includes running it through the Specializer. The specializer is built along with all the other tooling based on the specification that Ivan mentioned briefly.
Now for worse: because mill machines are specification-driven, there could be many more than just 4 ISAs. There could be more ISAs than they have customers depending all on the needs in each case. But it's no big deal because everything targets GenAsm and the machine code differences will be specialized away.
I'm pretty sure it's the Specification talk [2] that goes into the most detail about this.
[0]: http://millcomputing.com/wiki/GenAsm_(code_representation)
Yeah so this is the point I'm quibbling with. I have no doubt it's technically possible. I'm saying that it will hinder adoption, and they're probably underestimating the diversity of software components that generate native code, and underestimating the cost of modifying all those components.
Something like Xen succeeded because it designed up front to be a trivial modification to kernels -- i.e. paravirtualization.
This sounds like a whole new architectural element. They're not only changing the interface between the CPU and the compiler; they're also changing the relationship between the kernel and the CPU (aside from there being a different ISA.)
It's good to test assumptions, and I wish them luck. But after being somewhat excited about it, I feel it's just too ambitious. I'd love to be proven wrong though.
The mill is a new, novel, ISA, which will require compilers to support it as a target. There's no getting around that. But once the tooling is ready, typical programs written in high level languages like C (i.e. excluding inline assembly and architecture-specific assumptions) will be a compiler flag away from being able to distribute binaries that run on all mill chips.
If it's the specializer you're concerned about, it's intended to be very transparent to the typical user, integrated into the system. Most users are completely unaware that a thing called the linker even exists, this should be similar. By the way, just building the spec (defines belt size, available instructions, etc) generates a fully functioning specializer for that machine. I'm pretty sure it does that today.
It's more analogous to something like CUDA, where the you are completely abstracted away from the actual assembly that's running on the GPU. Or like targeting LLVM bitcode.
x86 has essentially a bidirectional 1:1 mapping between the assembly and the bytes directly executed by the cpu. The mill does not have that.
Other end is that with the instruction set not being binary stable, I'm curious how well the Mill would be useful for something like Singularity https://en.wikipedia.org/wiki/Singularity_%28operating_syste... or a hypothetical WebAssembly OS where userspace programs are an IR for the OS to compile. IIRC the Mill is suppose to have it's own IR for program portability
Binary translation viability is key if they want to support Windows-- see current Windows 10 for ARM
Have they made any indications about production time lines recently?
Whatever the case, I do hope they make it to market but maybe that's just my own morbid curiosity.
With that said, I'd love to see these things come off a fab line someday - there's a lot of potential in the ideas behind the Mill architecture, whether they'll pan out or not is to be seen, but if they fail I'd rather see another Itanic than to never make it to market in the first place.
They did an entire talk on how to translate a switch block into Mill assembly; that talk didn't really introduce any new hardware features, just describe how to use them.
It links to videos with longer explanations.
http://millcomputing.com/wiki/Architecture
EDIT:
But to summarize:
This is an exposed pipeline VLIW design except that the instructions are variable width so it doesn't really fit in the traditional conception of VLIW. There are a bunch of clever tricks for compressing the instruction stream and minimizing fetch bandwidth. Instead of registers recent results go into a static single assignment mechanism called the belt where the last N results are visible. To better handle memory speculation register contents are essentially wrapped in Maybe monads in a way that cleverly gets around the Itanium's problems with memory speculation. To better handle data pressure new pages can be declared in the cache hierarchy full of zeroes and are only assigned backing DRAM when ejected. The caches all work in terms of virtual addresses, translation to physical only happens at the DRAM interface. The stack is managed in hardware with its own dedicated queue, the Spiller, which will start pushing data to L2 on its own as it starts to fill up. Code for the Mill is first written to something called general assembly which assumes every possible instruction, infinite execution width, etc. A single pass specialization step goes through and performs necessary substitutions when an executable at the dynamic linking stage.
Edit: 14 years...