I think one could read too much into the "single-assignment" part of SSA. Sure, functional languages don't tend to rebind variables either, but that's really all SSA has in common with them.
I'm familiar with Appel's "SSA is a functional language" article, but the vibe I got was more than he was trying to build some bridges between the SSA and CPS camps to try to get some cross-pollination going and knowldge sharing between them.
I think he could have just as easily called it "CPS is an imperative language", but my hunch is the FP folks are more persnickety and wouldn't have swallowed that as easily as imperative folks would accept the other title.
CPS is typically used to handle imperative control structures by impure functional programming languages, so it should indeed be of interest to imperative compiler authors.
It seems as though C-like imperative languages and x86 are baked in forever... but the world is actually parallel.
By analogy, of course, I am suggesting that the Church model of computation might have served us better as a foundation than the Turing model has done, since it seems that everything's turning up lambdas as the technology S-curve for silicon computers continues to flatten out.
It's a bit more thrilling to think of it that way, as a matter of scale. When space and time resources are scarce, thinking of the computer as a big piece of RAM that we mutate in-place makes the most sense. As we push upwards and outwards into enough resources and a bigger need for parallelism, it suddenly makes more sense to switch perspectives and reason in terms of binding and substitution.
You can see the same pattern with IO models. We have programming languages descended from 50s-era concepts of computing in which the program drives the machine, but that just isn't true at all. A computer is a passive machine, not an active one: it reacts when you poke it with an interrupt, until it reaches a steady state and settles down again. Operating systems go to a great deal of trouble to simulate the kind of top-down flow control environment our imperative tradition wants to think it is operating in, but in order to get any real performance out of these systems, we all end up building asynchronous, reactive layers on top of the simulated batch-job anyway.
We'd all be better off if we flipped the paradigm, imagining the process of building a program not in terms of writing instructions for the computer to perform, but in terms of creating a structure which will react appropriately in response to whichever events may be brought to its attention.
Instead, we'll probably still be starting young programmers out with for-loops counting from 1 to 10 and printing the result on an imaginary console for decades to come, even though none of that is even remotely relevant to what's actually happening anymore, and once they've crossed that hurdle they'll promptly begin unlearning all that stuff in order to start getting real work done.
On an unrelated note, I spent a while last fall/winter playing around with the idea of a very low level functional language - could we use these techniques in memory-constrained environments too? I didn't find the model I was looking for, but I think it's probably out there, and it would be very interesting to develop a functional/type-safe/immutable language suitable for microcontroller programming, with no garbage collection or implicit allocation. It may be that the Turing model is a better fit for computing in the small, but I suspect that we may find dataflow and functional composition to be useful at all scales. That way of looking at the world has some profound philosophical strength, after all - "you cannot step twice into the same river" and all that.
In contrast, the Church model doesn't really have any notion of time, so it's not clear how to reason about I/O. What evaluation order is used?
In terms of computation, they're of course equivalent, so you can execute the Church model on the Turing-like machines (or more specifically a Von Neumann architecture). So maybe that was the right "choice" (if there ever was one).
There was an entire movement around dataflow processors AND dataflow languages in the 80's, where the instruction sets were not totally ordered (e.g. SISAL was single assignment). But these didn't fare well in the market.
Closer to the ground, there is a pattern of "throwing away information at interfaces" in computing. Interfaces like x86, C, Unix, etc. get solidified by evolutionary forces. The cost of that modularity is global inefficiency.
Another good example of that is the Java code gen architecture. People say "static types in Java code let you generated better code!" Well, no. The Java compiler throws out all the type information, generating Java byte code, which is dynamically typed. Then the JIT takes the byte code and has to re-infer all the type invariants to generate machine code.
My understanding is that the dataflow architectures didn't fare well for technical reasons. The instruction-level parallelism they worked with turned out to be too fine-grained, i.e. the cost of scheduling was greater than the win of parallelism for many programs, especially normal sequential programs that still need to be fast. This led to research in coarse-grained dataflow, but it doesn't seem that much has come of that.
https://www.cs.ucf.edu/~dcm/Teaching/COT4810-Fall%202012/Lit...
Wait another month and a half and I'll be getting around to writing my blog post in support of VLIW, and why previous attempts have failed miserably.
Another problem with Mill is that they have been around for a little over 10 years, and have not released anything publicly (though have given a number of very detailed talks). They have had a lot of trouble getting their design into FPGAs, and even more trouble with having working compilers. When I started REX, I believed in getting it into silicon as fast as possible, as that is the point when people would take us seriously. We closed funding in July of 2015, and are taping out our first silicon next month... While many have called us crazy (and we will be giving real information on why we are not that crazy), I definitely don't want to be called vaporware.
At least for myself and my computer architecture friends/colleagues, most of us just throw Mill into the stack machine box even if it is not entirely true... the benefits of the Belt architecture over the traditional stack machine are there in concept, but Mill has also not released any real concrete information on their compilers (at least from what I have seen).
We'll have actual hardware development/evaluation kits available this fall, so feel free to contact us or sign up for our mailing list to be alerted when we release details on all of that. As I said in a previous post here, I strongly believe that most people do not care about something that they can not physically touch, so I expect that most developers would much rather start to learn about and work on an architecture that actually has silicon available.
It doesn't matter how it's structured, just that it's faster.
SSA is fairly well studied, there are many algorithms to transform an IR into SSA and also lots of algorithms for optimizing programs in SSA form.