Stealthy startup Soft Machines launches virtual CPU cores
pcworld.com
pcworld.com
What seems pretty clear is that they use a proprietary ISA and have developed dynamic binary translators for ARM and x86 code. It's not clear if it's a VLIW architecture or if the ISA has any other properties required by the hardware.
From this report it sounds like they're essentially doing thread-level speculation in hardware.
Based on the linked article I would have been tempted to think they have reconfigurable pipelines, but based on the report I'm somewhat sure it was just a misunderstanding.
I generally think it's a bad sign when a company implements relatively well known concepts from research and then doesn't use the standard terminology and tries to pitch their creation as something 100% new and original. In any case, there's no need to rush a judgment, I guess they'll publish better technical information in time.
[0] http://www.softmachines.com/wp-content/uploads/2014/10/MPR-1...
Last I heard, Ebcioglu's execution timing simulations were getting 9:1 speedup on IBM's 370 code via 24-way VLIW.
My favorite old idea was to find and offer some programming language constructs that, in their implementation, could make good use of multiple threads without the programmer having to consider multiple threads.
Just to expand a bit on your apropos points: Michael Flynn is the originator of SISD/SIMD/MISD/MIMD computer architecture classifications. He also originated the notion Directly Executable Languages:
A Directly Executed Language (DEL) is the interface between the
output of a higher level computer language translator and the input
to the interpretive process of a computer. New developments in
computer technology (especially fast read-write control storage)
allow 'soft' computer architectures in which a range of flexible
'machine' or DEL's can be introduced. These DEL's are in many
respects unlike conventional instruction sets, especially in format
flexibility and specified operations.
Ref: http://books.google.com/books/about/Directly_Executed_Langua...(As an undergrad I was fascinated by the work, which was introduced to me by Bob Wedig and Gus Uht at CMU ECE, as well as their extensions for automated concurrency detection and the unit which kept track of which instructions had and had not executed: the Advanced Execution Matrix.)
The reason for failure is that VLIW is basically stuck with whatever the compiler decides can be done. OoO does it dynamically. In the end the only real drawback with OoO vs VLIW is the chip area and increased power consumption required for the logic.
The 9:1 speed up I reported was directly from Ebcioglu: For a while he was in our AI group at Watson. I don't recall just what he did, but he may have had an intermediate step that, for a given program, analyzed the stream of 370 instructions and rewrote them for the 24 way VLIW.
An idea I had him consider was very long addresses. Or, who the heck really wants the addresses in main memory to be 0, 1, 2, ...? Of course, no one! That's why we have lots of work, from old link edit relocation to collection classes, memory management (garbage collection), etc.
So, what do we actually do? Sure, darned near everything comes out of level 1, 2, 3 or so cache which works with just a hash of the main memory address.
So, since we are going to hash the main memory addresses anyway, let's just do that! So, have a very long address, say, 1024 bytes long. So, the 1024 bytes are immediate from the source code! That is, just concatenate, say, address space name, program name, function name, class name, instance name, member name, key name (as in key-value pairs). That's the address. Then hash it.
Ebcioglu checked how fast the execution logic could be, and it seemed okay.
I know; I know; I left out a lot and need a better explanation! I'm concentrating on my software for my project and have gotten away from such low level hardware issues; heck, I don't even remember if we hash the real or the virtual address!
But, the OP was a little light on a lot of history maybe relevant to the work of Soft Machines.
1. http://www.infoq.com/presentations/click-crash-course-modern...
"Industry skepticism is likely. The notion of abstracting software from chip hardware has been tried by companies such as Transmeta, a startup born in the mid-1990s that labored for years in secret on technology based on translating computing instructions in novel ways. The startup ultimately failed."
The thing I'm wondering the most is that how on earth can they get so many instructions per clock with short pipeline? Without knowing the details on how they compiled the SPEC benchmark it's really hard to say. Who knows, maybe they cheated and ran parts of the benchmark parallel on their cpu and not on others, with "It's the natural way for this chip!" as an excuse.
> Without knowing the details on how they compiled the SPEC benchmark it's really hard to say. Who knows, maybe they cheated and ran parts of the benchmark parallel on their cpu and not on others
That's most likely the case, but I wouldn't consider it cheating. As long as from the software perspective only a single thread is running (and SPEC CPU 2000 and 2006 are single threaded), I think it's fair game. The whole point of their project is to expose parallelism without requiring the programmer / execution environment to explicitly support it.
If the software is compiled into a normal single threaded program then what else there is left except instruction level parallelism?
And if you can compile it to work with two threads then we have Hyperthreading to take advantage of that even with a single core.
Their [Soft Machines] latest patent is about basically an OoO method in overdrive, it abuses only instruction level parallelism. And based on that their claim that their pipeline would be short is not really valid.
Trying to launch line of processors, even without any of the translation magic, seems like a very difficult venture all by itself.
EDIT: Jackpot! Google patent search to the rescue. Based on their patents in the last few year (latest one was published March 2014) one can get an understanding what the fuss is all about. Have to read those trough today.
Here is a better article:
http://www.pcworld.com/article/2838018/stealthy-startup-soft...