The Itanium processor, part 1: Warming up
blogs.msdn.com
blogs.msdn.com
http://www.intel.com/content/dam/www/public/us/en/documents/...
These are rarely discussed. However, Secure64's SourceT OS got plenty of mileage out of them. You might not be able to argue for Itanium on cost, performance, or ease of use. However, one could argue for it as a better start on secure OS's or appliances. Unlike academic prototypes, it's also in production with high speed and reliability. Let's also not forget you can do reverse stacks and other bug prevention tricks with less performance hit or clunkiness on a RISC architecture vs x86.
Personally, I'd just rather them have modernized the i960MX, minus other BiiN stuff:
https://en.wikipedia.org/wiki/BiiN
Targeting a robust, UNIX-compatible OS and C toolchain to it might have let it survive and get continually updated. Then, when HLL's got popular, we'd have a good hardware target for them that supported POLA inside applications. As usual, I'm speaking of great stuff in a past or never-happened tense. Least Itanium made it so far.
Lmao well-put. Maybe if they saw it worded that way on the drawing board they might have seen the folly. I keep thinking of checking on the Alpha ISA licensing situation to get PALcode, etc back. Last I checked Intel and Samsung had control of licensing but it kinda disappeared.
Might just stick with SPARC or RISC-V with modifications to use clever modifications from the past. Might even try to design a knock-off of i960's better features updates based on lessons learned over time.
I posted a link to Schneiers Squid thread today on a bunch of asynchronous chip work. One was a 180nm FPGA w/ several times Xilinx's performance and one a 40nm microcontroller. I imagine a combination of RISC-V with that async flow would produce one drool-worthy processor in price, performance, and NRE cost.
I confess to not being particularly conversant with async design (it became mainstream after my time), but I'll definitely check out your links and see what I can learn.
Interesting times and all that.
As hardware goes it's .. interesting. It's fast. But it uses huge amounts of power and requires massive cooling (if you disable any of the 4 or 5 fans in my 2U machine, it overheats in 5 minutes). It has early EFI which should be quite familiar if you've used UEFI on the command line. And it has excellent iLO / remote / serial support so it's great practice for learning about enterprise ops.
It's getting hard to find software that runs on it. The last Debian (Wheezy) runs, but current Debian has dropped ia64 support. RHEL dropped support years ago. You'll find there are lots of strange bugs because people no longer test their software on this platform.
https://rwmj.wordpress.com/2015/05/03/raise-the-itanic-part-...
https://rwmj.wordpress.com/2014/09/08/raise-the-itanic/#cont...
Huh, I would have expected better redundancy given how expensive the hardware must have been at the time.
Edit: Would love to know what the list price of my machine was back in 2006. Probably thousands ...
Try 15K euros!
"Time of introduction: 2002-2003 (rx2600)/December 2004 (rx2620) with prices at the time starting at $7,300 (entry rx2600), $16,000 (average rx2600) to $33,000 (large rx2600)."
My last gig had an Itanium running some HP-UX software that had been cobbled together over the past 20 (?) years. eBay provides cheap parts to upgrade the server as well as put together a machine for the failover datacenter.
CPU and RAM were both cheap. However, you're stuck with commodity SCSI drives which are still expensive for the largest sizes. 300GB drives, NIB and matching are like $150 USD a piece.
Huge power hogs. Even more than my ES40.
The iLO stuff is phenomenal. As is all the hot-swap support ... DIMMs, CPUs, fans, you name it. This goes back to 2003, too.
We had one of HP's first generation Itanium servers to test on, and man, was that SLOW.
setjmp()/longjmp() are not guaranteed to work with local variables not marked volatile.
Microsoft on RSE (2002) https://web.archive.org/web/20021018050724/http://portals.de...
Smotherman's notes (2002) http://people.cs.clemson.edu/~mark/subroutines/itanium.html
USENIX presentation on Itanium (2005) https://www.usenix.org/legacy/events/usenix05/tech/general/g...
I could see how you wouldn't expect it coming from another ISA and they could be a bit more explicit. It was weird. It was documented, though, by different people building on Itanium. Different workloads use it in a different way for efficiency.
"This also has a side-effect on function pointers. Since function pointers are generally used at some distance from allocation, they might be used in a module with a different gp value. The compiler gets around this by not compiling a function pointer to a single pointer-sized value; it compiles to a pair of pointer-sized values, one representing the address of the first instruction (bundle on IA64) in the function, the other being the correct gp value to use."
http://mikedimmick.blogspot.com/2004/01/ia64s-global-pointer...
http://www.cs.utexas.edu/~trips/overview.html
It tries to avoid some pitfalls of architectures such as Itanium. Seems fairly complex to me, though. Might be inherent in EDGE and VLIW's, though.
Talking exotic, look at No Instruction Set Computing (NISC) which was at least prototyped and published synthesis/compilers:
https://web.archive.org/web/20080302041756/http://www.ics.uc...
Really interesting stuff. Reminded me of Tensillica's tools that create a custom processor for your application. Need to accelerate your Hadoop, etc application? Run most of it on Intel CPU with an onboard FPGA & NISC tools doing the critical path. Intel's Altera acquisition might make something like that achievable in future.
Note: Used archive because their site is having a configuration error.
There are some consumer VLIW chips bumming around; I think the Nexus 9 has a VLIW chip - but it is essentially a hardware JIT compiler, translating ARM into its internal instruction set for hot paths.
Correction: AMD GPUs were VLIW when I took a class on GPGPU in 2011. Apparently, AMD subsequently switched from VLIW: <https://en.wikipedia.org/wiki/Graphics_Core_Next>.
http://www.embedded.com/design/prototyping-and-development/4...
Itanium code takes the form of fixed-width chunks, but they're a lot longer than a RISC instruction and contain multiple small instructions. All of these run in parallel. It's illegal to have any conflicts within a chunk (such as two instructions writing to the same register, or one writing from a register that another one is reading).
Itanium exposes this parallelism to compiler authors, so that a sufficiently smart code generator can take advantage of this. In practice, the sufficiently smart compilers never materialized.