The Itanium processor, part 2: Instruction encoding, templates, and stops
blogs.msdn.com
blogs.msdn.com
The failure of the Itanium because of cheap x86 chips doing "macro-op" execution, more efficiently than the Itanium explicit VLIW, is the other side of the same coin: no matter if explicit or implicit, the way to increase single-thread IPC (instruction per clock) rate is doing more in parallel and hiding execution latency. So, explicit, or as hidden abstraction, VLIW-esque execution seems to be unavoidable.
Let me interject that it has also not been successful for AMD's line of videocards - they used a VLIW architecture before transitioning to a SIMD architecture codenamed Graphics Core Next or GCN in 2011. nVidia had been using SIMD for quite a while before that, but some of their latest cards have a few VLIW features like the option to explicitly specify parallelism between instructions.
IA64 had wide "bundles" for superscalar issue, but it explicitly avoided exposing the individual instructions as "ports" and instead had a more flexible setup where, really, the bundles contained "just a bunch of independent instructions". There architecture assumed they'd all be issued in parallel (so there were rules about dependencies between instructions in a bundle), but made no promises about whether they actually could be, or even whether or not more than one bundle could be issued simultaneously.
For a very modern take, [shameless plug] the Mill CPU instruction encoding talk is very entertaining: http://millcomputing.com/topic/instruction-encoding/
(Mill team)
A good example is the Motorola 56K DSP family. Instructions are 24 bit wide (24 bit wide memory!). They come in a few formats, which can be broken down into two major types: complex instruction, or simple instruction plus up to 2 memory moves and address updates.
It's basically designed around performing a single-cycle 24x24->56b multiple-accumulate with two memory reads and address increments, with zero loop overhead. The rest of the ISA really, really suffers for it :) But it's impressive when it's doing exactly what it was designed to do! It's a great example of a domain-specific ISA.
Also worth noting that in 56K, there is no score-boarding or other dependency tracking to automatically insert pipeline bubbles. That's the job of the hand-writer/assembler/compiler. If you get it wrong, it simply does something undefined. It's "fun" when this happen.
Because of the nature of the patent system, it is impossible to state definitively that none of the currently-popular PLs currently have any patent encumberances, but we can say that none of the project leaders or main contributors to the design of any of the popular PLs have been accused of pursuing patents on their creations. Nor does anyone AFAIK complain or warn about any popular PL's being patent-encumbered.
Although most designers of instruction-set architectures that have seen significant economic use have pursued patents on this or that feature or technique, Itanium is AFAIK the only one where the patent lawyers were involved in the early stages of the design of the architecture.