Yet if I put on my software cap and think of how things should be, the whole point of a language is to create an abstraction. This allows a designer to think and express their intent at a higher level. If an abstraction is not complete then the designer can never stop thinking in the lower level terms, which nullifies the benefit.
Hardware description languages can never be as abstract as programming languages for the same reason that we can't create a single unifying abstraction for imperative multi and single threaded code: synchronization. Since each algorithm or data structure can have unique memory access/sync patterns, especially once optimized, there is no way to abstract away parallelization so we're forced to work with a large number of low level primitives like queues, locks, etc. In silicon, not only is everything running in parallel but each abstract unit of functionality has many more "degrees of freedom" like synchronization points, clock speeds, and physical location. It's simply impossible to cleanly abstract over all of the little tweaks you have to do in the implementation phase to get your blocks of transistors to synchronize properly.
I'm not sure what it would even mean for HDLs to be "as abstract" as programming languages. They're different paradigms, and they have different abstractions, so you can't really compare them directly.
What I do know is that you can do a lot better than Verilog, because I've experienced it. I worked in Bluespec for a few years and as a language it's so much better than Verilog it's not even funny.
Indeed, using it made me stop thinking of Haskell as a programming language, and start thinking of it as a declarative language for constructing imperative programs in the IO monad.
Yea, the parallels between hardware design and haskell are striking. (Not just haskell, but it is perhaps the easiest to see.) In undergrad, I wrote my basic cpu project (16-bit risc) in haskell, then got it to write the verilog for me. This had the benefit of having an incredibly fast iteration cycle (including testing) paired with the ability to output an actual hardware description at any point. Very powerful. It really hammers home the difference between computation and I/O, a common theme in both CPU design and functional programs.
Also—and this is just my opinion—I much prefer declaring the computation with logical combinators rather than manually specifying how the bits are routed. This is much easier in haskell (or any functional language) than it is in verilog.
http://people.eecs.berkeley.edu/~bora/Conferences/2015/ESSCI...
I'm a fan of state machines, but I don't think they're necessarily more performant. There's nothing about them that keeps you from doing anything crazy that will bog your program down, but I find them easy to reason about. I think they're a stylistic thing, and the only real advantage I can think of is in maybe helping the design of the program. Having to define your states and how you will move between them is usually a good idea.
I'm neutral about state machines. Obviously in things like microcontrollers and HDL, state machines are really necessary. Outside of those two examples, state machines can either make code much more stable or just add unnecessary overhead with all the wasted cycles. They definitely have their place but there are better options available sometimes like threading a process. I had a project recently where we needed to build a custom message broker. A teammate wrote the entire thing in python as one big state machine. Everytime we needed to add another queue, it involved adding so much extra code to handle it. Another teammate suggested just having everything threaded so we reduce our code down to a template queue and can dynamically create them as needed. The state machine might have been more efficient memory-wise if we were doing it all in C on a PIC or more conservative in PLU usage if it was an FPGA but we were using python on a raspberry pi. We ended up with a broker that could handle more bandwidth in terms of how many queues we could publish to and consume from.
Obviously we didn't really shrug off state machines. We just created many smaller ones (two states: wait for message, push message) rather than a huge one. I think the massive state machines we learn about in our freshman classes are unnecessary in many cases. Memory is so cheap that its not always worth it to be so conservative/old school.
We could have just used something like rabbitmq but we didn't want the extra overhead of a service like that when we had so much more resource heavy stuff running on the pi.
Personally I'm a fan of the smaller state machines you describe. You still got to deal with synchronizing them, but it just seems easier to grasp for me.
Essentially since Bluespec's model is a set of conditional transactions on your state, you can write an imperative sequence of transactions with a few basic flow control primitives (if-else blocks, while loops etc.), and it creates the state logic automatically.
Unfortunately it didn't seem to have received much attention, and I ran into quite a few problems when I tried to use it for anything serious (including one straight up behavioural bug).
Sure there are some cases where it might be a good idea, but if it's a requirement to understand the design, then your language isn't doing its job.
Optimization is important in hardware due to cost, speed, and power implications of hardware design. These tradeoffs are there for software design as well, but to a much less extent.
You simply can't do that when designing silicon. All of your transistors have to be physically connected on a 2d plane so you can't abstract over the physical location or layout of your "objects" (logical groups of transistors and how they interconnect in this case). This means that the vast majority of possible Verilog designs are impossible to physically implement because there is only a finite amount of space to route traces from one transistor to another. Writing Verilog before drawing a block diagram is like writing an embedded IoT framework in Ruby before checking whether your target microcontrollers have enough ROM/RAM to run the Ruby interpreter.
For example, if you had a stack allocated object containing a queue and you changed it to always heap allocate the queue, it would still do the same thing as before but with a different memory layout. If you did the equivalent in silicon (like move the location of a group of transistors), at best it would break your functionality and at worst it would make the design unsynthesizable. Swapping an interface to point to an implementation with a different size (I.e. switching an IQueue* to an MPSC Queue from a SPSC Queue) is a trivial change with software but when you swap a 100 transistor interface with a 1000 transistor one you have to move the blocks around it, often requiring significant changes to avoid breaking functionality.
Of course, we were in a research group, so performance was less important to us. And we were working with FPGAs, so we didn't need to get it right in one go (and I think the fabric is better able to cope with looser style).
That said, I think there are cases where you can afford to sacrifice a bit of design efficiency to increase development speed. And if you're working with FPGAs, you may even be able to optimise later.
EDIT: Also the impression I got is that while industry does spend a lot of effort on it, complex designs are nowhere near as optimal as they could be. The first clue should be that an Intel chip has clearly defined blocks on it, rather than being a mess of interleaved logic. They divide the design like that to make it tractable for the designers.
> If you did the equivalent in silicon (like move the location of a group of transistors), at best it would break your functionality and at worst it would make the design unsynthesizable.
I don't think this makes much sense unless you're presupposing a prescriptive physical layout. Sure, if you were to move a bunch of synthesized circuitry somewhere else and then try to reconnect it, it's not going to work. But that's not something you'd actually do in an abstract model.
What you might do is move a bunch of HDL logic from somewhere in the module hierarchy to somewhere else, in which case the synthesizer will move everything else around to compensate.
What were you researching/designing? The higher your clock speeds/bandwidth requirements are and the more complex your designs, the more locality matters. The last design I worked on, for example, was a 32x32 crosspoint switch with total bandwidth in the low hundreds of gigabits and we eventually had to tweak cell locations manually on the mask to get the design working to spec. In modern chip design, bandwidth on that scale is pretty common so i don't think we're talking about the extremes of hardware design.
>> I don't think this makes much sense unless you're presupposing a prescriptive physical layout. Sure, if you were to move a bunch of synthesized circuitry somewhere else and then try to reconnect it, it's not going to work. But that's not something you'd actually do in an abstract model.
Again, depends on a variety of factors. Some of the designs I've worked on have had internal clock signals that couldn't travel across a fraction of the chip before attenuating too much to switch a transistor. Sure a low speed design can be synthesized without caring about physical layout but once you're in the RF/high speed digital realm, the relative location of dependent high speed blocks matters much more and the HDL abstractions start to leak or breakdown entirely.
Obviously these are very different fields, to the extent that I wonder if Verilog is even appropriate for what you're talking about, except for legacy reasons. If you're working in a domain where layout is so important, it feels like you should be using a tool that represents it directly, rather than inferring it.
Doing some work with HDL in a university research setting is different than shipping a working product with the associated constraints.
EDIT: clarification
A littler bartering and I ended up owning display controllers, camera input peripherals, interrupt controllers and other such fun for 65nm and 40nm ASICs - mostly verilog (and, sigh.. TCL for test benches).
I found a singular insight got me all the way from zero to tapeout for all sorts of peripherals and infrastructure without any real support from other RTL peers in the team - design everything around a single clock (inc crossing all async signals into this singular domain as fast as possible) and then explicit pipelining of the peripherals functionality using a judgement of the process/clock speed to guesstimate how much complexity/gate density each pipeline stage could support before starting a new stage.
As I recall, my peripherals were marginally bigger than they needed to be but were extremely easy to push through the back end as the constraints were simple to define and synthesize.
I miss those days - the pressure to tape out a singular piece of art was a great way to focus the brain and the chip bring-up was exhilarating.
It makes no sense
I don't want to design from the bottom up and then if I'm "lucky" the compiler will give me what I wanted.
These are tools meant to scale out designs that you already have a good understanding of. Hardware is implicitly bottom up because you're dealing with physics at the bottom. Very few pieces of hardware are lenient to a single cycle of latency much less 10's of milliseconds.
In the early days of compilers, of course, you were mostly writing C as a macro language for your system's assembly language. If you wanted your program to perform well, you'd have to write C that was, more or less, a translation of assembly that you'd constructed in your head first. If you wrote bizarre C, you'd either get incorrect results, or if you were lucky, you'd get correct but inefficient results.
But that's also Dan's point: Verilog isn't a "high level language". You don't write programs with it, you describe hardware with it. (In fact, that is why it is called a 'hardware description language'!) So if you try to write a program, instead of describing hardware, you'll get something that isn't really either.
Maybe in really old compilers
> You don't write programs with it, you describe hardware with it
Which is fair enough, but it seems the "hardware description" pretends to be of a higher-level than it really is.
If you need the user to describe gates and flip-flops and how they connect then make them describe this.
You're talking about RTL, which is exactly what these languages output.
Fundamentally they're not programming languages, unfortunately the initial instinct is to treat them as such and it leads to a ton of confusion.
If VHDL/Verilog would output RTL, you could easily analyze it just as you analyze assembly output of your favorite compiler. Unluckily the output is some proprietary bitstream for the FPGA.
So why am I wasting time with verilog then if I have to "design" in low-level then translate it to verilog?
In the early days of compilers, we were mostly writing FORTRAN.
Then why doesn't the language let me describe what I want, and if reality disagrees, throws some compiling errors back?
Our current HDLs are lower level than block diagrams. And anything higher level they may claim to provide is iffy and won't work on practice. That's not a problem with the languages, but it is a problem for hardware development.
You mean typing random stuff into the editor without understanding what any of it does and hoping for the best? Because all of the online examples I've seen of "HDL done like programming" involved exactly that. In my experience, HDL design is very similar to functional programming (minus recursion) and the principles that work in one translate well to the other.