Or maybe not functions, one might consider duplicating code. There is DRY principle, but there is WET principle[1] also. Both of them have their proponents, so at least they are not stupid. How WET shows itself in a highly-optimized programs? Did you try it? I have no link, and may be wrong, but I believe that John Carmack argued in a favor of WET, and I wouldn't say that Carmack doesn't bother with a program speed.
I mean, I'm not convinced, that goto is the game changer, I see other ways around the issue, which in some ways are superior to the solution using goto. So I'm not so sure that your way to use goto is not a misuse.
> if a developer is already perfect with writing good code, but (s)he needs to do something special, like the state machine in the article, the restriction just makes the developer's life harder.
The trouble is there are no developers who a perfect. So it is not intrinsically bad thing to make live of a developer harder by introducing restrictions. If it forces developer to codify more of his assumptions about his program, then it is harder to be imperfect.
The thing is that we used hot/cold labels to minimize jumps over the code and improve instruction cache usage, so calling functions isn't good at all. We wanted one big function with is executed as linearly as possible to improve instruction cache prefetching.
I'm unable to find measurements. I mean, the amount of code duplication is not an optimization parameter, the amount of Mb/s of HTTP parsed is. It is obvious that the size of code leads to a greater cache use and slows down a program, but any engineer's decision should be an educated choice between alternatives with different upsides and downsides. Spagetti code with goto jumping between branches is a bad thing from the point of view of maintainability. From other hand constantly overflowing instruction caches also is a bad thing. Which one is worse? How to compare apples and oranges? I do not know how to quantify maintainability (pity), but throughput of a program could be quantified easily.
> The thing is that we used hot/cold labels to minimize jumps over the code and improve instruction cache usage, so calling functions isn't good at all.
Function in C/C++ is a high-level abstraction, would it lead to a CPU executing 'call' instruction or not -- is a choice of a compiler and a programmer.
If we return to a premise of your article, that rust isn't good for a systems programming, I might notice, that many wouldn't agree with you that what you do is a systems programming worth mentioning. It is nice to know how much C++ code could be tuned for a maximum speed, but I'm not sure that I'd prefer hand crafted state machine to a parser generator, even if the latter would be much slower. If I choose hand-crafted state machine, I'd probably do it by splitting code as much as possible into a 'static inline' functions (most of them would be called just once), with meaningful names and some common arguments to make it as easy to understand and to reason about as possible. Yes, I'd sacrifice some of speed by this way, but I need not the fastest program possible, I need program that is fast enough and have other properties as well. There are more then one optimization parameter: it is systems programming.
Of course, your way have it's rights to exist, and probably it has it's niche too, but as I see it, it is a small narrow niche were available computational resources are scarce while throughput needed is high. In the most cases system software needs to be highly reliable, which needs code enabling us to reason about it, audit it, change it. It needs code which could be maintained by other people, not only those who wrote it.
The measurements are covered in slides 23-24 in http://www.tempesta-tech.com/research/http_str.pdf
> goto jumping between branches is a bad thing from the point of view of maintainability.
In our parser we have DSL for the state machine, so we encode which state the machine should go. The HTTP parser is one of the most updated piece of a web accelerator code and we don't struggle on the FSM - every developer in our team updates the code, even newcomers.