This is like most of the reason I opened this article. I wish they'd spend more time talking about this. In fact, most of the README covers details that are important, but, irrelevant to their algorithmic improvements.
Maybe they expect you to open up the code, but even there the comments are largely describing what things do and not why, so I feel you need to have some understanding of the algorithm in question. For example, the parts where the state table gets constructed rather unhelpful comments to the uninformed:
/**
* Build an ASCII row. This configures low-order 7-bits, which should
* be roughly the same for all states
*/
static void
build_basic(unsigned char *row, unsigned char default_state, unsigned char ubase)
...
/**
* This function compiles a DFA-style state-machine for parsing UTF-8
* variable-length byte sequences.
*/
static void
compile_utf8_statemachine(int is_multibyte)
Even `wc2o.c` doesn't delve into the magic numbers it has chosen. I was hoping this repo would be more educational and explain how the state table works and why it's constructed the way it is.Does anyone have a good resource for learning more about asynchronous state-machine parsers, that also could hopefully help explain why this is better than just iterating over characters? I'm guessing maybe it's the lack of branching?