In order to look up glyph in your font, you'll typically need a 32bit unicode code-point. A utf-8 string encodes those code-points using a variable-number-of-bytes scheme, so you need to decode each code-point before you can render it.
Yes, but this code expands it to a 32-bit array. That seems a bit silly because you'd be wasting a lot of memory bandwidth there. I suspect the loss could even be greater than what you've gained from removing those branches.
Not quite, it's expanded to a 32-bit integer, one codepoint at a time.
Perhaps you're misinterpreting what a decoder does. From the article:
> This week I took a crack at writing a branchless UTF-8 decoder: a function that decodes a single UTF-8 code point from a byte stream…
To print it to the screen for example.
To get the unicode codepoint. For example, you need this when you want to render a character, or convert it to another representation.
Some languages represent unicode strings as objects quite separate to their binary encoding, e.g. Python.