It's just decoding that requires pipeline busting branches. Well, I guess it's fine if you take branch mispredict for the large string case.
Maybe 1 byte with highest bit zero, 0-127, 2 bytes for 128-32767, etc?
The irony is of course now we have variable length encoding for the string size...
I think just using size_t instead makes more sense than saving a byte or two per string. Memory is cheap, but CPUs aren't getting much faster.
When size is an issue and nothing else prevents it either, just LZ4 (or similar memcpy order of magnitude speed compression) the string.