7 karma · joined August 27, 2024
So it seems in regard to bit shifts, C++ behaves slightly differently (it seems to have less UB) than C.
Also, for UTF8 specifically, cutting code points in half is fine as long as all strings are valid UTF8. The UTF8 encoding is prefix free, i.e., no valid code point is a prefix of another valid code point, so for prefix matching we can usually just compare bytes.
It only gets more complicated if you add collations or want to match case-insensitively. But at that point you need to take into account all edge cases of the Unicode spec anyway.
How do I know they will be worth anything when the company exits? I probably won't get a contract with the company that tells me exactly how I'll be paid in cash if the company is acquired. Instead, I will get a token that will be bought back by the company at a price set by the company at a time set by the company.
Another similar issue I see is dilution: How do I make sure the company doesn't just issue a billion new tokens so that my tokens are diluted to nothing? How does the token price change when new real shares of the company are issued (presumably investors want to get real shares instead of tokens)?
I had planned to write about the Kahan summation, but I didn't want to overload the blog post with too many different concepts. I'm definitely considering writing about it in a follow-up post.