Pg-Emoji
github.com
github.com
I'm pretty sure you can already put emojis in the text fields of postgres. (or at least I'd be surprised if you couldn't)
That would certainly make more sense than anything I’ve been able to glean from it.
I’ve definitely used Postgres text columns to store user-provided text values that included emoji in the past. They’re part of my standard test case for any user input.
https://matrix.org/docs/spec/client_server/latest#sas-method...
Without digging into the source - and I intend to do that, if someone more familiar with this doesn’t chime in - it appears that it’s targeted at reducing resource consumption.
UTF-8 can encode emoji fine. Consider (“ grinning face with smiling eyes”), which is `\xF0\x9F\x98\x81` in bytes. That’s four bytes. From the pg-emoji Readme:
> A lookup-table is constructed from the first 1024 emojis from [https://unicode.org/Public/emoji/13.1/emoji-test.txt], where each emoji maps to a unique 10 bit sequence.
> The input data is split into 10 bit fragments, mapped to the corresponding emojis.
If my understanding is correct thus far, then instead of storing four bytes for each emoji, you’d only need 10 bits.
I don’t know where this would be worthwhile.
I’m further confused by the purpose of `to_text()` and `from_text()`. Their example shows a string composed of mostly Latin characters being encoded into a string of emoji and back.
This is meant to be used if you want to pass some text containing escape characters or perhaps JSON. Note also that the first emoji is a checksum, which might be useful if you want to make sure a user correctly copy/pasted a string, as opposed to sending a raw text string (without checksum).
I guess I don’t understand how this is an improvement. Perhaps it’s because I typically interact with the DB through a language-specific library/protocol like Python’s DB API, which handles escaping strings and parameterization without my really having to think about it.
Could you provide a specific example of when this might solve a real-world problem?
For instance, if you instruct the user to copy "this text string" and paste it somewhere, some users might copy the text string with the double-quotes and some without them. By instead emoji encode the string, the receiver of the copied emoji string can detect if not all emojis were copied.
Further... if the receiver can validate the encoded string itself, they implicitly already have the string. Why require the user to copy/paste at all? If you meant "Ensure that the user hasn't copied quotation marks as well", then we're back to it being application logic.
If I'm understanding correctly that the primary benefit is that there is a checksum, then there are already many solutions for this in common use - base58checksum, as used to ensure the validity of Bitcoin addresses, comes immediately to mind. I wrote an implementation of that quite a while ago: https://github.com/lyndsysimon/cryptocoin/blob/primary/crypt...
Please don't misunderstand, I'm in no way intending to be argumentative. I don't understand the practical use of this project, which leads me to believe that there is a problem being solved that I lack the context to identify.
Its a system for encoding data as emojis. "this is a string" => some emojis. ie baseemoji or base1024
(Admittedly there are still issues within the emoji space such as some of the "faces" are quite similar in appearance in many fonts and still easily confused. Plus in the larger emoji space the subtle differences of skin color/gender can be easily confused if you have to rely on them for distinction. Restricting to only 1024 emoji and fewer ZWJ sequence variations presumably takes care of most of those issues.)
https://matrix.org/docs/spec/client_server/latest#sas-method...
Data encoded in base1024 (here using the 1024 safe chars represented by emojis) gives much more efficient storage usage.
16 kB raw data encoded in base64
ceil(16*1024/3)*4 = 21848 bytes long ~= 21.8kB.
16 kB raw data encoded in base1024 ceil(16*1024/9)*10 = 18210 bytes long ~= 18.21kB.
So base64 needs about 19.7% more data storage than base1024 and both can be used anywhere utf8 is supported.Let the baseEmoji revolution begin...
Long ago, I made http://github.com/strayptr/memes
If you scroll down to Kappa, you'll see a Lambda, which is pg in the style of twitch.tv's Kappa emote.
https://cloud.githubusercontent.com/assets/12214175/7581578/...