Teemoji: Like tee but with emojis
github.com
github.com
Please consider making it available on MacPorts (for those who don’t use Homebrew). Anyone else here who can bring this to MacPorts?
If anything, it's a bandaid that highlights the biggest flaw of UNIX philosophy: everything is passed as unstructured text. Because of it, half of shell programming is just piecing together ad-hoc, buggy parsers that interpret the input, possibly rearrange it, and then dump it down the pipeline as unstructured text, so the next step can do it all over again. And then, of course, every CLI program has to do that too.
This program is using a machine learning model to parse its input; while this may be the only reasonable way to go about guessing emojis from arbitrary text, I can easily see people doing the same to parse outputs of Linux CLI tools at runtime, because it looks much more pleasant and might be even more reliable than writing input parsers by hand. Let's pause here to consider the absurdity of that situation.
That all said, I don’t really disagree with any of your individual points. But teemoji is clearly not meant to be considered a serious tool so I wouldn’t be too critical of UNIX that this tool exists.
Of course this is another standard that tools would have to be incredibly careful to keep track of, but JSON is decently mature and a bunch of modern tools can operate in it, so at least there's slow progress away from the plaintext quagmire.
Besides, what format should the data be structured in? In the 1960s and 70s, it might've been GML. In the 80s, maybe SGML. In the 90s, XML, and later maybe JSON, or... shudder YAML? Or, no, it obviously should've been some binary format.
OK, so let's say everyone agrees on a single format for several decades. Should all programs written for all those shells that support it also support the data format? Even if that's handled by shells and not the programs, there will surely be bugs and incompatibilities in each shell implementation, leading to fragmentation.
Furthermore, will the data format support versioning? Will it be backwards compatible? Will all programs or shells need to support this? Will users have to mix and match, and experience all sorts of compatibility issues?
So, clearly, unstructured text is the only "format" that is both flexible and future proof enough, at the expense of convenience for the user to glue programs together. I'd rather have that than have to depend on the structured format du jour that will have to be continuously maintained and supported by all programs in the ecosystem.
Besides, there are projects like Nushell and Murex that wrap the input and output of existing programs or write their own to support piping structured data. I'm not a big fan of them, but it points out that unstructured data is not some UNIX flaw, but a deliberate design decision which is IMO partly what enabled it to proliferate in the first place.
Maintaining a shell that supports piping structured data can only be done by a single entity that oversees the entire ecosystem, like Microsoft does with PowerShell, or the above mentioned projects. That is the antithesis of the UNIX philosophy where disparate tools can be written by anyone, yet somehow still be made to interoperate by the user. It's also a gargantuan project that realistically not even mega corporations can successfully maintain without limiting the user or introducing bugs. Plus it will always be a bottleneck for adding new tools to the ecosystem. For these reasons I don't foresee any of the structured data shells to be nearly as popular as UNIX shells have been and will continue to be 30 years from now.
simply impossible. when a input parser breaks, it breaks. with a LLM, it breaks... sometimes, for different reasons. Let alone the enourmous computing power needed compared to simple input parsers.
tee does something very specific, it makes an unmodified copy into a file (one branch of the plumbing tee-joint) as it passes the stream to stdout (the other branch); as opposed to sed or awk or even grep, et al, which modify the stream. How in hell is this inspired by tee which does not modify its inputs?
and who capitalizes Tee?
I still absolutely love this project. It’s inspired me to play around with CoreML.
By the way, I use that emoji to test whether astral planes are handled correctly.
I need to recover from the cognitive dissonance now...
:sadface:
When I was in kindergarten there were tasks where you'd get a paragraph of text where you had to fill in blanks, and next to those blanks were pictograms of whatever noun was expected.
Whenever I see people overuse emojis, for example "Yesterday we flew [plane emoji] to Japan [Flag of Japan] and took the train [bullet train emoji] and saw Mt. Fuji [emoji of Mt. Fuji]", I always think, "This person is still in kindergarten."...
https://www.apple.com/voiceover/info/guide/_1121.html
On Linux there is
https://help.gnome.org/users/orca/stable/index.html.en
I have no idea how good that one is, but VoiceOver is pretty good.
I did a ShowHN because I thought it was a good idea but it flopped. Maybe I should pivot my idea toward kindergarteners!
That said, why is a Japanese sentence written using one set of pictographs respectable and proper, but another set of pictographs means the writer is in kindergarten?
I was doing `echo cat | teemoji` tests and it would work, but ironically 'echo happy face | teemoji` and the like didn't work so well for many other obvious single-word emojis. But it did a "checkered flag" for "I got the job done".
It's very slow in comparison because it uses llama3.2 model but it's a very common model and the emojis generated are more related to the lines.
Next feature would be to add an emoji model to this tool to make it faster.
You could start with something naive using pattern matching, common formats, common words, etc. and apply the most relevant emoji. It'd be fast and small but not as flexible and no cool ML in there to enjoy. It'd probably be how I'd go about it though - you could actually use AI at the dev stage to get the dataset for this pretty tight I reckon..
Or.. you could use a tiny embedding model (which might even be fast on CPU alone), assign emojis to the best appropriate regions, then pick the 'closest' emojis for embeddings of fresh text. Potentially inaccurate but could be fast-ish.
Or.. all the way up to using something like TensorFlow and training a classifier, but it'd probably take a fair bit of space and not be super fast if it fell back to CPU. Would be a fun experiment! (I haven't done this on Linux before so I might be wrong in 2025.)
One big benefit of Core ML is it Just Works™ if you meet the hardware and OS requirements. On Linux you'd potentially need to fall back to CPU or have a lot of various dependencies to lean on (maybe ONNX takes care of all this nowadays?)
I've been thinking about training a model, maybe a T5 to automate the task. I've tried asking Microsoft's Copilot to do it, and it is ok but makes decisions that I wouldn't make. I argued with it a lot and couldn't get it to draw ⁰₀⁰₀⁰ as an imitation Olympics logo and had to do it myself.
if I may offer a small nitpick for feedback, I am seeing wrench emojis on empty lines which create a lot of noise.
If you are interested in PRs (and making the change) I might try to take a look at the code and see if it's possible for me to contribute this to it.