Terminal escape sequences in Git commit email field
twitter.com
twitter.com
Terminal emulators are, far and away, the most archaic computer technology we still use daily.
Believe me when I say, it's a very hard problem to solve since everyone has subscribed to Unix's philosophy of "everything is text".
As much as I dislike Microsoft's philosophies on software, they had the right idea with Powershell. Its execution is just abysmal.
Expect this stuff to change drastically in the next 5 years.
Bwahaha... is this your first encounter with backwards compatibility?
Changing this would be very hard, but this has nothing to do with somebody embracing some grand philosophy, but with tons of existing software relying on Unix like systems to, well ... be Unix-like, because that's what the software was written for.
Invert the shell and terminal: every shell command (with unredirected streams) gets it's own pty.
- Like the those repl "notebooks" you can scroll each output separately, without hacks.
- Backgrounded processes won't spew garbage when you are trying to type. They just continue running in their terminal. You can even background something like vim or a game that takes the whole terminal, and it's just fine!
- Rather than hacking scroll back onto processes that might use tty stuff in expected ways, just straight-jacket them by giving them plain pipes instead! The OP exploits are then opt-in impossible.
- Having powershell-esque "post-text" commands that dispense with the bullshit are far easier to manage: also give them regular file descriptors and negotiate protocol or whatever, just like the above, rather than a pty.
- stdin can be a nice form with submit button rather than fixed (usually line) buffering policy.
- shell might even be simpler as there is less careful juggling of shared resources except where the programmer asks for it.
I also use ssh multiplexing instead of tmux for extra shells. When I do need to use tmux, I hate that read only mode prevents scrollback. The above builds nicely upon this:
- The persistent remote shell has as many ptys as needed before for chained commands which might still be running, or separate "sessions"
- The local client can use native UI elements for everything, no TUI jank.
- Scrolling over many (live or saved) terminals, the normal scrollback case, is possible and very side-effect free.
Anyone see any issues? I don't for anything I do daily (git, coretuils, etc.)! I don't think I have time to make this, so please someone else do.
The problem, of course, is this either breaks existing software or requires new software to be written from scratch to utilize the new frameworks. This was (one of) the reasons plan9 never gained much traction, and the big hurdle for any contender to the TTY/shell throne.
(It's also worth pointing out microsoft themselves reworked their terminal interfaces to be more unix-like for this very reason -- software compatibility)
Focusing on this point, you can background vim just fine right now - I have tons of vim processes backgrounded all the time.
Perhaps for most use cases involving vim it doesn't make that much difference—what would it be doing in the background anyway?—but there are other TUI programs that could be doing useful processing in the background if they didn't need to share the pty with other processes.
Mentioned this comment here: https://blog.williammanley.net/2021/03/31/seen-on-hn-invert-...
You’d need to implement a new protocol between terminal and shell. Maybe the shell would be in charge - allocating PTYs and multiplexing the output back to the terminal. Or maybe the terminal would be in charge where it would allocate PTYs for children and send them to the shell using FD passing. There are advantages both ways.
Integral to its success though would be getting the support for this new protocol upstream in bash. Bash is the most widely deployed shell and the only way to get to the point where you can log in to a new machine (or over SSH) and expect this to “just work”. Without that it would remain a niche and could join the graveyard of other improved shells/terminals that approximately no-one uses.
As such the protocol would have to be minimal and easily implementable in C. Systemd’s protocols could be an inspiration. I’ve implemented both sides of sd_notify and LISTEN_FDS before, in more than one programming language and it was very straightforward.
As long as this interface remains undisputed, "everything is text" is a feature, not a bug. Arguably, one of Unix's greatest.
Terminal sequences, unicode all violate the identifier rules. There is no u8ident library to sanitize this input or output.
Unicode standardises identifier rules. The default identifier as per UAX31-R1 can be checked with a regex:
perl -E'say "my_identifier" =~ /\A \p{XID_Start} \p{XID_Continue}* \z/x'
I guess that's not what you mean, so an explanation with some details is needed here.Btw your snippet does not work for identifier validation. There is much more needed to pass the identifier security guidelines.
But here we just need a simple terminal escape sequence stripping library. Unicode is a bit harder. Only Java, rust and me did that.
I do not expect text to be significantly replaced any time in the next 5 years. Accessibility - screenreaders, increased text size with text reflow - is one massive reason.
Isn't the real solution to do away with the idea that we need to remain compatible with VT100 video terminals?
We've learned a lot since the unix philosophy was created. One of those things is that, if you want software that lasts, software that performs well, you need to be strongly typed.
Plain text is as far from strongly typed as you can get. Passing plain text between programs in 2020 goes against everything we've learned in the last 50 years.
In theoretical terms, strong typing every, always is preferred. In practical terms, this devolves into a Java-esque nightmare of what's-the-type-for-the-parent-type-of-the-thing's-type absurdity. Where people stop using your system because they don't want a second job learning its Borges-esque comprehensive type encyclopedia.
In practical terms, bare minimum typing that prevents 80% of errors, but doesn't require gyrations to cover the remaining 20% works. And given a choice between the two, more software is going to be built in systems that function this way.
Does it break all the damn time? Yes. Does it still get more done? Yes.
I know but my point is that the interface itself is still text.
I would have to say the opposite is true.
It's text based interfaces that have stood the test of time, and now even flourish in HTTP. We have more widely used text based interfaces than ever before where REST,HTML, XML and JSON seem to rule most machine-machine communication.
Strongly typed binary RPC such as DCOM and CORBA however are now nearly gone.
Unix itself, for example, won't last even 10 years. Look for it to be gone by the mid-late 1970s as it gets replaced by strongly-typed objects passed directly between applications to control our fusion-powered flying cars.
What will computing look like in 500 years? Do we really believe that we reached the pinnacle of OS development in the 1970s?
I'm also not extremely familiar with the concept, so at the moment, I'm reading this document to understand it better https://www.varonis.com/blog/how-to-use-powershell-objects-a...
Look at it the other way round. Embrace text. Text is not the problem, non-text is. You'll take text from our cold, dead hands!
If you are parsing "ls -l" output, you are doing something wrong.
Use your language built-in features, every language has them (for example in bash, use *-expansion and [-commands). If they don't work for some reason, there is "stat -c" and "find .. -printf", which both produce text which is absolutely trivial to parse.
Of course, ls does more than just list file names so it can be tempting to utilise its features. coreutils ls has (relatively) recently received an additional output format that can be unambiguously parsed, but that's as far as the portability goes.
> Use your language built-in features
The issue is that current shells _don't_ have a built-in structured data type for e.g. a `struct stat`. But why not?
`ls` calls some APIs and gets some in-memory `struct stat`s populated by the kernel. Then it throws away 90% of the data that the kernel copied to user space and then serializes it as text. Why not pass the structs themselves to the next process? We can't currently (except with powershell?), so you have to write actual code.
This is a "bright line" between code and shells that could very well be blurred, but it would take an agreed-upon serialization format for posix-y data structures.
This is why in shell, if you want programmatic "stat" output, you use "stat" tool, not "ls". "ls" prints all at once. "stat" has a custom output format, which means you print exact fields you want, so your intermediate results are concise and readable. And since every element of the "struct stat" is an integer, a simple space-separated format works very well.
(that said, I would not mind seeing more tools print JSON or JSONlines. I add this functionality to many of the tools I write, and it is pretty powerful in conjunction with jq)
Because then the output of ls can be used by programs that don't know what a list is, or what a file is; and this is breathtakingly beautiful.
Structural data can still be serialized easily, take json for example which’s spec fit on a napkin.
It's not that uncommon for ol' C hackers to directly write those stack and heap locations out to disk and call it a "file format." Trouble is, you're almost entirely at the whims of your platform and compiler as to what the actual layout of that is.
If you're thinking, "well that's dumb, why doesn't C have a standardized representation for those in-memory objects that hides platform differences", it does: printf and scanf.
Text isn't necessarily a great answer to that problem, but it definitely is an answer. Others include packed structs with htonl and friends and low-overhead serialization formats like protobuf, Thrift, and Avro. Inside, say, Google, you have "everything is a protobuf" instead of "everything is text," and it does end up working roughly as well as you might expect. That is to say, reasonably well, but with its own sets of problems that people won't ever stop complaining about.
`ls` is for you, the human.
system("ls");The moment you want to do something else with the text, you have to parse it in some way. At that point it's no longer text, it's a poorly specified serialization format.
I also worry about a bad standard becoming popular a la most of the web
It's not really good, but calling it abysmal is just too negative in my opinion :) It does get a lot of things right after all, especially in the "not everything is text" department. Also a lot of things wrong (my main gripes would be a lot of the syntax in general, case insensitive, usually more than a couple of ways to achieve the same thing). So in the end the learning experience/curve for me was the same as for bash (or languages like C++) and just as long/steep probably: feels like suffering until you've learned all the caveats by heart and in the process of doing that also learned how to actually use it, then it becomes ok to use. After that process I'm leaning towards favoring PS over bash though. Now it's possible there are already other shells which get more things right but I really doubt that in my lifetime I'm still going to spend (waste?) time on learning yet another shell. Or it must be convincingly good and provably fast to learn :P
- the pattern that was searched for
- the search options from the invocation (case, line folding, etc)
- the text of the matched segment
- the byte range of the matched segment within the file
- the line number(s) of the matched segment within the file
- the file name
- contents of any (named) capture groups from the search patternAdding a --json flag is just the start. Ideally, other grep-like programs would need would want to emit the same format. But what happens when some grep-like programs have additional features?
Suggesting a structured representation as an output format isn't a panacea. It's just the start and it's not clear at all to me that it would be better than what we have now.
Also the commands were super long. Sure there was the ability to alias them to something shorter but still....
That being said... powershell has a lot of potential.
(Also I’m woefully out of date... I haven’t touched powershell in a decade. I hope somebody tells me it’s all better now)
Windows Terminal is gr8 and removes all those problems. You have alternatives too like ConEmu.
Long commands are OK, use aliases. Bash has long command arguments too and nobody complains.
Saying Powershell has a lot of potential is kinda lame - its best of them all.
Our continuing mission: To explore strange new platforms. To seek out new bugs and new software. To boldly shitpost where no one has shitposted before.
Mission accomplished, LMAO.
One of the previous times I hit the front page of HN it was for inserting snark into an X509 certificate and then using it on my site.
The Unix way seems to be to not solve the problem at all, unless you're writing a text editor, and then you reinvent the wheel and solve it yourself.
Apache 1.3 (way back when) would escape escape sequences before printing them to the log, to prevent this exact same thing.
It is the responsibility of the program dumping data to the terminal to escape things prior to dumping to the tty.
It also probably doesn't help that the problem is unfixable in general without rewriting so much software... and that things mostly work okay, most of the time.
As opposed to what, the ẃ͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́ͥ́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́̂́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́ͥ́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́᷅́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́ͥ́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́̂́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́ͥ́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́ͅͅͅͅͅͅͅͅé͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́ͥ́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́̂́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́ͥ́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́᷅́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́ͥ́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́̂́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́ͥ́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́ͅͅͅͅͅͅͅͅb͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔͔̩͔́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́ͥ́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́̂́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́ͥ́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́᷅́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́ͥ́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́̂́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́ͥ́͗́́͗́́͗́́͗́́͗́́͗́́͗́́͗́ͅͅͅͅͅͅͅͅ? :)
(note: appearance of above may depend on browser)
If anything is text, nothing is text.
There must be some libraries or something that filter Unicode...
It has good test coverage, fails safe (if it doesn't understand a sequence, it still strips the escape character), and can preserve colors in cleaned up text. It is directly and indirectly used by quite a few projects. It receives extremely few bug reports, because I took the time to do it right.
If you want a good way to do this, probably start here. It's three short regular expressions and a simple state machine (to clean up single-line edits).
My only complaint about the code at this point is it operates on string instead of []byte, but I don't want to break the API and nobody has complained.
There are many reasons users might not want/be able to use Go, or any other language this might be published in. Given the terminal's provenance I can't help but be absolutely certain (despite having no substantive evidence) a lot of people working with C in Interesting™ situations would love something like this, for example.
(I must admit I skimmed it so fast (on my phone) I didn't realize the comment about RGB colors was only referring to the line directly underneath... and then didn't stop to think "RGB can't need all that", woops.)
The nice thing is this is definitely simple enough to be able to just directly port somewhere else, which I guess would make extracting a higher level representation premature-optimization-level superfluous complexity.
*Adds to bookmarks*
Done -- as simple as it gets.
I'll definitely need to add support for that. Possibly annoying to strip that with utf8 involved. Any idea how many terminals support 0x9b?
For example this shows red on xfce4-terminal in UTF-8 mode:
printf "\xc2\x9b31mHi\e[0m\n"Of course your code is a library so you have the library author's prerogative to punt that problem to the users :)
I thought that this was what isprint() was for?
https://twitter.com/cmmrc2/status/1329991505524690944
I'd been looking for an artist for months at that point, and looked through his work and saw some nice portraiture and dropped him a DM.
I can't figure out who shared it, but I think it was another pixel artist who I liked and was hoping would open up commissions again.
I wish it were more common for artists to auction off at least some commission slots.
This is gold.
And "git" for example applies "less" to pretty much all output by default, which makes most of the git-based attacks, including this one, irrelevant.
(Unless sixel required nul bytes in the encoding)
Though I suppose it's not technically a git commit email, it's still a git command
https://news.ycombinator.com/from?site=twitter.com/ryancdoto...
i've used ansi sequences in my zsh prompt for 20 years to make colors and move the cursor; it's just in-band ascii that is interpreted by the terminal emulator, no?
With default options, you only get colors, and this is pretty safe.
git log --format=%ae | sort -u
the cursor movement sequences will be preserved.
I haven't delved too deeply into what git actually does here in terms of processing for "pager" output vs not.