Linux terminal emulators have the potential of being much faster
phoronix.com
phoronix.com
While clang is still in my experience faster than gcc, it turns out handling that last 5% correctly can knock a lot off your performance. Of course with terminals it might turn out you don’t need that last 5%, or you really can get all the way and keep the speed, but this isn’t the first project dedicated to making “the fastest terminal”.
Not sure how this got off terminals and on to compilers, but here we are.
Having written the profiler I used to optimize this (Sysprof), and a large portion of the OpenGL renderer that GTK uses, I know a thing or two about what code runs in the hot path from VTE. So I could target something that does enough of that to see what should be possible as a best case / upper bound.
That's why I'm fairly confident that adding the rest of the features that VTE has (which is one of the most comprehensive terminal emulator code bases out there) wouldn't affect this particular test case.
That said, I have no interest in doing that. I'd rather just go make VTE fast at drawing and call it a day. I have literally dozens of other projects in GNOME that require my time.
I hope you are right, and also someone else will pick this up -- I've just seen an awful lot of projects over the years that provide a huge speedup for 90-95% of the functionality of other products, and very few manage to keep the speedup once they hit full functionality.
Compilers could be faster, but at this point they've become optimal enough that a lot of the improvements would be trade-offs for something else. In my opinion, not so for terminals: I'm pretty sure you could analyze the amount of time spent rendering in a terminal emulator and come to the conclusion that a lot of it is essentially wasted, with exception of a few more modern terminal emulators that considered this more deeply. This is in part due to the nature of the outputs of terminal emulators vs compilers; like for high throughput, it would be good if the terminal emulator spent as little time as possible processing things that won't ever actually make it to a frame. Compilers are mostly not interactive though, and don't have the luxury of deciding to simply not do something because nobody is looking.
I don’t think that there’s any evidence that this is true. In fact, the performance of the Rust compiler has continually improved over the last five years while _adding_ new features, and not by removing anything. <https://perf.rust-lang.org/dashboard.html>
All a C++ compiler needs to do to get similar results is to have a team that does nothing but performance work, and is actually rewarded for it.
Rust is different. It has a completely different design for translation units and modules that C++ can't adopt without hard breaking changes, due to e.g. the way its macro system works.
There are teams at all of Google, Microsoft, et al that spend time on C++ build performance. At this point, it's not even necessarily just about faster build times. Their builds could be measured in environmental impact at that scale.
[0]: https://www.gem5.org/project/2023/02/16/benchmarking-linkers...
Since VC++ modules support got good enough, that all my side projects only use modules.
A few breaking changes to the language have been made over the years (such as the introduction of async/await), resulting in several different editions of the language. The compiler supports all editions equally, and different crates can even be in different editions, even if they are linked into the same program. That means that code written for Rust 1.0 still works perfectly, and there is no need to update it to the latest edition in order to continue using it even in programs that are written in the latest edition.
For large C codebases, compile time is dominated by lexing and parsing, not optimization passes.
Headers have become too crufty; they tend to only grow over time, resulting in an amplification attack on the lexer and parser.
Lazy parsing is maybe interesting. There's a hazard related to the mangling of lambdas. Carbon has an interesting idea of not constructing an AST.
I think we also need tools to help us automate header refactoring.
It reminds me of the massive Fast kernel headers patchset by Ingo Molnar. I hope it got in. I don't know anything about it, just marvelled at the scope of it.
Imagine 30 years of headers. They tend to grow and grow and become a tangled mess.
Ingo's patch set did not.
I asked what was being done to automate this so that we wouldn't backslide.
I received no response.
That said, I think that work is fascinating. I have an intern starting Monday that's going to start research into automating this.
>I believe what you’re doing is describing something that might be considered an entire doctoral research project in performant terminal emulation as “extremely simple” somewhat combatively…
Casey then spent a weekend and produced a couple thousand line terminal emulator which seemed feature complete(? probably missing some bells and whistles), but blew away Windows performance.
A year later, Microsoft released a patch, roughly incorporating some of Casey's suggestions, and did not credit him for anything[1].
Edit: This[2] seems to be the much bigger Hacker News discussion about the topic instead of [1]. If nothing else, [2] has a repeatedly edited post by the Microsoft author where attitudes are on full display.
[0] https://github.com/microsoft/terminal/issues/10362
Where's the flak, exactly? (FYI, no "c" in "flak")
Linux being an OS where the terminal is as powerful as it is, it's really good to see a high level of attention to detail going to terminal emulators finally. Terminal emulators should be extremely robust and extremely efficient; and finally, they're getting pretty close.
I started using computers with a vt100 terminal back in the 300/1200/2400 baud days though, any modern terminal feels basically infinitely fast in comparison.
On one hand, these latency figures are so small that they're usually difficult to directly perceive; who can even notice the difference between e.g. 25ms and 30ms of latency? That aside though, I believe it likely still negatively impacts the experience. Personally, it feels extremely pleasing to use a pre-compositing desktop environment on a CRT display; the lack of latency is visceral. I am so used to modern machines that it typically feels like it's updating before I have even performed the action.
I think that modern terminal emulators and text editors in modern desktop environments on modern computers with modern peripherals run the risk of having latency bad enough that it actually, if only subtly, impacts typing. USB polling rates, relatively high latency 60hz panels, buffering from compositing, it all does add up over time for sure. Add SSH to the mix and it could legitimately make the difference between something feeling quite usable and something that is irritating.
Earlier this year, I went through every single terminal emulator in Arch's repos and tested using them while compiling AOSP (Android). AOSP's Soong build system does not produce a lot of scrollback, but it does overwrite a line containing the progress ~150k times (once per source file). Alacritty was the only one that was both responsive and didn't eat up a CPU core (staying under 10%). Wezterm, Kitty, Foot, and Konsole stayed responsive, but used 100% CPU on a single core, causing a measurable impact to the build speed (though Konsole 23.08 seems to have regressed and now sometimes locks up). Most of the others became unresponsive or had multiple second delays trying to do anything (eg. ^C, opening new tab, etc.).
This was a fairly straightforward implementation of VT parsing and uses the retained render tree in GTK properly. There is no need to sacrifice CPU for GPU or vice-versa. You can literally do better in both directions while rendering at full frame rates.
Every time I upgrade my desktop I notice measurably higher stats.
And yeah I find it kind of funny this post is by a gnome developer when gnome-terminal is my main example for this.
I find that a bit hard to believe. Try setting your monitor's refresh rate to 30hz instead of 60hz, and the difference is huge. Going from 60 to any higher is also a very obvious improvement to the user experience (not power usage).
But I agree about the terminal emulator case. I've personally never had an issue with terminal performance on Linux, but someone in the comments mentioned that building AOSP in a slow terminal actually slows the build down due to it taking up an entire core just for rendering the display. Although, I think that's an extreme edge case.
But this is about the physical refresh rate of the display. What I meant in the previous comment is about the rendering FPS a program would produce.
But I'm pretty sure that code pre-dates reliable access to vertical sync from applications. Either via XProperty notifications from the compositor, XPRESENT, or actual glSwapBuffers() and/or equiavalents.
The purpose of motion blur is first to make things seems smoother, but mostly to indicate how things are moving on the screen with lower framerates. I think with motion blur you could make text jump way further frame by frame while understanding how text was moving. If you have a word that is bold on screen, you could make it jump half a screen with within a single frame while immediately understanding where it went. Wouldn't that help navigating e.g. huge logs (esp. if they're colored) faster while still having a sense of how far you moved frame by frame?
If your terminal outputs so much that nothing of the screen is the same frame by frame, you could also still get a feeling of how much is actually outputted, by how blurry the screen is.
Never got around implementing it though, curious what you think!
This is complicated by computer graphics being gamma corrected. If you average multiple temporal samples together in sRGB colorspace, brightness will vary unrealistically with motion speed. You need to calculate the motion blur in linear colorspace then convert to gamma colorspace before displaying it. And for best results, the image to be blurred should be HDR, with the blurred version mapped to the lower dynamic range of the screen. This best replicates in-camera motion blur, giving you clear trails behind bright points.
However, any deliberately introduced motion blur is inferior to just increasing framerate. Your eyes already provide motion blur. Anything added by simulation or camera is unrealistic additional motion blur on top of this.
However, with high refresh rate displays ~120Hz+ (and today's performant terminals), I can read logs incredibly well. I can pick phrases out of absolutely constant, non-pausing, streams.
These displays are typically touted as for twitch-like FPS gamers, but they're honestly very useful for consuming a large amount of logs.
There's of course natural blur that comes with this - the quality of the panel decides how clear.
Using Wayland probably helps, output is synchronized by default IIRC.
Related personal observation: A higher refresh rate increases the max scroll speed where text tracked with the eyes is clear enough to read, but decreases the legibility of fast output read by persistence of vision. When I tested this – the store employee was rather puzzled at this use case – 60Hz seemed to be optimal. At 120Hz or higher, scrolling text was (edit) less comprehensible at the same output rate. This applied to various panels and drive modes.
I've never used a high refresh-rate CRT, but I surmise they don't t differ here, since CRT flicker only reduces motion blur that is itself caused by eyes smoothly pursuing onscreen objects made of stationary frames.
Edit: Speaking of scrolling UX, I wish browsers had the option to "overscroll" when paging down, so the unread portion always starts at the top of the screen, as implemented in https://news.ycombinator.com/item?id=36825027. You can make less do this with `-c' (not at all evident from the documentation). A userscript, perhaps...
But for GTK 4, we can do things much differently and faster.
Gnome terminal / libvte is and has always been slow, and alacritty might have good throughput, but sadly is high latency.
A couple of days ago I spent a while trying to figure out why a piece of hardware appeared to be outputting serial data in bursts. After eliminating issues with the underlying hardware and buffering in the USB stack, it (too slowly) dawned on me that the issue was entirely due to how slow terminal emulators have become (this was on a fastish PC).
Updating a screenful of text at 60/120Hz, why/how is it so much slower these days? (and what benefits have we gained in return for the slowdown?).
> "I don't care too much because creating your own terminal is like 20 lines of code these days. People who really care can just create one as easy as configuring an existing one."
These sentences sounds conceited and arrogant.
My commitment to "FOSS" is not up for debate.
I've written in multiple places what it takes to get stuff that fast. If that isn't enough for any of those projects, I'm happy to take money in exchange for my time.
Kitty came closest to being usable for me, but still it involved throwing away ease of access to various things to have a boost to performance that was only something I was told I was getting, not something I was experiencing.
The author of Zutty has a writeup about it: https://tomscii.sig7.se/2020/12/A-totally-biased-comparison-...