All Sliders to the Right
queue.acm.org
queue.acm.org
A popular approach for Python web serving is to launch a number of "workers" (eg via gunicorn, etc), that hang around waiting to serve requests.
Each one of these workers in recently running code (here) idled using ~250MB of non-shared memory. With about 40 workers needed to handle some fairly basic load. :(
Rewrote the code in Go. No need for workers (just using goroutines), and the whole thing idles using about 20MB of memory, completely replacing all those Python workers. o_O
This doesn't seem to be all that unusual for Python either.
That said, Python devs are some of the worst engineers I encounter, so it’s not surprising things are being implemented incorrectly.
There seems to have been some recent-ish work on improving this, though: https://peps.python.org/pep-0683/#avoiding-copy-on-write
https://news.ycombinator.com/item?id=21831951
AFAIK part of the Teams API has been RE'd and is sufficient for writing an IM-only client, so it would certainly be an interesting exercise to do the same as above but for Teams.
The only space free from this appears to be videogames. Gamers regularly complain about "low" framerates that'd put a lot of (at least Windows-based) text renderers to shame. A load screen that you notice is points docked in the review score. I can open Steam, launch CS:GO and get into a game before Teams has finished its morning coffee. [1] https://en.wikipedia.org/wiki/Jevons_paradox
Fundamentally games are actually very different problems amenable to very different kinds of optimization (3d being easier to render than 2d). That’s why all the text you see in 3d games are static menus that are hyper optimized for the limited specific content the ever show in a single font. And the 3d itself has a massive accelerator coprocessor that you need to run at high frame rates. And that accelerator one really works well for extremely parallelizable work which font rendering (and 2D in general) tends not to be (although like I said, there’s research work on making 2d run on GPUs).
Are you perhaps on the Microsoft team responsible for the new Widows Terminal who claimed that drawing text fast is a PhD research-level topic? I remember it didn't turn out well
Consider that Raph and many others like him have spent decades thinking about this problem and working on it off and on and still haven’t fully cracked it. See the PathFinder for example as an attempt to research new techniques to make this work. There’s other work extending this and trying to put it into a real world system. I’ve briefly worked in related areas so I know a little bit about how hard it is and have talked to experts in the GPU space, but it’s not my particular area of interest / expertise. I’d call myself an expert generalist so a lot of the expert-level techniques that domain experts employ can be beyond my skill set (in general vector processing like SIMD and GPUs aren’t areas I’ve delves in too deeply because of how difficult it can be to get really good).
Even if "dynamic high quality text" is hard, I would like to remid you that we're essentially running supercomputers. And most of "dynamic high quality text" on our computers is incredibly static for signinficant amounts of time.
Even on CPU we should be able to render any quantity of any quality of any dynamic text at speeds that cannot be perceived by the human eye. (Well, as it turns out, we do exactly that, even on smart phones)
And, of course, people have done that, with very little effort. Reason? They actually looked at what a modern computer is capable of, and used that instead of hiding behind high brow "oh, it's beyond PhDs my dear Watson for sho"
You may want to reread what I said. I was responding to someone talking about games and game performance in the context of text rendering. And the PhD research topic is how to use the GPU to accelerate text rendering.
You seem to perhaps be misunderstanding what I’m saying and somehow thinking I’m defending
> Microsoft Teams takes 9 seconds to display under a kilobyte of text, including 3 seconds to just display a splash screen)
I’m not. Clearly it shouldn’t take that long to render that text. But achieving a screen full of text at 120fps (or is it even just 60fps) on the highest resolution displays is actually a difficult problem for CPU renderers to achieve (even for the greatly simplified case of a single mono space font). If you think otherwise, you may want to sync up with Raph and publish your contributions to this space. It’s possible I’m getting some of the specific details wrong (again - not my area of expertise) but there’s definitely a push to find HW acceleration within the GPU to achieve those frame rates (and not just text rendering - all 2D rendering is difficult to HW accelerate).
Notice in [3] how CPU renderers are still lower latency than GPU ones because it’s an area of active research that hasn’t yielded winners yet.
Also I’m well aware we have super computers in our pocket. But don’t forget that the screen resolution is really frickin high and DPI scaling on top (so that text is actually legible) adds additional overheads (+ rendering at 60hz instead of 120). There’s also all sorts of tricks being played (eg it’s not going to be rendering the full screen I think because most UIs are retained not immediate).
[1] https://raphlinus.github.io/rust/graphics/gpu/2020/06/01/pie...
[2] https://raphlinus.github.io/rust/graphics/gpu/2020/06/13/fas...
[3] https://medium.com/@raphlinus/inside-the-fastest-font-render...
I had to check the date to make sure that we weren't in the 80s. Accelerated 2D has been the norm since at least the early 90s, and every integrated GPU that's absolutely horrible at 3D performance has a perfectly adequate 2D accelerator.
The hardware isn't the problem. It's the disgustingly bloated software on top.
The computer scientist Niklaus Wirth wrote in his 1995 article "A Plea for Lean Software" that software is getting slower more rapidly than hardware is becoming faster.
Constraints sometimes can lead to more innovation and efficiency. Limiting the level of hardware that the dev team can design and test against could be a good start.
As the demoscene has shown, we are still probably on average a few orders of magnitude away from pushing the limits of what even decades-old hardware is able to do.
For one, “screens are changing too fast and I can’t keep up with it” is a common complaint from less tech-inclined people. Screens transitioning from one state to another state, then staying frozen until user completes decision making process, may not be a universally ideal mode on interaction. For me it is but I’m likely not normal.
For another, I’ve just come across an anecdote yesterday that some developers encounter pressures from bad managers to use phones and tablets in place of a desktop OS, or resistance against engineering workstation upgrade plans, with “computers are always slow and that’s normal” as a rationale.
There must be elements that are normalizing slowness on computers, even preferences over fast computers and software. Blaming “bad” programming won’t work if there are incentives against improving it: bad incentives must be identified and removed first.
And I've noticed that I'm not immune to it myself. I sometimes catch myself wondering "what am I not doing right?" when I've made a thing that is perceptively instant compared to a competitor taking 5 to 10 seconds to do the same thing.
Indeed, it seems to contradict the studies that companies like Amazon have done to show that "every 300ms of latency leads to X% of users bouncing out." I've only seen people complain about the most egregious of performance issues. People seem to have been trained to expect software to be slow and have come up with an ex post facto rationalization for why that is the case.
It's particularly frustrating because I've only ever worked on small teams, which already receive a lot of bias of the form, "what could a handful of people do compared to the might of FAANG?” When often the reality is that the organizational structure of large corporations frequently leads to the lack of cooperation necessary to make quality software.
I think a likely culprit might be the massive amounts of logging that I don't do, either of user or system activity. When using tools like Logcat to try to debug issues on-device, I notice hundreds of log entries being recorded per second from all corners of different apps and system layers before I finally re-learn how to use the arcane filtering system to get to the stuff from only my app. And that's just system level logging. Throw on top all the Google Analytics stuff some apps do for every click, every page transition, every single thing a user could do, and I suspect it adds up to quite a bit of UI latency, network bandwidth saturation (some ostensibly simple sites become unusable on bad network connections while my own apps seem barely affected), and energy usage.
I'd speculate that this has more to do with a lack of feedback in the UI than anything.
Looking at the bigger picture:
- Moore's law is slowing down
- The cost for ever shrinking transistors is exploding
- Energy costs are soaring
- Planetwide, we need to get rid of non renewable power sources ASAP
Thus, efficient hardware and software should be on everyone's priority list, even if developers are expensive.
You can’t hot loop optimize a Pinto into a Ferrari, you have to be mindful of performance from the start. Unfortunately performance-related compromises net you bad code reviews and dismissive attitudes from your peers
And yet, we still see the common misconception that CPU speed == clock speed.
> It is true that the days of upgrading every year and getting a free—wait, not free, but expensive—performance boost are long gone, as we're not really getting single cores that are faster than about 4GHz.
This made me go check on the date of the article. Two CPUs? But I suppose for CDN servers that makes sense, they're simply big caches.
Most people dealing with hardware don’t refer to “CPUs” anymore, because there’s a lot of confusion around sockets, cores, and hyperthreads, so it’s better to be more specific in any conversation.
We have a bunch of Go and Rust tooling coming out for JS, for example.
If you just pull teeth instead of fixing them you need less dentists. Everybody wins.
Why bother with foundations if you can just tape together whatever works and call it a day? If it crashes you just add tape.
But the comparison here is more like the other way around, isn't it? Fixing the problem by just throwing more resources at it is like making up for an inefficient construction workflow by hiring more construction dudes.
Compare: Computer game from 1995 to computer game from 2023, to chat client from 1995 to chat client in 2023.
One of these demonstrates many orders of magnitude improvement consistent with the improvements in the hardware, and it isn't the chat client.