Qt signal is ten times slower than a virtual function
developernote.com
developernote.com
This is a non-issue and anyone who is worried about it is having a severe case of premature optimization.
I think the "This is a premature optimization" mindset is what leads to slow pieces of software that I like to avoid (React, Qt, electron). But I guess it's fine, as most users don't care as much as I do.
I think it's less true now, but I remember KDE (Qt-based) being slower and buggier than gnome (Gtk-based) a couple of years ago, just to cite one thing. It just felt like Qt-based stuff was in general more of a pain to use & heavier than GTK based stuff. It's really a matter of personal preference here and Qt is a nice project, just that I have some criticism here regarding performance choices. I feel like bad performance decisions tend to snowball and get multiplied when people make library choices and add their own performance issues on top.
Because you included it in a list that otherwise only contained web tech frameworks. It’s not clear to me why you think those are peers of Qt and implied to me that you think they are near equivalents.
What I'm trying to understand: are you talking about when you ask Fusion 360 to do something expensive the UI stops responding for a bit? I certainly see that, but it's not Qt that is the source of problems.
Do you work with huge models?
For hobbyist stuff at least Fusion is by far the least painful option, but the bar is set really, really low. FreeCAD is a gigantic clusterfuck to put it charitably (akin to using an awl to carve a drawing out of cardboard versus pencil and paper). OpenSCAD is neat but it really suffers because OpenSCAD is basically developed and maintained by a single person.
I just dislike people bashing technology for the wrong reason (like it being a "js framework").
In qt+qml usually you connect slots to signals in js, even if they are implemented in c++.
Now a virtual function call, a JIT compiler can devirtualize more often than an AOT compiler. But the vast majority of function calls in your typical C++ app are not virtual.
We're talking about Qt 'signals' though - they're sort of heavy-weight virtual call, reimplemented in C++, not regular function calls.
It should also be noted that Qt signals are far from the optimal way this can be implemented in C++. On top of that, C++ itself makes it more complicated than it needs to be by not providing a bound (to receiver) member function pointer as a primitive; but even then, this can be done in two indirections. A compiler for a language that supports such a facility directly - say, Delphi - can compile it down to a single indirection.
But again... a JS JIT could unroll that loop, leaving with just a couple of instruction per call and no loop.
Being mindful enough to identify when you're feeling the urge to optimize something too soon, will let you step back and optimize what will have the biggest impact once it's finished.
I have seen developers spend hours optimizing some functionality, pick the fastest technique, and it turned out by designing everything to work with their earlier optimizations they made the overall system much slower.
10x the latency of a virtual function call for a signal is very very small beans compared to where you're actually spending CPU cycles, for any reasonable software.
I agree that we shouldn't always think "this is premature optimization". However, we should focus our optimization efforts where they matter, and I really struggle to think of a place where signal latency is really crucial. In any well architected software that's going to be really rare and it makes sense to focus your efforts on optimizing other parts of the software, which seems to be what the Qt devs did.
As for the crazy contrived example they have provided: "Assume I receive 100,000 trades per second from some crypo exchange and if one trade triggers 100 signals than I have 10.000.000 signals per second that is something comparable with the maximum." - They'd have to process the incoming messages on a separate thread (or process) anyway, so they wont drop packets because of the event loop that also processes user input. Then update the GUI once every event loop tick because that's how frequently the UI gets updated anyway.
Is it really though? It's used in embedded systems. I wouldn't call those that can run Qt UIs microcontrollers.
https://www.qt.io/microcontrollers-st .
As per my old boss, the difference in prices when it comes to these microcontrollers and "less powerful" linux socs is negligible that it wasn't worth trying to target them. Not sure if/how the economics changed over the last 2-3 years either.
> Qt Quick Ultralite is designed to serve as a rendering engine for the application's graphical user interface (UI). Its implementation is different from the standard Qt, and it does not depend any Qt libraries such as Qt Core or Qt Gui. Hence Qt Quick Ultralite applications need to use standard C++ containers and classes instead of those from Qt. For example, instead of using QObject or QAbstractItemModel, Qt Quick Ultralite provides a simple C++ API to expose objects and models.
> It does not include the following from the Qt world:
> The Qt C++ APIs. The non-graphical modules such as Qt Core and Qt Network. The Add-on modules such as Qt Multimedia, Qt Bluetooth, and others Qt Addon Modules. The non-MCU embedded platforms such as embedded Linux or the mobile platforms.
https://doc.qt.io/QtForMCUs-2.1/qtul-integratecppqml.html
Which is kind of interesting ...
Qt signals are absolutely fine for like 99% of their usage.
Even if the callback were infinitely fast, if it's called 10e6 times per second, the work already can't take longer than 100 ns on average (10e6 * 100ns = 1e9ns = 1 second). So a more useful way to frame this would be in terms of the ratio of the callback overhead to the work done by the callback (both measured in time), but there's no mention of the latter.
In this particular case, this is a pretty obvious case of 'this is the wrong tool for the job', at least with the stated requirements. Also, at the point where you need to do something 10e6 times per second it's usually appropriate to think about how you might distribute that work across multiple cores.
https://doc.qt.io/qt-5/signalsandslots.html
"Compared to callbacks, signals and slots are slightly slower because of the increased flexibility they provide, although the difference for real applications is insignificant. In general, emitting a signal that is connected to some slots, is approximately ten times slower than calling the receivers directly, with non-virtual function calls."
The documentation literally says ten times slower.
I guess what I'm saying is, the example feels contrived.
A good programmer ought to have read that sentence and instinctively observed that between 100ms and 10ns is a full seven orders of magnitude. For two numbers that at the human level may seem not terribly far away from "zero", there's a lot of space between them.
After doing a bit of tuning I was able to steer the microscope with no visible latency, which means I'm handling user events at ~25FPS or higher and not seeing any high variances. The only problems I have are when the handler that receives a signal takes longer than I have budgeted (IE, more than 1000/25).
The retro gaming community is obsessive about input and display latency and even there anywhere between 5-16ms (16ms being one frame of 240p content) is considered acceptable for even the most hardcore twitch response games.
That’s not saying that other processes aren’t happening faster than that, just that human input and subsequent visual feedback maxes out somewhere between 200-300 times per second and for the vast majority of humans, it is far, far lower.
If you measure the response of individual photoreceptors, it takes 25-50 ms to peak after a flash of appropriately-colored light; the precise number depends on the color and intensity of the light. After that, the signal still needs to propagate through a bunch of visual brain areas, and then even more needs to happen to somehow influence behavior. With everything tuned just so, you can complete that whole process in 100 ms or so, but the conditions have to be perfect; otherwise, 300+ ms between (simple) stimulus and (simple) response is more typical.
Obviously, a lot of this is happening asynchronously, and high refresh rates can help in other ways (e.g., by smoothing out movement), but it astonishing how laggy our visual system is.
As for download progress etc.. I don't think I have ever had to worry about speed of a function call ever - as long as I was leaving it to the event loop take care of it.
You wouldn't want to use ObjC's message sending OR Qt's signalling mechanism in a tight inner loop – hell, you probably don't want to deal with the indirection of the vtable incurred by a virtual function in a tight inner loop. But all of these are more than fast enough for interactive UI work.
objc_msgsend is slower than a virtual function, but like 1.1x-2x, not 10x. (In rare cases it can even be faster due to it being a little more predictor-friendly.)
https://www.mikeash.com/pyblog/friday-qa-2016-04-15-performa...
https://www.mikeash.com/pyblog/objc_msgsends-new-prototype.h...
A cached IMP is faster than a C++ virtual method call (both of which are slower than a non-virtual method call, inline code, or a C function call of course).
[1] https://www.mikeash.com/pyblog/performance-comparisons-of-co...
It's when you start to have to use signals thousands of times per event that it becomes a problem.
Signals for raw user events won't make a difference, but signals as the main mechanism of API interfacing is a problem.
I would say it's a serious problem.
I avoid using signals in general.
You're going to be using up, what, 300 microseconds?
But even then, you're missing the point, because you don't have to use signals to detect the mouse moving if you don't want to.
Like, I’m fairly sure you could have decent latency if your callback function for onMouseMove made a local network call to another process.
Also, how fast do you think download progress should update? Animating it to the next keyframe every second is much more than enough.
Computers work faster than people who click manually.
Sounds like the guy who wanted an event to be fired for each audio sample and required 44100 samples per second. Interestingly, qt signals maybe even fast enough for just that.
Signals are good for rare, unpredictable events. If you have a firehose worth of data coming in, you're not sitting around waiting very often. Your data probably also has some commonality to it, and can be handled in bulk. If you have 100K events coming in, they're probably events of the same type.
Just like if you're writing gigabytes to a file, a program written with the most minimum amount of intelligence isn't going to do it one character at a time.
Also to process important incoming messages on the same thread as the main event loop - where other things like user input, application drawing is handled, is kind of a bad idea.
But for clicking the mouse, scrolling a page, kicking off a network request - you could go up to a full ms and you'd still feel that the app is (in a vague general sense) "performant".
Having written a few kioslaves decades ago, I feel pretty qualified to comment on Qt's signal/slot performance. Similar to select and poll, QSocketNotifier fires off a signal when there's data ready on an fd. Typically you won't see an event per byte with either interface (unless you're deliberately writing obtuse code).
I think that a lot of the "Qt is slow" nonsense came from C++'s awful reputation especially when it came to templates. The reality was that signals/slots were always a slick and performant wrapper around callbacks. It's worth noting that Qt in all its glory is available on QNX (a well known RTOS) and has been since 4.6. Performance is a non-issue.
Insofar as other overhead sensitive environments go, Trolltech ported Qt (including the signal/slot paradigm) to a realtime OS (QNX) and features it prominently in safety critical contexts (e.g. automotive dashboards).
I wonder who thought signals were fast to begin with?
You're surprised signals are fairly low overhead (and that's fine) and and wondered if anyone actually thought they were performant. Having dabbled with Qt off and on for decades I wasn't surprised. The beauty of moc is that it front loads a lot of the magic. Compiling it sucks (less now), running it typically doesn't.(I believe the reason why it doesn't quite work that way in practice is because slot dispatch is ID-based, not pointer-based.)
>In general, emitting a signal that is connected to some slots, is approximately ten times slower than calling the receivers directly, with non-virtual function calls.
But before that their benchmark was for regular signal vs. observer pattern no?
``` void onValueChanged() override { ++m_count; }
void onValueChangedSignal()
{
++m_count;
}
```
QT signal: 00:00:01.841
Virtual function: 00:00:00.179I dunno how someone is skilled enough to do this benchmark but still thinks that this is how you are supposed to use signals and slots...
FWIW, I did some pretty intense real time animation in Qt about 10 years ago, and even on the crap hardware at the time I had no issue using signals to hook into mouse move events (which happen often and quickly sometimes).
But for most of the use-cases like UI interaction handling, the difference is negligible.
The hypocrite in me thinks flutter is my next choice.
Descriptivism vs prescriptivism, or something
I've used it once since then on an embedded vehicle guidance system, and it seemed fine, once you learn all it's quirks.
English is stupid, but there are rules…
…and that isn’t one of them.
“Twice as slow” is perfectly acceptable, “slow as fuck” is a standard measure of time and “ten times slower” means exactly what it says.
I fail to see any relevant insights here.
How do you deduce this without studying it? Or someone else studying it for you?
It's just inherently slower than doing something equivalent that requires zero graphical processing or monitoring of user input. It is common sense to anyone who knows the tech stack that underlies the pixels on the screen.
This is no fault of Qt by any means. It's just the nature of the tech.
Example pleas ? That "mission critical" and "in GUI" is kind of strange... But if you just talking about "human reaction is needed" time then 10 x virtual function call time don't looks so bad...
Just hope they (Qt Co.) hear about that ship with touch interfaces accident. But still, "GUI _reaction_ time" is not a problem while interacting with humans. I would more worry about replacing mechanic/hydraulic subsystems with electronis - just in case EMP bomb in neighborhood. Or last years favourite - pilot's decisions overriding... Because it's captain/pilot/driver who should have final decision to make - nothing better invented so far and not I would not recommended to change that. Some limited Skynet-like incidents are not hard to imagine...
Just hope they (Qt Co.) hear about that ship with touch interfaces accident.
Yeah I'm in camp mechanical user interfaces, but… But still, "GUI _reaction_ time" is not a problem while interacting with humans.
Sure it is. People interact with this crap while driving all the time. They shouldn't but they do. A sluggish interface could absolutely be a major distraction.* road vehicles functional safety
* railway software applications
* electrical / electronic safety-related systems
* medical device software- software life-cycle processes