How Rust 1.64 became faster on Windows
tomaszs2.medium.com
tomaszs2.medium.com
The subtle win is the space savings, no more Jackson pollock inlining
I'd like compiler writers to embed a 'default profile' into the compiler, which uses data from as much opensource code as they can find all over github etc.
This default profile will improve the performance of lots of libraries that everyone uses, and will probably still help closed source code (since it will probably be written in a similar style to opensource code).
* application type (e.g. client, server, batch process, parser, etc.) * target architecture, vendor, model, etc. * target resources like RAM, HD Types, Network interfaces, etc.
I would think you could get very close to an actual PGO level of performance with just a handful of parameters and lot of data.
What would be the point? The whole thing about PGO is that it measures which paths of _your_ code are "hot".
(e.g. trying to convert a 16-bit unsigned integer into a 32-bit signed integer can't fail, that always works so its error type is Infallible, whereas trying to convert a 32-bit signed integer into an unsigned one clearly fails for some values, that's a core::num::TryFromIntError you need to handle)
So we're left only with errors which don't happen. But who says? On my workload maybe the profile image file doesn't exist 0% of the time since I'm actually making the image files, so of course they exist, but in your workload the user gets to specify the filename and so they type it wrong about 0.1% of the time, and in somebody else's workload the hostile adversary spews nonsense filename values like "../../../../../etc/passwd" to try to exploit bugs in some PHP code from 15 years ago, so they see almost 10% errors. What would we learn from a "general profile"? Nothing useful.
$ process Some Image Name.png
Could not find file “Some”
$ process “Some Image Name.png”
Done.
... Urg, that.
If I ever implement a bespoke file system format, it is going to be encoding-level impossible to represent file names with spaces. Not FAT-style[0] "the spec says to replace that with a underscore" or something, but more "the on-disk character encoding does not contain any sequence of bits that represents space".
0: (non-ex-)FAT stores filenames in all caps, but the data of disk is ASCII, so you can just write lowercase letters in the physical directory entries. (I've seen at least one FAT implementation that actually uses that to 'support' lowercase filenames.)
(I had someone ask for the possibility of being able to choose older versions of Unicode in my library to handle his use case with clusters in a terminal app, but on further investigation to what he was trying to do, I discovered that it was a misunderstanding about how grapheme clusters work and in fact would not do what he wanted it to do.)
And lots of your code has similar hot paths to everyone elses code. It turns out that `for x in pixels { }` is probably going to be a hot loop... But `for x in serial_ports { }` probably isn't a hot loop...
I'm imagining for example a 'profile server', which anyone can upload profiler data to, and that the compiler queries to get profile data for any given file it wants to compile.
The most modern ones that is (Java, .NET, Android).
The difficulty with "as much open source code as they can find" is that we need to execute the code to make a profile. And unless we're running the code under real-world conditions, there's no guarantee that we'll generate a useful profile. So we need to be a little careful about which code we look at from a performance perspective. Even when we have a profile, it's a count of branches taken for the specific code that was compiled, and it's not normally applicable to either a different version of the compiler or any input that's not identical to the input used for profiling. With link-time optimisations, even a "common" profile for library code isn't necessarily going to be useful: which bits of a library we'll try to inline will vary according to the code that's calling it.
The default profile is a nice hack. We do this by default for C++ builds at [company], it works great. Teams that care can build a custom profile which performs better, but most don't.
> I'd like compiler writers to embed a 'default profile' into the compiler, which uses data from as much opensource code as they can find all over github etc.
Working out how to build, let alone profile all that code is no joke. And the result will be large, and maybe not that much overlap with the average program. As a sibling points out, maybe using ML to recognize patterns instead of concrete code would help?
I'd settle for profiling of the standard library. In an ecosystem like Rust, per-crate default profiles that you could stitch together would be amazing.
What sort of work is OS specific, or language specific?
I’ve used PGO before, but I’m not familiar with the details.
FTA: But there is one problem: PGO was up until now available only on Linux.
I think they couldn’t use a profile generated on Linux because of differences in ABI and standard library.
I also think generating a profile is OS dependent because you want it to not have much of a performance impact.
That would be a lot of work to set up!
Not to be snobby, but I wish everyone used a standard date format.
Edit: I'd like to add that if the 10-20% mentioned is measured on the benchmark that was used to do the pgo, then that figure might indeed not be representative of the real gain.
If you manage to overfit against that, it's still probably an amazing general purpose solution.
From what I understood: Criterion was the gold standard, but there is a built-in benchmark suite but that is only supported in Nightly.
What’s the difference?
Pros for Criterion over the stdlib: https://github.com/bheisler/criterion.rs#features
Downsides of Criterion: https://bheisler.github.io/criterion.rs/book/user_guide/know...
Though there may be a lot of opportunity with some database subsystems that have a more consistent usage pattern.
Edit: also, PGO is closely related to JIT techniques, which are based on current runtime information rather than profiles generated a long time ago on a workload that may or may not be representative.
Have you actually seen otherwise?
Look at the [dead] section
So compiling Rust code on Windows is faster with 1.64 than 1.63.