https://learn.microsoft.com/en-us/cpp/build/reference/stack-...
373 karma · joined June 12, 2022
https://learn.microsoft.com/en-us/cpp/build/reference/stack-...
> We are running different workloads including rustc perf suite. In general the runtime performance is on par with llvm.
Not what I would've expected!
> has 2-3 wide multiplications compared to 1 for newer methods.
As written, the `shortFloat()` function always calls `uscale()` two times, followed by an optional third call. Each `uscale()` does two wide multiplications (one full 64x64->128 and one 64x64->hi64, in case we want to make that distinction), so that works out to either 4 or 6 wide multiplications in total. I think `shortFloat()` could be rewritten to always do exactly 2 wide multiplications (both 64x64->128) at the cost of some more ALU operations. However, I don't see how that could be further reduced to only one wide multiplication.
EDIT: Going back to look at Zmij's to_decimal, I just realized that it also doesn't do just one wide multiplication in the sense I originally meant, so what you're saying is likely correct in the first place. I overzealously used a different definition of "wide multiplication", which I probably should've realized, given the fact that my numbers are exactly double yours, but alas. My apologies.
Zmij and xjb are in a league of their own. Broadly speaking, dtoa first has to find the shortest decimal representation of the floating-point input, and then format that decimal representation into a string. Zmij and xjb pull far ahead of the others mostly by speeding up the second part of that process.
uscale is quite good without the stringification, as are many other algorithms. I would say uscale's main strength isn't its speed, but rather its simplicity and, more importantly, the fact that it does both formatting and parsing using a single ~11 KiB table, which no other state-of-the-art algorithm offers (although yy comes close).
https://github.com/golang/go/blob/go1.27.0/src/internal/strc...
Make of that what you will.
That being said, the parent commenter is actually referring to other recent proposals as opposed to existing `@Vector` functionality:
- https://www.mdsec.co.uk/2022/04/resolving-system-service-num...
- https://klezvirus.github.io/RedTeaming/AV_Evasion/NoSysWhisp...
- https://whiteknightlabs.com/2024/07/31/layeredsyscall-abusin...
Though I'm not sure which of these techniques, if any, would be most favored by a game DRM as I've never looked into it.
[1] https://learn.microsoft.com/en-us/windows/win32/api/dcommon/...
[2] https://learn.microsoft.com/en-us/windows/win32/api/dwrite/n...
[1] https://github.com/Microsoft/DirectXTK/wiki/throwIfFailed
On x86, yes. There is no performance penalty for misaligned loads, except when the misaligned load also happens to straddle a cache line boundary, in which case it is slower, but only marginally so.
Also, you can link against ntdll.lib directly. Manually calling GetProcAddress for a few functions isn't a tragedy by any means, but in this case, why bother?
Great article nonetheless!
1. Go to https://store.rg-adguard.net.
2. Paste in https://apps.microsoft.com/detail/9n4wgh0z6vhq.
3. Change ring to "Retail".
4. Download the file with an "appxbundle" extension.
5. Install it (might need to enable developer mode for this step; don't remember).
Selected in a reasonable order by default, but can be overridden.
There are three ways to do so:
- Set the SDL_HINT_GPU_DRIVER hint with SDL_SetHint() [1].
- Pass a non-NULL name to SDL_CreateGPUDevice() [2].
- Set the SDL_PROP_GPU_DEVICE_CREATE_NAME_STRING property when calling SDL_CreateGPUDeviceWithProperties() [3].
The name can be one of "D3D11", "D3D12", "Metal" or "Vulkan" (case-insensitive). Setting the driver name for NDA platforms would presumably work as well, but I don't see why you would do that.
The second method is just a convenient, albeit limited, wrapper for the third, so that the user does not have to create and destroy their own properties object.
The global hint takes precedence over the individual properties.
[1] https://wiki.libsdl.org/SDL3/SDL_HINT_GPU_DRIVER
[2] https://wiki.libsdl.org/SDL3/SDL_CreateGPUDevice
[3] https://wiki.libsdl.org/SDL3/SDL_CreateGPUDeviceWithProperti...
One of the developers made an interesting blog post motivating this decision [2] (although some of the finer details have changed since that was written).
There is also a "third party" solution [3] by another one of the developers that enables cross-platform use of SPIR-V or HLSL shaders using SPIRV-Cross and FXC/DXC, respectively (NB: It seems this currently wouldn't compile against SDL3 master).
[1] https://github.com/libsdl-org/SDL/blob/d1a2c57fb99f29c38f509...