HNHacker News
TopNewBestAskShowJobs

e4m2

373 karma · joined June 12, 2022

submissionscomments
e4m2··on A Design Space Exploration of Async/Await
Likewise, on Windows, it's 1 MB of reserved memory but only 4 KB of initially committed memory.

https://learn.microsoft.com/en-us/cpp/build/reference/stack-...

e4m2··on Rust is tier-1 language at Microsoft
From the linked Zulip thread (https://rust-lang.zulipchat.com/#narrow/channel/131828-t-com...):

> We are running different workloads including rustc perf suite. In general the runtime performance is on par with llvm.

Not what I would've expected!

e4m2··on Go 1.27
Right, that's basically what I was trying to say (in so many words). I learned a lot from your dtoa blog posts and Zmij's implementation. Thank you!

> has 2-3 wide multiplications compared to 1 for newer methods.

As written, the `shortFloat()` function always calls `uscale()` two times, followed by an optional third call. Each `uscale()` does two wide multiplications (one full 64x64->128 and one 64x64->hi64, in case we want to make that distinction), so that works out to either 4 or 6 wide multiplications in total. I think `shortFloat()` could be rewritten to always do exactly 2 wide multiplications (both 64x64->128) at the cost of some more ALU operations. However, I don't see how that could be further reduced to only one wide multiplication.

EDIT: Going back to look at Zmij's to_decimal, I just realized that it also doesn't do just one wide multiplication in the sense I originally meant, so what you're saying is likely correct in the first place. I overzealously used a different definition of "wide multiplication", which I probably should've realized, given the fact that my numbers are exactly double yours, but alas. My apologies.

e4m2··on Go 1.27
The upstream fmtlib dtoa-benchmark integrates uscale (https://fmtlib.github.io/dtoa-benchmark/results/). It uses C code from Russ Cox's original fpfmt repository (https://github.com/rsc/fpfmt/tree/main/bench/uscalec), which is slightly different from the Go code upthread.

Zmij and xjb are in a league of their own. Broadly speaking, dtoa first has to find the shortest decimal representation of the floating-point input, and then format that decimal representation into a string. Zmij and xjb pull far ahead of the others mostly by speeding up the second part of that process.

uscale is quite good without the stringification, as are many other algorithms. I would say uscale's main strength isn't its speed, but rather its simplicity and, more importantly, the fact that it does both formatting and parsing using a single ~11 KiB table, which no other state-of-the-art algorithm offers (although yy comes close).

e4m2··on Go 1.27
Not mentioned: Floating-point parsing and formatting now uses Russ Cox's uscale algorithm.

https://research.swtch.com/fp

https://github.com/golang/go/blob/go1.27.0/src/internal/strc...

e4m2··on Mojo 1.0
https://news.ycombinator.com/item?id=48057901#48068126

Make of that what you will.

e4m2··on Everyone should know SIMD
FYI you linked to a really old version of the GCC documentation. Google apparently loves those old docs, so they often show up near the top of search results despite being ancient. For posterity, here's the latest version: https://gcc.gnu.org/onlinedocs/gcc-16.1.0/gcc/Vector-Extensi....
e4m2··on Everyone should know SIMD
Historically, it was a bit more complex than that, and applies to more than just AVX-512: https://gist.github.com/rygorous/32bc3ea8301dba09358fd2c64e0....
e4m2··on Zig Structs of Arrays (2024)
Zig vectors do not necessarily force data into SIMD registers; a scalar implementation would work equally well. This is not just a theoretical argument, because Zig code that uses `@Vector` also has to compile for architectures that do not have SIMD instructions.

That being said, the parent commenter is actually referring to other recent proposals as opposed to existing `@Vector` functionality:

https://codeberg.org/ziglang/zig/issues/32032

https://codeberg.org/ziglang/zig/issues/35376

e4m2··on Windows 7 marketshare jumps to nearly 10% as Windows 10 support is about to end
It's not designed for IoT devices per se, the naming is just terrible. A comparison to OpenWrt is not warranted here, although to reiterate, the naming is terrible.
e4m2··on Windows 7 marketshare jumps to nearly 10% as Windows 10 support is about to end
You're "supposed" to acquire LTSC through non-official means, not using the evaluation ISO.
e4m2··on Windows 7 marketshare jumps to nearly 10% as Windows 10 support is about to end
See also: https://betawiki.net/wiki/Windows_8_build_8172
e4m2··on Microsoft is open sourcing Windows 11's UI framework
Explorer uses XAML Islands. Parts of it are WinUI, while the rest is still Win32.
e4m2··on Microsoft is open sourcing Windows 11's UI framework
C++
e4m2··on 7-Zip for Windows can now use more than 64 CPU threads for compression
"Good enough" is not good enough.
e4m2··on Windows x86-64 System Call Table (XP/2003/Vista/7/8/10/11 and Server)
There's a million and one ways to do it, here's just some of the ones I remember:

- https://www.mdsec.co.uk/2022/04/resolving-system-service-num...

- https://klezvirus.github.io/RedTeaming/AV_Evasion/NoSysWhisp...

- https://whiteknightlabs.com/2024/07/31/layeredsyscall-abusin...

Though I'm not sure which of these techniques, if any, would be most favored by a game DRM as I've never looked into it.

e4m2··on The radix 2^51 trick (2017)
On modern enough x86 CPUs (Intel Broadwell, AMD Ryzen) you could also use ADX [1] which may be faster nowadays in situations where radix 2^51 representation traditionally had an edge (e.g. Curve25519).

[1] https://en.wikipedia.org/wiki/Intel_ADX

e4m2··on Growing Buffers to Avoid Copying Data
According to the standard `realloc(NULL, size)` should already behave like `malloc(size)`. You shouldn't need that special case unless you're working on a system with a very buggy/non-compliant libc.
e4m2··on Better text rendering in Chromium-based browsers on Windows
It just uses GDI_CLASSIC for DWRITE_MEASURING_MODE [1] and DWRITE_RENDERING_MODE [2] in that case. No actual GDI in sight.

[1] https://learn.microsoft.com/en-us/windows/win32/api/dcommon/...

[2] https://learn.microsoft.com/en-us/windows/win32/api/dwrite/n...

e4m2··on Tilde, My LLVM Alternative
https://github.com/RealNeGate/Cuik/blob/5c6f6ef9bfa983eb358a...
e4m2··on C: Simple Defer, Ready to Use
The much dreaded Annex K functions are perhaps the worst possible example of an attempt at "fixing" anything safety related in C. A waste of ink.
e4m2··on C++ exception performance three years later
Are there any benchmarks of Windows EH? The implementation is very different from DWARF/SJLJ EH and it would be interesting to see how it fares. I've seen some pretty exceptional claims for both sides of the argument (e.g. "[...] has no code impact for x64 native, ARM, or ARM64 platforms" [1]), but there's rarely if ever any data to back this up.

[1] https://github.com/Microsoft/DirectXTK/wiki/throwIfFailed

e4m2··on ChibiHash: Small, Fast 64 bit hash function
> Is a single MOV instruction still fast when the 8 bytes begin on an odd address?

On x86, yes. There is no performance penalty for misaligned loads, except when the misaligned load also happens to straddle a cache line boundary, in which case it is slower, but only marginally so.

e4m2··on Windows Memory Mapped File IO
I think you should be able to get rid of most of the undocumented API usage with the newer CreateFileMapping2/MapViewOfFile3 APIs. Though that does require a higher minimum OS version and the crucial NtExtendSection function still has no documented equivalent as far as I can tell, so it's kind of a wash.

Also, you can link against ntdll.lib directly. Manually calling GetProcAddress for a few functions isn't a tragedy by any means, but in this case, why bother?

Great article nonetheless!

e4m2··on The Byte Order Fallacy
Be aware that if you actually want to do as the article prescribes, don't just copy and paste -- you shan't take anything at face value in C: https://news.ycombinator.com/item?id=31718292.
e4m2··on Comic Mono
There is also Comic Code by the same author: https://tosche.net/fonts/comic-code. Feels like a more polished version of Comic Mono (but it is paid, too).
e4m2··on Keyhole – Forge own Windows Store licenses
You don't need this exploit. You could use a media player that doesn't need MS codec packs, but assuming this is not an option:

1. Go to https://store.rg-adguard.net.

2. Paste in https://apps.microsoft.com/detail/9n4wgh0z6vhq.

3. Change ring to "Retail".

4. Download the file with an "appxbundle" extension.

5. Install it (might need to enable developer mode for this step; don't remember).

e4m2··on SDL3 new GPU API merged
> is this possible or is the backend selected e.g. based on the OS?

Selected in a reasonable order by default, but can be overridden.

There are three ways to do so:

- Set the SDL_HINT_GPU_DRIVER hint with SDL_SetHint() [1].

- Pass a non-NULL name to SDL_CreateGPUDevice() [2].

- Set the SDL_PROP_GPU_DEVICE_CREATE_NAME_STRING property when calling SDL_CreateGPUDeviceWithProperties() [3].

The name can be one of "D3D11", "D3D12", "Metal" or "Vulkan" (case-insensitive). Setting the driver name for NDA platforms would presumably work as well, but I don't see why you would do that.

The second method is just a convenient, albeit limited, wrapper for the third, so that the user does not have to create and destroy their own properties object.

The global hint takes precedence over the individual properties.

[1] https://wiki.libsdl.org/SDL3/SDL_HINT_GPU_DRIVER

[2] https://wiki.libsdl.org/SDL3/SDL_CreateGPUDevice

[3] https://wiki.libsdl.org/SDL3/SDL_CreateGPUDeviceWithProperti...

e4m2··on SDL3 new GPU API merged
The current SDL GPU API does not intend to use this shader language. Instead, users are expected to provide shaders in the relevant format for each underlying graphics API [1], using whatever custom content pipeline they desire.

One of the developers made an interesting blog post motivating this decision [2] (although some of the finer details have changed since that was written).

There is also a "third party" solution [3] by another one of the developers that enables cross-platform use of SPIR-V or HLSL shaders using SPIRV-Cross and FXC/DXC, respectively (NB: It seems this currently wouldn't compile against SDL3 master).

[1] https://github.com/libsdl-org/SDL/blob/d1a2c57fb99f29c38f509...

[2] https://moonside.games/posts/layers-all-the-way-down

[3] https://github.com/flibitijibibo/SDL_gpu_shadercross

e4m2··on Techniques for safe garbage collection in Rust
Just to add one more data point, the .NET GC is also written in C++.
Page 1 of 3Next →