HNHacker News
TopNewBestAskShowJobs

ack_complete

815 karma · joined January 8, 2022

submissionscomments
ack_complete··on The state of SIMD in Rust in 2026
>So? AVX-512 shouldn't exist because it can't be retroactively enabled on older hardware? or because Intel can't say goodbye to Skylake? I don't understand the logic.

No one's saying it shouldn't exist, but it's a different story to ship software that requires it unconditionally, particularly if you are targeting the consumer market instead of only servers.

Not sure why you keep referring to Skylake since Intel hasn't shipped a Skylake-based CPU in years, and the most problematic CPUs currently shipping that lack AVX-512 support are several architectural generations ahead of Skylake, for both the P-cores and E-cores.

> Good thing Intel hasn't shipped a viable consumer-level CPU in years either. What should never have happened is Intel removing AVX-512.

Agreed on not removing, but saying that Intel hasn't shipped a viable consumer-level CPU in years is silly. They still ship huge volumes of CPUs, especially in laptops where AMD is still underrepresented.

> At one point, Intel integrated graphics were literally the most common individual GPU in the Steam Hardware Survey - Intel HD 4000 held the #1 spot until the GTX 970 overtook it in late 2015. Does that mean developers shouldn't have targeted discrete GPUs either, just because most survey participants didn't have the hardware they wanted to target?

Targeting a discrete GPU is different than not supporting integrated GPUs at all, which would be analogous. And many games do have to support integrated graphics even if the performance is not great, precisely because they don't aim high enough in the market to be able to ignore iGPUs entirely.

I also mention the Steam Hardware Survey because it tends to overrepresent users with higher end rigs. If you look at the non-gaming market, the hardware level tends to be considerably worse, and as a result a lot of productivity programs still ship plain SSE2 as their baseline required target.

ack_complete··on The state of SIMD in Rust in 2026
AVX2 is much more prevalent than AVX-512 and thus tenable to require. RHEL is switching to x86-64-v3 baseline, for instance, which requires AVX2. AVX-512, on the other hand, has been moving backwards since Intel has not shipped any consumer-level CPUs with it for years now. The Steam Hardware Survey has AVX2 at 95.4% while AVX512F is still far behind at 23.9%.
ack_complete··on The state of SIMD in Rust in 2026
Both. The specific documentation I'm referring to is their optimization guides, such as the Cortex-A72 optimization guide:

https://support.arm.com/documentation/uan0016/a/

This has detailed information on latencies and throughput, which are important when optimizing SIMD code. But ARM doesn't publish optimization guides for all their cores.

On top of that, the cores are often modified in significant ways. Snapdragon CPUs, for instance, have used modified Cortex cores in the past and can have performance differences from the original core.

To be fair, Intel's been slacking a lot on this too, not even bothering to update their own optimization guide for their latest cores. But that's made up for the community mining this information in great detail on sites like uops.info, and also being a lot less x86 cores to deal with. On the ARM side, there's practically not much analogous other than the Apple M1 microarchitecture analysis.

ack_complete··on C's Flexible Integer Sizes Were Not a Design Mistake
The 68000 was a practical example of where mapping int as the "most convenient" size isn't very clear even within a single CPU. The instruction set handled 32-bit integers well, but execution-wise 16-bit integers were faster. Some compilers like Lattice C defined int as int32, while others like Aztec C mapped it to int16.

One result of this is that a fair amount of C code targeted at 68000 systems used short instead int everywhere, and then had to be subsequently updated for CPUs that preferred int32 to int16. Occasionally collateral damage can be seen from this in mysterious comments like "check if the buffer is too int."

ack_complete··on The state of SIMD in Rust in 2026
Pretty sure it only actually currently needs SSE4, since they started compiling it with POPCNT for x86. The official system requirement is Intel 8th gen CPU, but Intel made Kaby Lake and Coffee Lake CPUs that lack AVX.
ack_complete··on The state of SIMD in Rust in 2026
AArch64 definitely has a much more comprehensive baseline than x86-64, but there are some optional extensions that are situationally impactful, including the Crypto extension and some of the newer accumulation / dot product instructions. And unlike Intel, ARM has no portable equivalent to CPUID for querying feature flags and is terrible at documenting which intrinsics require specific FEAT_* flags.

The ARM-based CPU manufacturers make this worse by posting almost no low-level documentation for their CPUs. For basically any mainstream x86 CPU, it's trivial to find documentation listing what ISA level it supports and general execution widths and latencies for common operations. For the majority of ARM CPUs, there's absolutely nothing. ARM only has optimization guides for selected Cortex cores, and NVIDIA published info for their Olympus core. But execution details had to be reverse engineered for Apple M1, and there is nothing for Oryon. This is especially bad for in-order cores, which unfortunately is still relevant because new CPUs are still being shipped with in-order efficiency cores.

ack_complete··on Amiga Screens: A Primer
No, there is actually a difference between an interlaced and non-interlaced signal. A non-interlaced signal has an integral number of lines per field without the half-line that an interlaced signal would have. With a non-interlaced signal, a classic CRT will scan all of the fields with the same alignment instead of interleaving even/odd fields vertically. There was a definite visual difference between a non-interlaced mode and an interlaced mode repeating the same screen for both fields.
ack_complete··on C++26: Trivial infinite loops are no longer undefined behaviour
Yes, but the compiler does have to preserve any observable behavior produced by the call to the standard library function. Being able to omit this inserted yield() by the as if rule would mean that it isn't observable, which would also mean that the compiler could already add or not add it anywhere as needed without changing the behavior of the program. Which would seemingly make the inserted yield() pointless as it would have no effect.
ack_complete··on C++26: Trivial infinite loops are no longer undefined behaviour
What I'm most annoyed at with the variable initialization change is that:

  - It's potentially a performance change in every single function, especially ones that have sizable fixed-size buffers
  - If you have regressions you have to spray [[indeterminate]] everywhere, because there is no coarser way of suppressing it.
  - While the language says unrecognized attributes are ignored, compilers frequently warn on unrecognized attributes. Clang, for instance, currently warns on [[indeterminate]].
  - There is no defined macro name for backwards compatibility.
Which means that libraries are going have to all declare their own macros for [[indeterminate]] and pepper their code with it.
ack_complete··on Subnormal floating-point numbers are expensive on Intel processors
> AFAIK ARM NEON always works in that mode

This was only true for ARMv7 NEON (32-bit). ARMv8 / AArch64 NEON is IEEE compliant.

ack_complete··on Accurate Models of AMD Matrix Cores
Pretty sure the internal precision setting works the same way on the 8087, it's documented in the original 8087 datasheet. Windows sets it to 53-bit so there are no excess precision surprises with doubles, and you could run Windows 95 with an 80386 + 80287.
ack_complete··on On Binary Translation and Its Consequences
> though it seems to be off by default since MSVC2022

This seems to be a typo in the docs, VS2022 is 17.x and still generates volatile metadata. VS2026 is 18.x.

Last time I tested it, the penalty in Prism for running x64 code without volatile metadata was ~25% on Snapdragon X.

ARM64EC code is essentially x64 code pre-translated to ARM64. It's built against the x64 emulation ABI conventions but runs directly as native ARM64 code. Translation thunking conventions allow for cross-calling between ARM64EC and emulated x64 code.

A lot of the OS libs are shipped in Windows 11 ARM as ARM64X, so they're hybrid ARM64EC+ARM64. x86 libs like MSVCRT.DLL do still seem to be compiled with volatile metadata. DUMPBIN /LOADCONFIG reveals if volatile metadata has been included.

ack_complete··on Microcode in Intel's 8087 floating-point chip: the scale instruction
Strange, because .NET was specifically designed to be a JITted environment and taking advantage of SSE2 when available would ordinarily be an advantage of a JIT. But sure enough, .NET 4.0 x86 still uses x87 instructions for math. It's not even good x87, this is surprisingly bad:

  01b2086a 8975e4          mov     dword ptr [ebp-1Ch],esi
  01b2086d db45e4          fild    dword ptr [ebp-1Ch]
  01b20870 d95de4          fstp    dword ptr [ebp-1Ch]
  01b20873 d945e4          fld     dword ptr [ebp-1Ch]
  01b20876 d80dcc08b201    fmul    dword ptr ds:[1B208CCh]
  01b2087c d95804          fstp    dword ptr [eax+4]
And that should be with optimization enabled, I didn't start it from the debugger.

But clearly it has some support for using SSE2 when available, because it does use it for zeroing memory:

  01b20884 0f57c0          xorps   xmm0,xmm0
  01b20887 660fd607        movq    mmword ptr [edi],xmm0
  01b2088b 660fd64708      movq    mmword ptr [edi+8],xmm0
ack_complete··on Microcode in Intel's 8087 floating-point chip: the scale instruction
There is a significant difference between a stack-based ISA and a stack-based bytecode. In bytecode, it's fine or even a requirement to empty the stack between loop iterations. The JIT will then enregister variables across the loop as appropriate.

With x87, however, that causes extra overhead from loads and stores that's best avoided. Unused stack space can be used to cache frequently used variables, but as operations must use ST(0) as one parameter, FXCH instructions must be used to swap around variables. Matching the x87 stack state on entry and exit of the loop is tricky and compilers historically have had trouble doing it. Different FPUs also differed on the efficiency of FXCH so there were often situations where a particular arrangement would double the speed of a routine on one CPU model and halve it on another.

ack_complete··on Show HN: Compute polynomials twice as fast
The answer's all over the place with each successive CPU generation. Originally Intel CPUs had adds faster than multiplies, then both went through the FMA unit so they were the same, then they added a fast FP adder, etc. And current timings on uops.info now show FP fma 4c and mul 3c over two multiply units, and add 2c over two separate addition units.
ack_complete··on Implementation of GCC's Nested Functions (vs. C++ Lambdas)
Sadly, I've never seen a C++ compiler use this (old) technique for lambda reference captures. The main compilers all seem to just use individual references for each capture instead of a single reference to the stack frame, which makes the lambdas with a lot of reference captures more expensive.
ack_complete··on Paint.net 5.2 alpha now runs on Linux
Direct2D has some complex algorithms because it was designed for high-quality antialiased rasterization on DX9-class hardware, particularly tesselation, and suitable not only for vector graphics but also text glyphs. This bleeds over into the API, which is quite complex and somewhat annoying at times. It revives the GDI-style interface of needing to create and select a solid brush object instead of just passing a color, and drawing text with colored spans is embarrassingly involved (it involves writing a custom rendering bridge between Direct2D and DirectWrite).

The other problem is that Wine has only implemented a minimal amount of the Direct2D API to get specific programs working that aren't heavy graphics programs. Two major omissions, the last time I looked, were that Wine's implementation does not support antialiasing at all, and its ArcTo() draws a line.

ack_complete··on GUIs should be fully keyboard-driven
The keyboard support in "Modern" Windows apps is so random. In the new Notepad, the dialog that appears asking if you want to create a new files shows Yes and No buttons, but it doesn't accept Y and N keys, only Esc and Enter. However, the dialog that appears when exiting without saving does accept S and N for Save and Don't Save. These inconsistencies are everywhere in UWP.
ack_complete··on SIMD in the 90s: Programming Intel's Pentium MMX
MMX did help significantly with software 3D rendering. I worked on a software rasterizer that benefited significantly from it. But it didn't take long before even a well-optimized software renderer on a high end CPU couldn't keep up with a low end GPU on a low end CPU.
ack_complete··on SIMD in the 90s: Programming Intel's Pentium MMX
MMX in did not become irrelevant with the GeForce 256. Hardware video decoding was only in its infancy at the time and even the highest end GPUs only supported motion compensation acceleration for decoding only at best. Non-display image processing on the GPU was heavily bottlenecked by very slow read-back speeds from the GPU to the CPU across the AGP bus.
ack_complete··on SIMD in the 90s: Programming Intel's Pentium MMX
> I believe that only the first Pentium 3 core, Katmai, did this.

No, all Pentium 3s as well as the Pentium M. Pentium 4 notably didn't suffer from it, but it of course had many, many, MANY other performance issues.

> I have some faint memories of hearing somewhere that Long Mode didn't support x87. I wonder if it is related to this early info you mention and it being Microsoft specific.

It was VM86 mode that Long Mode didn't support, which was one of the rumored reasons for removing 16-bit NTVDM support (among many). x87 and MMX were always supported in long mode and notably some libraries like OpenBLAS still use x87 instructions. Windows does prohibit use of x87/MMX in kernel mode where the need is negligible.

ack_complete··on SIMD in the 90s: Programming Intel's Pentium MMX
MMX had heavy adoption in image and video processing. IDCT, motion prediction/compensation, YUV/RGB conversion, and alpha blending all benefited from it.
ack_complete··on SIMD in the 90s: Programming Intel's Pentium MMX
I did extensive MMX and SSE2 optimization of audio and video codecs in the 2000s. MMX made a large difference, but it was a pain.

MMX optimization practically required assembly language. The Pentium MMX was an in-order dual pipe CPU, and while compilers supported MMX intrinsics, their code generation for it was abysmal. Visual C++ 6, for instance, would emit code that was 2/3rds register-to-register moves, with values being unnecessarily moved between two registers between each ALU op. This was also a problem with SSE/SSE2 intrinsics. The worst case I saw was the _mm_set_epi8() intrinsic, which was used to construct a 128-bit vector from 16 inputs. When used with all constants, it should have generated a single 128-bit constant load; instead, Visual Studio 2008 generated ~80 instructions to compute it from byte loads. Microsoft didn't fix it until VS2010.

The latency of MMX instructions combined with the in-order dual pipe architecture also made asm loops messy. Simple ops were single-cycle, but multiplies had 3 cycle latency, stores required data an additional cycle in advance, and computed load/store addresses were also needed a cycle in advance. Simply running 2-4 iterations in parallel wasn't an option as there were only 8 vector registers and you'd still get bottlenecks on functional units. Getting peak performance thus often required interleaving loop iterations with special entry and exit code around the loop.

The issue with EMMS is understated. When the CPU switched to MMX, it marked the entire x87 stack as full. If you forgot the EMMS instruction, it wasn't just some strange floating-point bugs that would happen -- the next few x87 floating point calculations could just outright produce NaNs due to FP stack overflow. Furthermore, as these NaNs propagated, the CPU required microcode assists to handle them. So, even if the program didn't crash, an entire calculation domain would get poisoned and slow down by ~20x.

Ultimately, I don't think SSE2 was what killed MMX, but rather SSE, and specifically floating point. MMX not only didn't support floating point, but was also highly concentrated on 16-bit signed integers and secondarily 8-bit unsigned integers. Support for 32-bit integers was particularly lacking and pack/unpack conversions were a bottleneck. Trying to do 3D was cramped because doing so required fixed-point and MMX didn't have the same affordances as DSPs or NEON for rounding or implicit narrow/widen in operations, or even swizzles. SSE, on the other hand, was just straight floating point with standard automatic IEEE rounding and also had important added operations like swizzles and insert/extract. Thus, when 3D took off, SSE was far more useful than MMX.

MMX, however, still remained useful for a while for image and signal processing. SSE2 being twice as wide didn't help algorithms that couldn't use the greater width, such as 8x8 block motion prediction. Additionally, some CPUs at the time only had a 64-bit data path and had to split SSE2 ops, but because of their 4-1-1 decode template, could only decode one such instruction per cycle. The result was that code using the MMX registers could still run noticeably faster than with the SSE registers. This caused some confusion with the 64-bit version of Windows since Microsoft tried to say that x87/MMX shouldn't be used in long mode, but after queries from video processing companies had to document that the x87/MMX registers were enabled and context switched for user mode code.

ack_complete··on Why tiny JPEGs look different in Chrome
I wonder if this is a gamma correction related. JPEG encodes directly in a non linear color encoding (full-range YCbCr), so the partial decoding may be effectively scaling without gamma correction.
ack_complete··on How Claude marks AI-generated content
Moreover, what if you quote text that happens to have been generated by Claude, does that bump up the AI-ness score of your source file or document?
ack_complete··on How Claude marks AI-generated content
We already have one, our Claude setup already requires output to be 7-bit ASCII clean and scans it for such.
ack_complete··on Windows 11's built-in Weather app wastes more than 1 GB of RAM
The stacks are allocated gradually, but there's a minimum few committed pages for each thread and there can be a lot of threads. The graphics driver alone will typically spawn a thread per CPU core unless you specifically disable threading optimizations, and that's separate from the thread pool.
ack_complete··on Windows 11's built-in Weather app wastes more than 1 GB of RAM
Hardware texture compression formats aren't really great for UI and result in visible artifacts, even with a good compressor -- mostly flat areas or gradients divided by sharp edges makes the compression loss more visible. BC3/DXT5 in particular can only encode interpolated colors between two endpoints in each 4x4 block, so two colored regions divided by a line in a third color gets smashed.

Alternate encodings like YCoCg tend to do better.

ack_complete··on Windows 11's built-in Weather app wastes more than 1 GB of RAM
No interior pointers, no embedded arrays, and only a subset of Java definitely seems like an odd target. I would have expected at least two different full languages with acceptable porting overhead to be the MVP.
ack_complete··on Windows 11's built-in Weather app wastes more than 1 GB of RAM
> The thing that the task manager doesn't tell you is whether or not these are shared components. It may be that the 662 MB used by the "Renderer" is shared between many Windows components, so killing that Weather app may not reclaim as much space as you may hope, instead, it would require killing every user of the component, some may be core system apps.

It is the other way around, shared memory causes Task Manager to _underestimate_ memory usage. Task Manager's default views report the process private working set, no shared memory included. This means that 662MB is the _minimum_ amount of memory commit that would be released by ending the process.

> the OS can overcommit

Windows does not allow overcommit by default. It may compress or optimize memory allocations to reduce the physical working set, but the kernel will start failing memory allocations once physical + swap is exhausted regardless.

Page 1 of 10Next →