HNHacker News
TopNewBestAskShowJobs

icelusxl

37 karma · joined December 2, 2021

submissionscomments
icelusxl··on Z80 – The 1970s Microprocessor Still Alive (2021)
ARM1 has no cache memory, no hardware multiply and division, no MMU and no cache.
icelusxl··on SIMD in the 90s: Programming Intel's Pentium MMX
Compilers like GCC and Clang treat C's "long double" type by default as 80-bit wide and result in x87 generated code. This can be overridden to use either 64-bit or 128-bit floating point values.
icelusxl··on Win16 Memory Management
Memory mapping/bank switching was fairly common on 8-bit and 16-bit systems, where a small memory window was used to select different memory banks, allowing a program to access more memory in chunks.

Game consoles like NES, SNES and Game Boy had additional hardware built in the cartridge to support memory mapping/bank switching.

For PCs, EMS (memory) provided a similar concept. It reserved a 64 kB window divided in 16 kB pages in the first 1 MB and allowed to map up to 32 MB.

icelusxl··on Windows GOG DOS Games on M-Series Macs
> There are only a handful of different instructions that account for 90% of all operations executed, and, near the top of that list are addition and subtraction. On ARM these can optionally set the four-bit NZCV register, whereas on x86 these always set six flag bits: CF, ZF, SF and OF (which correspond well-enough to NZCV), as well as PF (the parity flag) and AF (the adjust flag).

> Emulating the last two in software is possible (and seems to be supported by Rosetta 2 for Linux), but can be rather expensive. Most software won’t notice if you get these wrong, but some software will. The Apple M1 has an undocumented extension that, when enabled, ensures instructions like ADDS, SUBS and CMP compute PF and AF and store them as bits 26 and 27 of NZCV respectively, providing accurate emulation with no performance penalty.

https://dougallj.wordpress.com/2022/11/09/why-is-rosetta-2-f...

icelusxl··on Let's compile Quake like it's 1997
Yes, Turbo Pascal 5.0 introduced those features in 1988.

https://www.youtube.com/watch?v=UNx4dxXptUg

icelusxl··on macOS 27 won’t be supporting Intel anymore
Virtualize macOS 26 for testing purposes: https://eclecticlight.co/2025/01/21/what-can-you-do-with-vir...
icelusxl··on Too much discussion of the XOR swap trick
It still is. The CPU's register renamer can detect these instructions to not have data dependencies and can zero the register itself. It doesn't send the instruction to the execution engine meaning they use no execution resources and have zero latency.
icelusxl··on Microsoft hasn't had a coherent GUI strategy since Petzold
The "Performance Improvements in .NET" blog lists the new JIT support for instruction sets each year.

https://devblogs.microsoft.com/dotnet/performance-improvemen...

icelusxl··on AVX Bitwise ternary logic instruction busted
Also supported in .NET 9

* https://devblogs.microsoft.com/dotnet/performance-improvemen...

* https://github.com/dotnet/runtime/pull/91227

icelusxl··on Bit Twiddling Hacks (2009)
Maybe http://0x80.pl -- this site features mainly SIMD/SWAR code.
icelusxl··on Chrome: Heap buffer overflow in WebP
Seems to be reported by Apple and looks a lot like this security update: https://support.apple.com/en-us/HT213906
icelusxl··on Popcount CPU instruction (2019)
> "rotate and mask" series of instructions

Sounds a lot like the PowerPC's rlwinm instruction.

The PowerPC 600 series, part 5: Rotates and shifts

https://devblogs.microsoft.com/oldnewthing/20180810-00/?p=99...

icelusxl··on RISC-V Instructions
"rep stosb" has been optimized since Ivy Bridge CPUs.

> Beginning with processors based on Ivy Bridge microarchitecture, REP string operation using MOVSB and STOSB can provide both flexible and high-performance REP string operations for software in common situations like memory copy and set operations.

> Beginning with processors based on Ice Lake Client microarchitecture, REP MOVSB performance of short operations is enhanced. The enhancement applies to string lengths between 1 and 128 bytes long.

* https://www-ssl.intel.com/content/www/us/en/architecture-and...

* https://stackoverflow.com/a/33485055

icelusxl··on Dynamic bit shuffle using AVX-512
* Visualization: https://www.officedaytime.com/simd512e/

* Book: https://link.springer.com/book/10.1007/978-1-4842-4063-2

icelusxl··on Amazon EC2 M1 Mac Instances
€49/month at Hetzner.

https://www.hetzner.com/dedicated-rootserver/matrix-apple