HNHacker News
TopNewBestAskShowJobs

Cold_Miserable

14 karma · joined May 16, 2023

submissionscomments
Cold_Miserable··on How much slower is random access, really?
Worst case scenario for random access is a multiple level TLB miss, a memory refresh cycle and then a system management mode interrupt all occurring consecutively.
Cold_Miserable··on Denuvo Analysis
From the "analysis" I gather it works by encrypting the .exe and the key's are server-side. The hardware info is used to further encrypt it.

I think the goal should be to fool the checks rather than remove the encryption which would be a nightmare. CPUID can output whatever you want, it just reads MSR's. I'm sure there are possibilities to use kernel drivers to make windows functions also read out whatever you want.

Cold_Miserable··on Why Northern England is poor
The UK is a poor country with a money laundering capital tacked-on.
Cold_Miserable··on Optimizing uint64_t Digit Counting: A Method that Beats Lemire's by up to 143%
mov eax,64 lzcnt r8,rcx sub eax,r8d imul eax,1233 shr eax,12

Accurate to within 1.

Cold_Miserable··on Dividing unsigned 8-bit numbers
Alderlake supports AVX512-FP16. Still only 9.6x faster than div. Most likely reciprocal is just too slow.
Cold_Miserable··on Dividing unsigned 8-bit numbers
This is ~9.6x faster than "scalar".

ASM_TestDiv proc ;rcx out, rdx A, r8 B mov rax,05555555555555555H kmovq k1,rax vmovdqu8 zmm0,zmmword ptr [rdx] vmovdqu8 zmm4,zmmword ptr [r8] vpbroadcastw zmm3,word ptr [FLOAT16_F8] vmovdqu8 zmm2{k1},zmm0 ;lower 8-bit vmovdqu8 zmm16{k1},zmm4 ;lower 8-bit vpsrlw zmm1,zmm0,8 ;higher 8-bit vpsrlw zmm5,zmm4,8 ;higher 8-bit vpord zmm1,zmm1,zmm3 vpord zmm2,zmm2,zmm3 vpord zmm5,zmm5,zmm3 vpord zmm16,zmm16,zmm3 vsubph zmm1,zmm1,zmm3{rd-sae} ;fast conv 16FP vsubph zmm2,zmm2,zmm3{rd-sae} vsubph zmm5,zmm5,zmm3{ru-sae} vsubph zmm16,zmm16,zmm3{ru-sae} vrcpph zmm5,zmm5 vrcpph zmm16,zmm16 vfmadd213ph zmm1,zmm5,zmm3{rd-sae} vfmadd213ph zmm2,zmm16,zmm3{rd-sae} vxorpd zmm1,zmm1,zmm3 vxorpd zmm2,zmm2,zmm3 vpsllw zmm1,zmm1,8 vpord zmm1,zmm1,zmm2 vmovdqu8 zmmword ptr [rcx],zmm1 ;16 8-bit unsigned int ret

Cold_Miserable··on Dividing unsigned 8-bit numbers
Heh? Surely fast convert 8-bit int to 16-bit FP,rcp+mul/div is a no-brainer? edit make that fast convert,rcp,fma (float 16 constant 1.0) and xor (same constant)
Cold_Miserable··on AMD's Ryzen CPUs might be slower in PC games due to a weird Windows 11 bug
How is the built-in Administrator account faster than a named account? I wonder if its a bloatware effect. They need to test with all bloatware removed. No control flow guard, defender, firewall, appearance to performance, no sounds, no background, microcode dll's deleted, mitigations disabled in regedit, storport disabled, every service disabled, every app deleted, edge deleted etc.
Cold_Miserable··on AMD records its highest server market share in decades
Indeed. Intel is trash right now. No AVX512 thus AMD is infinitely faster.
Cold_Miserable··on Zen5's AVX512 Teardown and More
Does Zen 5 have cldemote or senduipi ?
Cold_Miserable··on Fast Multidimensional Matrix Multiplication on CPU from Scratch (2022)
250 is nonsense. 2xFMA per cycle @ ~4.5Ghz = 32*4.5 = ~144 Gflops

Beating cuBlas is unlikely. You probably made a mistake. Last I tested it, it was even better than MKL in efficiency.

Cold_Miserable··on How does a computer/calculator compute logarithms?
Its not possible to automate. Log is one example. You can't just use log(x) but log(x+1). There are other problems like the "zero problem" and its still possible to devise better approximations with other elementary operations such as |x| or ABS(x).
Cold_Miserable··on Intel's anti-upgrade tricks defeated with Kapton tape
Intel makes money selling motherboard chipsets. There's no excuse, just greed.
Cold_Miserable··on Eye exercises for myopia prevention and control: comprehensive systematic review
Even a thicko can figure out this "research". You can't exercise your eyes to change their physical shape.
Cold_Miserable··on The Paradox of x86
Biased article. 14nm was delayed and 10nm was delayed by 4 years.
Cold_Miserable··on AI made these movies sharper – critics say it ruined them
??? The best action sci-fi ever made is a "goofy 80's movie".
Cold_Miserable··on Thermoelectric Cooling
TEC's are no good. The best bet is to bury copper tubing deep underground and have 2 loops.
Cold_Miserable··on The return of the frame pointers
Not interesting. Enter/leave also does the same thing as your save/restore rbp.

Far more interesting I recall there might be an instruction where rbp isn't allowed.

Cold_Miserable··on Avoiding register spills in vectorized code with many constants
"experts in performance".
Cold_Miserable··on What it takes to pass a file path to a Windows API in C++
Windows uses unicode because NTFS uses unicode instead of efficient ascii.
Cold_Miserable··on Breakfast cereal is in long-term decline
Austrian organic dark chocolate granola is delectable.
Cold_Miserable··on Kill dwm.exe on Windows for less input lag and better performance
Turning off "fullscreen optimizations" doesn't always work. It may not even be possible for DirectX 12. If you press Start and the taskbar appears, its not fullscreen. I wrote a simple Direct3D test program and entering fullscreen using alt-enter prevents start from appearing so DirectX 11.0 at least still supports fullscreen, for now.
Cold_Miserable··on Dynamic bit shuffle using AVX-512
I knew about the instruction that Daniel Lemire missed too. It seems useless to me.
Cold_Miserable··on Sam Bankman-Fried fails to dismiss criminal charges related to FTX
US justice system failing again nicely I see. The average man would be rotting in jail and slowly wasting away.
Cold_Miserable··on Do you know how much your computer can do in a second? (2015)
Using fibre-optic cable where light travels in a zig-zag instead of using lasers increases latency too.
Cold_Miserable··on A performance analysis of Intel x86-SIMD-sort (AVX-512)
AVX512 is scarcely faster than radix at 1 million elements while being vastly more complicated.
Cold_Miserable··on Deepmind Alphadev: Faster sorting algorithms discovered using deep RL
They haven't even found an optimal sort for a ymm of 8 or zmm of 16 integers? Just 5 elements? That's completely and utterly useless.
Cold_Miserable··on Taking a new look at fundamental tech for quiet undersea propulsion
Archimedean Dynasty and the Entrox drive is surely your first thought? That awesome underwater game made by BlueByte software.
Cold_Miserable··on Please do not require AVX support for your software
No. AVX512 was disabled because the money-men want to force you to buy the far more expensive Sapphire-Rapids which are the same chip. Same old Intel. Learned nothing. They didn't just fuse it off, they also forced microcode updates that permanently disable it for working alderlakes.
← PreviousPage 2 of 2