HNHacker News
TopNewBestAskShowJobs

vient

315 karma · joined October 20, 2020

reverse engineer
submissionscomments
vient··on Ask HN: Why hasn't x86 caught up with Apple M series?
Weren't we comparing CPUs though? Those Blender benchmarks are for GPUs.

Here is M4 Max CPU https://opendata.blender.org/devices/Apple%20M4%20Max/ - median score 475

Ryzen MAX+ PRO 395 shows median score 448 (can't link because the site does not seem to cope well with + or / in product names)

Resulting in M4 winning by 6%

vient··on The first Media over QUIC CDN: Cloudflare
Oh, I see - hard refresh consistently shows HTTP/2 but after one or two soft refreshes it becomes HTTP/3 for me until next hard refresh.

Edit: it is always second soft refresh for me that starts showing HTTP/3. Computers work in mysterious ways sometimes.

vient··on The first Media over QUIC CDN: Cloudflare
Limited to macOS? Does not reproduce in FF 141 and 142 on Windows.
vient··on Test Results for AMD Zen 5
AMX is indeed a very strong feature for AI. I've compared Ryzen 9950X with w7-2495X using single-thread inference of some fp32/bf16 neural networks, and while Zen 5 is clearly better than Zen 4, Xeon is still a lot faster even considering that its frequency is almost 1GHz less.

Now, if we say "Zen5 is the leading consumer CPU for AI" then no objections can be made, consumer Intel models do not even support AVX-512.

Also, note that for inference they compare with Xeon 8592+ which is the top Emerald Rapids model. Not sure if comparison with Granite Rapids would have been more appropriate but they surely dodged the AMX bullet by testing FP32 precision instead of BF16.

vient··on Transition to using 16 KB page sizes for Android apps and games
True, but that has nothing to do with tagged pointers.
vient··on Transition to using 16 KB page sizes for Android apps and games
Memory page size should be transparent for tagged pointers (any pointers, really), I don't see how they can be affected. You have an object at address 0xAB0BA, does the size of underlying page matter?
vient··on SIMD.info – Reference tool for C intrinsics of all major SIMD engines
Would also be nice to remove empty categories from tree view. For example, right now you can uncheck VSX and still see "Memory Operations - VSX Unaligned ..." full of empty tags.
vient··on Are polynomial features the root of all evil? (2024)
But what's the point in acknowledging numerical issues outside of [-1,1] if polynomials do not even work there, as author explicitly notes?
vient··on What the hell is a target triple?
> Kalimba, VE

> No idea what this is, and Google won’t help me.

Seems that Kalimba is a DSP, originally by CSR and now by Qualcomm. CSR8640 is using it, for example https://www.qualcomm.com/products/internet-of-things/consume...

VE is harder to find with such short name.

vient··on OpenVINO AI effects for Audacity
They have some NVIDIA support in the form of external project: https://github.com/openvinotoolkit/openvino_contrib/tree/mas...
vient··on WebGPU Tech Demo
For me fails on identical setup with

  GET https://gnikoloff.github.io/webgpu-sponza-demo/assets/Sponza.bin
  net::ERR_CONNECTION_CLOSED 206 (Partial Content)
vient··on Common misconceptions about compilers
LLVM_USE_SPLIT_DWARF may help with this, some recent measurements: https://www.tweag.io/blog/2023-11-23-debug-fission/#an-examp...
vient··on The Wes Cook Archive
Sounds right for a panic co-founder.
vient··on Honey, I shrunk {fmt}: bringing binary size to 14k and ditching the C++ runtime
It is a featureful formatting library, not simply a library for slow printing of ints and strings without any modifiers. You can't create a library which is full of features, fast, and small simultaneously.
vient··on Counting bytes faster than you'd think possible
Wow, changing `count` type from uint64_t to uint32_t or int radically changes results - now gcc gets 26500 and clang gets 25000, that's just 1.7 times slower than current best solution.

So you can get 25k with following code, clang -Ofast -std=c++17 -march=native -static

    #include <iostream>
    #include <cstdint>
    #include <sys/mman.h>
    #include <unistd.h>

    int main() {
      auto file_lo = (const uint8_t*)mmap(0,250000000ull,PROT_READ,MAP_PRIVATE|MAP_POPULATE,STDIN_FILENO,0);
      int count = 0;
      for (uint32_t i = 0; i < 250000000; ++i) {
        if (file_lo[i] == 127) {
            ++count;
        }
      }
      std::cout << count << std::endl;
      _exit(0);
      return 0;
    }
vient··on Counting bytes faster than you'd think possible
Nice. Did some quick tests with your code on site, got score of ~34000 - best solution is around 14700, so this one is only 2.3 times slower.

Used clang with -Ofast -march=native -static. Funnily, gcc gets only 54000 with the same options, 1.6 times slower.

vient··on Counting bytes faster than you'd think possible
Note that "standard C++" solution uses std::cin while optimized one uses mmap - completely different things, a lot of speed comes just from that. Would've been nice to compare with solution having optimized input and otherwise standard summing loop.
vient··on Optimizing a bignum library for fun
> I've been wanting a language with that feature

Python, which realization CPython is mentioned in the article, has arbitrary precision integers.

vient··on Flame Graphs: Making the opaque obvious (2017)
I like speedscope.app for viewing flamegraphs. It is more interactive than traditional SVG flamegraphs, and what is relevant here is a "sandwich" view - basically a sorted list of all functions, you see what function was spent the most time in, click on it and see all calling stack traces like a mini flamegraph, filtered and centered on this function. Speedscope supports several popular trace formats, really useful.
vient··on Llama.ttf: A font which is also an LLM
And, as expected, the very first reference in the post is to the Tom7's video. Of course it would be.
vient··on CRIU, a project to implement checkpoint/restore functionality for Linux
There is a QEMU fork used by Nyx fuzzer, may be interesting to you https://github.com/nyx-fuzz/QEMU-Nyx

Basically, for the fuzzing purposes speed is paramount so they made some changes to speed up snapshot restoring. Don't know the limitations but since it is used to fuzz full operating systems, there should not be many.

I believe it should be faster than forking because why even patch QEMU otherwise.

vient··on A venture capitalist walks into a bar
> don't randomly hit up bloggers

Not sure if satire or you just don't know who lcamtuf is.

vient··on Intel Meteor Lake's NPU
Looking at OpenVINO repo activity, seems that most of them got relocated and still work at Intel.
vient··on Intel Meteor Lake's NPU
It worked only on first batches/microcode versions of hybrid CPUs, since then Intel had completely disabled AVX-512 on them.
vient··on Intel Meteor Lake's NPU
Not the best time for testing OpenVINO at least, they plan to release NPU support in 2024.1 as I understand, current version 2024.0 has very limited NPU support if any.
vient··on Predictive CPU isolation of containers at Netflix (2019)
Also sched-ext which seems close to be mainlined and is already a default scheduler in CachyOS:

https://github.com/sched-ext/scx

vient··on Llamafile 0.7 Brings AVX-512 Support: 10x Faster Prompt Eval Times for AMD Zen 4
Huh, seems you are right. I've read about "not really real" AVX-512 implementation in Zen 4, and also saw that on my workload turning on AVX-512 on Zen 4 indeed gives almost nothing (compared to x2 speedup on Intel). But now that you've written it, I understood that my only test was very specific, full of 32-bit FMAs, and I was comparing Zen 4 with Skylake-X which has 2 FMA units.
vient··on Llamafile 0.7 Brings AVX-512 Support: 10x Faster Prompt Eval Times for AMD Zen 4
Not clear how AVX-512 can provide 2x speedup on Zen 4, even more 10x (if they are comparing with AVX2 which is obvious assumption). Zen 4 does not really have proper AVX-512 units, and 10x means that there was no vectorization at all before?
vient··on The rev.ng decompiler goes open source
Huh, for me as a malware analyst previously and a reverse engineer in general, decompilation is the most important part of such tools. It's all about speed, pseudo-C of some kind lets you roughly understand what's going on in a function in seconds. I guess you can become pretty fast with assembly too, but C is just a lot more dense.

Regarding reliability, I would say that Hex-Rays is pretty reliable (at least for x86) if you know its limitations, like throwing away all code in catch blocks. Usually wrong decompilation is caused by either wrong section permissions, or wrong function signature, both of them can be fixed. It can have bad time when stack frame size goes "negative" or some complex dynamic stack array logic is involved, which are usually signs of obfuscation anyway.

It was less reliable 10 years ago though.. Also even now hex-rays weirdly does not support some simple instructions like movbe.

vient··on Cockpit mishap seen as likely cause of plunge on Latam Boeing 787
Guess the real reason may also provoke anger in some passengers. You can argue with attendants if you know that crew did it, you can't really argue with a plane or forces of nature. For those passengers providing fake reason would ensure a safe flight.
← PreviousPage 2 of 4Next →