HNHacker News
TopNewBestAskShowJobs

mshockwave

1,083 karma · joined September 22, 2015

Interested in Compilers, Code Generation, Static Analysis, System Software and System Security. Familiar with LLVM.

meet.hn/city/33.6856969,-117.8259810/Irvine

submissionscomments
mshockwave··on Beating Google's kernelCTF PoW using AVX512
Nitpick: static link won’t give you inlining but only eliminates the overhead of PLT. LTO will give you more opportunities for inlining
mshockwave··on Improving performance of rav1d video decoder
> In the case where one of the structs has just been computed, attempting to load it as a single 32-bit load can result in a store forwarding failure

It actually depends on the uArch, Apple silicon doesn't seem to have this restriction: https://news.ycombinator.com/item?id=43888005

> In a non-inlined, non-PGO scenario the compiler doesn't have enough information to tell whether the optimization is suitable.

I guess you're talking about stores and load across function boundaries?

Trivia: X86 LLVM creates a whole Pass just to prevent this partial-store-to-load issue on Intel CPUs: https://github.com/llvm/llvm-project/blob/main/llvm/lib/Targ...

mshockwave··on A leap year check in three instructions
sometimes also known as superoptimization, which many of them also use SMT solvers like Z3 mentioned in the article
mshockwave··on Forget IPs: using cryptography to verify bot and agent traffic
I might be wrong, but it seems like Anubis asks the _client_ to solve the cryptography challenges while the approach Cloudflare described here asks the server to verify the (cryptography) signature?
mshockwave··on JEP 515: Ahead-of-Time Method Profiling
in addition to storing profiles, what about caching some native code? so that we can eliminate the JIT overhead for hot functions

EDIT: they describe this in their "Alternative" section as future work

mshockwave··on Gorgeous-GRUB: collection of decent community-made GRUB themes
judging from some of these repos, I think the answer is: they don't. It seems like you have to manually pick images with the correct resolution
mshockwave··on OrangePi RV2: The New Reference SBC for RISC-V Enthusiasts – Full Review
nit: it shouldn’t be RV64GC if it has vector extension
mshockwave··on Compiling C++ with the Clang API
I think this is exactly how Zig compiler does under the hood for C/C++ sources. So I guess you can do a similar thing for your own programming languages to support interop.
mshockwave··on The FFT Strikes Back: An Efficient Alternative to Self-Attention
yes, almost all DSPs I know have native HW supports for FFT, since it's the bread and butter for signal processing
mshockwave··on Wild – A fast linker for Linux
It’s pretty common to roll your own linker script in embedded software development
mshockwave··on Tilde, My LLVM Alternative
> Here's my question: It seems inevitable that people will eventually port all LLVM's smarts directly into MLIR, and remove the need to shift between the two. Is that right?

Theoretically, yes -- taking years if not decades for sure. And set aside (middle-end) optimizations, I think people often forgot another big part of LLVM that is (much) more difficult to port: code generation. Again, it's not impossible to port the entire codegen pipeline, it just takes lots of time and you need to try really hard to justify the advantage of moving over to MLIR, which at least needs to show that codegen with MLIR brings X% of performance improvement.

mshockwave··on Tilde, My LLVM Alternative
modular in terms of using only some of the LLVM libraries without the need to pull the entire compiler into your project. In fact, many of the LLVM libraries have absolutely nothing to do with LLVM IR and have zero dependency on it. For instance, LLVMObject and LLVMDebugInfoDWARF. You can use those libraries to build useful tools, like your own objdump or just use it to read debug info.
mshockwave··on Tilde, My LLVM Alternative
I’m definitely happy to see this happening. But I would like to point out two ingredients that constitute LLVM’s success beyond academic merits: License and modularity. I’m not a lawyer so can’t say much about the first one, all I can say is that I believe license is one of the main reasons Apple switched to LLVM decades ago. Modularity, on the other hand, is one of the most crucial features of LLVM and something GCC struggles to catch up even nowadays. I really hope Tilde can adopt the modularity philosophy, provide building blocks rather than just tools
mshockwave··on Execution units are often pipelined
> ALUs in Recent Apple cpus can actually start a new division every other cycle (in addition to having an abnormally low latency)

That's indeed impressive.

I'll argue that we're definitely capable of making fully pipelined divisions, it's just that it's usually not worth the PPA.

mshockwave··on Execution units are often pipelined
shameless self plug of modern uArchs and how LLVM models it: https://myhsu.xyz/llvm-sched-model-1
mshockwave··on Execution units are often pipelined
divisions, regardless of integer or floating point, are usually NOT pipelined though
mshockwave··on Show HN: Instantly visualize any codebase as an interactive diagram
> Of course I have to test it with https://gitdiagram.com/torvalds/linux

I really wish to see how well (or bad) it works on mega projects. Because those are usually the ones I need diagrams like this the most.

mshockwave··on Qualcomm wins licensing fight with Arm over chip designs
Zb extension is in both RVA22 and RVA23 profiles, meaning application cores (targeting consumer devices like smartphones) designed in the past few years almost certainly have shXadd instructions in order to be compatible with the mainstream software ecosystem
mshockwave··on RISC-V HiFive Premier P550 Development Boards with Ubuntu Now Available
SpacemiT K1 on BananaPi is another commonly seen RVV 1.0 capable chip. IIRC both Kendryte K230 and SpacemiT K1 are in-order cores.
mshockwave··on RISC-V HiFive Premier P550 Development Boards with Ubuntu Now Available
> 2 others that I honestly have no idea what they do

CSR (i.e. status) register and instruction fence extensions. Instruction fences are most useful in cases where you modify text section during runtime (e.g. JIT or code hot reload) such that you need to ensure the consistency of code across different harts

mshockwave··on Common misconceptions about compilers
Re: Why does link-time optimization (LTO) happen at link-time?

I think maybe LLVM's ThinLTO is what you're looking for where whole-program optimization (more or less) happens in middle end.

mshockwave··on Common misconceptions about compilers
> As an example, in my desktop, it takes ~25min to compile a Debug version of LLVM, but ~23min to compile a Release version.

oh I think I know what might cause this: TableGen. The `llvm-tblgen` run time accounts for a good chunk of LLVM build time. In a debug build `llvm-tblgen` is also unoptimized, hence the long run time generating those .inc / .def files. You can enable cmake variable `LLVM_OPTIMIZED_TABLEGEN` to build `llvm-tblgen` in release mode while leaving the rest of the project in debug build (or whatever `CMAKE_BUILD_TYPE` you choose).

mshockwave··on Common misconceptions about compilers
> The compiler optimizes for data locality

> So, we have a single array in which every entry has the key and the value paired together. But, during lookups, we only care about the keys. The way the data is structured, we keep loading the values into the cache, wasting cache space and cycles. One way to improve that is to split this into two arrays: one for the keys and one for the values.

Recently someone proposed this on LLVM: https://discourse.llvm.org/t/rfc-add-a-new-structure-layout-...

Also, I think what you meant by data locality here is really optimizing data layout, which, as you also mentioned, is a hard problem. But if it's just optimizing (cache's) locality, I think the classic loop interchange also qualifies. Though it's not enabled by default in LLVM, despite being there for quite a while.

mshockwave··on LLVM-powered devirtualization
I wouldn't recommend using the term "devirtualization" here, as that term has been used to refer simplifying C++ virtual function calls (into normal function call) in LLVM. And such optimization has been turned on by default for quite some time.
mshockwave··on Show HN: I made an ls alternative for my personal use
This is also one of the reasons I write C++ with vim without any auto-completion nor fancy plugins (I do use syntax highlighting though, but I think it comes by default with vim nowadays), as well as using GNU screen -- not every machines install tmux by default, surprisingly. In case I need to login into some random Linux box, I'm sure I'll be almost as productive as I am on my own machine.
mshockwave··on Assembly Optimization Tips by Mark Larson (2004)
> No one's doing ASM programming on x86 CPUs these days

I don't think that's entirely true...it's still pretty common to write high-performance / performance sensitive computation kernels in assembly or intrinsics.

mshockwave··on Assembly Optimization Tips by Mark Larson
> (Intermediate)1. Adding to memory faster than adding memory to a register

I'm not familiar with Pentium but my guess is that memory store is relatively cheaper than load in many modern (out-of-order) microarchitectures.

> (Intermediate)14. Parallelization.

I feel like this is where compilers come into handy, because juggling critical paths and resource pressures at the same time sounds like a nightmare to me

> (Advanced)4. Interleaving 2 loops out of sync

Software pipelining!

mshockwave··on AMD outsells Intel in the datacenter space
Reminds me that's also many people's speculation on how Qualcomm builds their RISCV chips -- swap an ARM decoder for a RISCV one.
mshockwave··on AMD outsells Intel in the datacenter space
RP2350 is using a hybrid of ARM and RISCV already. Also it's not really hard to use RISCV not as the main computing core but as a controller in the SoC. Because the area of RISCV cores are so small it's pretty common to put a dozen (16 to be specific) into a chip.
mshockwave··on Wasp Flamethrower Drone Attachment
I think this is indeed invented for controlled burns in fighting bush fires
← PreviousPage 2 of 10Next →