HNHacker News
TopNewBestAskShowJobs

mshockwave

1,083 karma · joined September 22, 2015

Interested in Compilers, Code Generation, Static Analysis, System Software and System Security. Familiar with LLVM.

meet.hn/city/33.6856969,-117.8259810/Irvine

submissionscomments
mshockwave··on MongoDB CEO resigns to join Meta
in case anyone who doesn't get the reference: https://youtu.be/b2F-DItXtZs?si=gxrqufTAy-YQ88wW
mshockwave··on Platform-independent SIMD in Go
> constrain the length to the maximum vector size supported by the lowest platform you want to support or you lose the portability and actually end up with code that performs much worse than the scalar alternative on some platforms.

Or, put a dynamic factor into your vector size and design everything around it. Such that every platforms can plug in their own factor and _scale_ the size of vectors. This is basically what LLVM IR does for SVE and RVV: `<vscale x 4 x i32>` where vscale is the said dynamic factor. Though the exact value of vscale is only known during runtime, it doesn't matter -- we still can design compiler optimizations and lowering around it. The generated binaries can then be portable across platforms with different vscale values.

mshockwave··on Platform-independent SIMD in Go
Just want to say among many portable SIMD solutions I’ve seen recently (e.g. Fearless SIMD), this is the first that makes non-fixed vectors like SVE and RISC-V vector (RVV) easier to support. Glad to see they made this decision
mshockwave··on A third world engineer responds to “RISC-V: They should have known better”
I disagree, compiler optimization and software ecosystem in general already take a really really really long time to build, and I think it'll take longer without a standardized low-level binary interface i.e. ISA to enable rapid distribution.

And the approach you mentioned here:

> a family of ISAs that are ABI compatible such that one can compile down to a semi-pre-optimized portable IR, and just do the last bit per ISA

I think this is basically WebAssembly and PTX, and one may argue, Java bytecode. Yet look at how much efforts and time it took for WASM runtimes and JVMs to actually produce performant machine code (the "last bit per ISA" you mentioned) for just a couple of architectures! (e.g. X86 and ARM). And I wouldn't surprised if NVIDIA pour even more money on building optimization pipeline from PTX to each of their different uArchs.

mshockwave··on Deep-sea vehicles spot 'alien' sharks deep beneath the waves in the Pacific
or have little vision to begin with
mshockwave··on Qualcomm to Acquire Modular
one of the reasons I rarely read press releases is that I don't believe in promises -- I believe in _incentives_. In this case, what will Qualcomm be incentivized to do? What are in their interests?
mshockwave··on Qualcomm to Acquire Modular
indeed, open sourcing is only half (or even less) of the picture: who is driving the open source community and how it is driven (i.e. governing structure) are probably more important IMHO. There are countless of cases where an open source project is either killed by slow death, or dictated by a single entity. Chris's previous projects like LLVM and MLIR are fortunate enough to grow and thrive organically, and that takes years if not decades to cultivate
mshockwave··on How Mark Klein told the EFF about Room 641A [book excerpt]
it's a book excerpt
mshockwave··on GCC 16 has been released
> Does gcc use LLVM anywhere under the hood

No

mshockwave··on The RISE RISC-V Runners: free, native RISC-V CI on GitHub
One thing I observed is that RVV code is usually slower in QEMU
mshockwave··on Emacs internals: Tagged pointers vs. C++ std:variant and LLVM (Part 3)
LLVM now has another way to implement RTTI using the `CastInfo` trait instead of `classof`: https://llvm.org/doxygen/structllvm_1_1CastInfo.html

But it's really just an implementation difference, the idea is still to have a lightweight RTTI.

mshockwave··on We tasked Opus 4.6 using agent teams to build a C Compiler
how did it do regalloc before instruction selection? How do you select the correct register class without knowing which instruction you're gonna use?
mshockwave··on Banned C++ features in Chromium
> I don’t know many good reasons for extrusive linked lists

for one, its iterator won't be invalidated

mshockwave··on Helion: A high-level DSL for performant and portable ML kernels
Is it normal to spend 10minutes on tuning nowadays? Do we need to spend another 10 minutes upon changing the code?
mshockwave··on Swift on FreeBSD Preview
It's likely that Swift compiler is using LLVM LIT (https://llvm.org/docs/CommandGuide/lit.html), which is implemented in python, as the test driver
mshockwave··on RISC-V Conditional Moves
> In the end, programs will want probably to stay conservative and will implement only the core ISA

Unlikely, as pointed out in sibling comments the core ISA is too limited. What might prevail is profiles, specifically profiles for application processors like RVA22U64 and RVA23U64, which the latter one makes a lot more sense IMHO.

mshockwave··on The Weird Concept of Branchless Programming
yes, it has been done for at least a decade if not more

> Even more of a wild idea is to pair up two cores and have them work together this way

I don't think that'll be profitable, because...

> When you have a core that would have been idle anyway

...you'll just schedule in another process. Modern OS rarely runs short on available tasks to run

mshockwave··on The Weird Concept of Branchless Programming
The article is easy to follow but I think the author missed the e point: branchless programming (a subset of the more known constant time programming) is almost exclusively used in cryptography only nowadays. As shown by the benchmarks in the article, modern branch predictors can easily achieve over 95% if not 99% precision since like a decade ago
mshockwave··on Machine Scheduler in LLVM – Part I
yes, the short answer is LLVM uses RegPressureTracker (https://llvm.org/doxygen/classllvm_1_1RegPressureTracker.htm...) to do all those calculations. Slightly longer answer: I should probably be a little more specific that in most cases, Machine Scheduler cares more about register pressure _delta_ caused by a single instruction, either traverses from bottom-up or top-down. In which case it's easier to make an estimation when some of other instructions are not scheduled yet.
mshockwave··on What if every city had a London Overground?
The scenario suitable for TBM is surprisingly limited so even nowadays many tunnels are still dig using the good’o way, mostly with explosives
mshockwave··on How can AI ID a cat?
Came to say Apple also did a great job on tagging my bois who are both grey-ish cats, even in pictures they faced backward, no idea how they did that
mshockwave··on Quickshell – building blocks for your desktop
I thought the original comment meant “_for_ who doesn’t use X or Discord, here is the github mirror link”. There’s a “for” missing, and thus I think they agree with you
mshockwave··on Global Trade Dynamics
second this, it didn't show anything when I hovered over North America
mshockwave··on lsr: ls with io_uring
and grep / ripgrep. Or did ripgrep migrate to using io_uring already?
mshockwave··on Preliminary report into Air India crash released
tangent: I believe Air Force and Navy have loosen the vision requirement quite a bit. IIRC the new rule only measures the corrected vision, regardless how you corrected it, including wearing glasses.
mshockwave··on SIMD.info – Reference tool for C intrinsics of all major SIMD engines
> RVV is the hardest though (20k intrinsics)

A bit late to this comment but most of these intrinsics are overloads of different LMUL and SEW on a single instruction. I'm pretty sure the actual number of RVV instructions is way less. So maybe you could consolidate overloads of the same instruction into the same page or something.

mshockwave··on SIMD.info – Reference tool for C intrinsics of all major SIMD engines
This is pretty useful! Any plan for adding ARM SVE and RISC-V V extension?
mshockwave··on Performance Debugging with LLVM-mca: Simulating the CPU
> that modeling instruction scheduling doesn't matter all that much for codegen on OoO cores.

yeah scheduling quality usually has a weaker connection to the performance of OoO cores. Though I would also like to point out:

  1. in-order cores still heavily relies on scheduling quality
  2. Issue width is actually a big thing in MachineScheduler regardless of in-order or out-of-order cores. So the problem you outlined above w.r.t different implementations of uops cracking is indeed quite relevant
  3. MachineScheduler does not use the BufferSize -- which more or less mirrors the issue queue size of each pipe -- at all for out-of-order core. MicroOpBufferSize, which models the unified reservation station / ROB size, only got used in a really specific place. However, these parameters matter (much) more for llvm-mca
mshockwave··on Performance Debugging with LLVM-mca: Simulating the CPU
> This means I can't feed any function as-is (even a tiny one). I need to manually cut out the loop body.

> It doesn't support branches at all. I know it's a very hard problem, but that's the problem I have

Shameless self-plug: https://github.com/securesystemslab/LLVM-MCA-Daemon

mshockwave··on Why Android can't use CDC Ethernet (2023)
It was later reverted[1] because "there are devices in the field using usbX interfaces for tethering". Shortly after that, it got re-landed but only supported Android V+[2]

[1]: https://android-review.googlesource.com/c/platform/packages/...

[2]: https://android-review.googlesource.com/c/platform/packages/...

Page 1 of 10Next →