HNHacker News
TopNewBestAskShowJobs

celrod

1,143 karma · joined February 4, 2018

SIMD and performance enthusiast. https://github.com/JuliaSIMD https://spmd.org/
submissionscomments
celrod··on Suckless.org: software that sucks less
I use Wshadow personally. I highly recommend it. I think code that violates it (even if correct) is harder to understand.
celrod··on Linux: We need tiling desktop environments
> I don't want my web browser or video player to be resized because I open a new program

I've been using niri (a tiling WM) recently. This is their very first design principle: https://github.com/YaLTeR/niri/wiki/Design-Principles Maybe other PaperWM-inspired WMs are similar. niri is the first I've used.

If your windows within a workspace are wider than your screen, you can scroll through them. You also have different workspaces like normal. I'll normally have 1 workspace with a bunch of terminals, and another for browsers and other apps (often another terminal I want to use at the same time as browsing, e.g. if I'm looking things up online).

celrod··on Papersway – a scrollable window management for Sway/i3wm
Do you not often quickly look between files? If so, odds are you're using tiles within tmux, vim, emacs, vscode, or something.

I use kakoune, which has a client/server architecture. Each kak instance I open within a project connects to the same server, so it is natural for me to use my WM (niri) to tile my terminals, instead of having something like tmux or the editor do the tiling for me. I don't want to bother with more than one layer of WM, where separate layers don't mix.

celrod··on ICPP – Run C++ anywhere like a script
Can confirm, this works for me in my actual examples, thanks!
celrod··on ICPP – Run C++ anywhere like a script
I've defined a few pretty printers, but `operator[]` doesn't work for my user-defined types. Knowing it works for vectors, I'll try and experiment to see if there's something that'll make it work.

  (gdb) p unrolls_[0]
  Could not find operator[].
  (gdb) p unrolls_[(long)0]
  Could not find operator[].
  (gdb) p unrolls_.data_.mem[0]
  $2 = {
`unrolls_[i]` works within C++. This `operator[]` method isn't even templated (although the container type is); the index is hard-coded to be of type `ptrdiff_t`, which is `long` on my platform.

I'm on Linux, gdb 15.1.

celrod··on ICPP – Run C++ anywhere like a script
How feasible would it be for something like gdb to be able to use a C++ interpreter (whether icpp, or even a souped up `constexpr` interpreter from the compiler) to help with "optimized out" functions?

gdb also doesn't handle overloaded functions well, e.g. `x[i]`.

celrod··on AMD Ryzen 5 9600X and Ryzen 7 9700X Offer Excellent Linux Performance
Yes, I see now that while not advertised on seller's websites, Asus's product pages do indeed say that.
celrod··on Zen5's AVX512 Teardown and More
Skymont little cores have 4x 128-bit execution. They could quadruple-pump.

But looks more like they're giving up on people writing code for wide vectors, instead settling on trying to make the existing code faster.

celrod··on AMD Ryzen 5 9600X and Ryzen 7 9700X Offer Excellent Linux Performance
Zen5 appears to officially support up to DDR5 5600, but unfortunately all of the ASRock Rack or Supermicro boards I looked at only supported DDR5 5200.

I may wait for new Zen5 boards, or maybe take a gamble on something like the Asus ProArt, where I saw comments online indicating that ECC is (unofficially?) supported.

Looking forward to Ryzen 9000 ECC benchmarks.

celrod··on AMD Ryzen 5 9600X and Ryzen 7 9700X Offer Excellent Linux Performance
Any suggestions for ECC?

Would you suggest going with an ASRock Rack motherboard, even for desktop use, like you used here? https://www.phoronix.com/review/amd-ryzen9-ddr5-ecc

I'm strongly tempted to get a Zen5 CPU, but am unsure of the motherboard.

celrod··on Clang vs. Clang
Signed integer overflow being undefined has these two consequences for me: 1. It makes my code slightly faster. 2. It makes my code slightly smaller. 3. It makes my code easier to check for correctness, and thus makes it easier to write correct code.

Win, win, win.

Signed integer overflow would be a bug in my code.

As I do not write my own implementations to correctly handle the case of signed integer overflow, the code I am writing will behave in nonsensical ways in the presence of signed integer overflow, regardless of whether or not it is defined. Unless I'm debugging my code or running CI, in which case ubsan is enabled, and the signed overflow instantly traps to point to the problem.

Switching to UB-on-overflow in one of my Julia packages (via `llvmcall`) removed like 5% of branches. I do not want those branches to come back, and I definitely don't want code duplication where I have two copies of that code, one with and one without. The binary code bloat of that package is excessive enough as is.

celrod··on "We ran out of columns"
I don't think you'd even necessarily need to ignore. Roll it out in phases. You aren't going to have to deliver the final finished solution all at once.

Some elements are inevitably going to end up being de-prioritized, and pushed further into the future. Features that do end up having a lot of demand could remain a priority.

I don't think this is even a case of "ask for forgiveness, not permission" (assuming you do intend to actually work on w/e particular demands if they end up actually continuing to demand it), but a natural product of triage.

celrod··on rr – record and replay debugger for C/C++
Thanks for the clarification.
celrod··on rr – record and replay debugger for C/C++
C++20 added `[[no_unique_address]]`, which lets a `std::is_empty` field alias another field, so long as there is only 1 field of that `is_empty` type. https://godbolt.org/z/soczz4c76 That is, example 0 shows 8 bytes, for an `int` plus an empty field. Example 1 shows two empty fields with the `int`, but only 4 bytes thanks to `[[no_unique_address]]`. Example 2 unfortunately is back up to 8 bytes because we have two empty fields of the same type...

`[[no_unique_address]]` is far from perfect, and inherited the same limitations that inheriting from an empty base class had (which was the trick you had to use prior to C++20). The "no more than 1 of the same type" limitation actually forced me to keep using CRTP instead of making use of "deducing this" after adopting c++23: a `static_assert` on object size failed, because an object grew larger once an inherited instance, plus an instance inherited by a field, no longer had different template types.

So, I agree that it is annoying and seems totally unnecessary, and has wasted my time; a heavy cost for a "feature" (empty objects having addresses) I have never wanted. But, I still make a lot of use of empty objects in C++ without increasing the size of any of my non-empty objects.

C++20 concepts are nice for writing generic code, but (from what I have seen, not experienced) Rust traits look nice, too.

celrod··on Do not taunt happy fun branch predictor (2023)
Multiple accumulators increases accuracy. See pairwise summation, for example.

SIMD sums are going to typically be much more accurate than a naive sum.

celrod··on GCC's new fortification level: The gains and costs (2022)
C++23 added `allocate_at_least`: https://en.cppreference.com/w/cpp/memory/allocator_traits/al...

I'm not sure if any standard libraries have an implementation that takes advantage of the "at least" yet.

celrod··on I found an 8 years old bug in Xorg
Yes! It annoys me when a scene with characters shouting is much louder than a scene where characters are talking with hushed voices, as an example.

We know a shout was louder at the source, but the decibel level at our ears is proportional to 1/distance squared, meaning hushed voices aren't necessarily any quieter "in real life". I'd prefer suspenseful and dramatic scenes both play at similar, comfortable levels. I don't want to have to adjust the volume up and down so I can understand one scene and then not have it be disturbingly loud in the next. In practice, I just use subtitles to circumvent the "difficult to understand" problem.

celrod··on Leveraging Zig's Allocators
Nontemporal writes are substantially slower, e.g. with avx512 you can do 1 64 byte nontemporal write every 5 or so clock cycles. That puts you at >= 640 cycles for 8 KiB. https://uops.info/html-instr/VMOVNTPS_M512_ZMM.html
celrod··on Leveraging Zig's Allocators
This memory is now the least recently used in the L1 cache, despite being freed by the allocator, meaning it probably isn't being used again.

If it was freed after already being removed from the L1 cache, then you also need to evict other L1 cache contents and wait for it to be read into L1 so you can write to it.

128 cycles is a generous estimate, and ignores the costs to the rest of the program.

celrod··on Intel's Lion Cove Architecture Preview
Different vector widths for different cores isn't currently feasible, even with SVE. So all cores would need to support 1024-bit SIMD.

I think it's reasonable for the non-SIMD focused cores to do so via splitting into multiple micro-ops or double/quadruple/whatever pumping.

I do think that would be an interesting design to experiment with.

celrod··on AMD Unveils Ryzen 9000 CPUs for Desktop, Zen 5
I read that comment as "the wider, the sweeter" (which I agree with), but that we're now (as you say) at the end of the road, and thus the sweetest point.

But an increase in cacheline size would be nice if it can get us larger vectors, or otherwise significantly improve memory bandwidth.

celrod··on Thoughts on Skymont Slides
Our software ecosystem doesn't work well with an army of ants. I think we'd need a paradigm shift to get there.

Also, FWIW, Xeon Phi hit 244 threads in 2012 and 256 threads in 2016, although it used 4 threads/core.

celrod··on Unexpected anti-patterns for engineering leaders
Yeah, I don't think it's useful (except to score [office] political points) to read the least generous interpretation you can. Trying to understand a position lets you better decide whether or not it actually makes sense, e.g. Chesterton's fence. Although, if we tried to fairly assess every argument we see, we'd spend all our time doing that instead of being productive, so I get why it's often wise to dismiss things out of hand, at least initially.

You can disagree with the article, but I think your parent comment is not a fair summary, e.g. w/ respect to bullets 1 and 2:

> At first, I thought, ‘This is a really unreasonable person.’ But later as I dug into it, I discovered he was right...This process of conflict mining served Larson a key lesson. “I could have just ignored him, But then I would have missed the key learning, which is that I was the one who was missing context, and needed to refine my approach.”

Maybe the parent comment was "conflict mining" itself. Often the quickest way to learn about something is to make a statement about it and then let others correct you if you're wrong.

On the other hand, "what is a power dynamic" is a fair counterpoint to that. Jeff Bezos, in contrast, said he generally withheld his opinion until the end so others wouldn't be afraid to contradict him; a powerful person sharing their idea can prevent others from sharing better ones.

celrod··on Why not just do simple C++ RAII in C?
> We live in mortal fear of compiler writers smiting us for innocent things like punning through a union.

C++20 introduced `std::bitcast`, so I appreciate alias analysis getting all the help it can.

celrod··on New exponent functions that make SiLU and SoftMax 2x faster, at full accuracy
Yes, I have an AVX512 double precision exp implementation that does this thanks to iperm2pd. This approach was also recommended by the Intel optimization manual -- a great resource.

I just went with straight math for single-precision, though.

celrod··on Qualcomm's Oryon LLVM Patches
The Apple M4 is rumored to be ARMv9, even featuring SME on top of SVE2: https://wccftech.com/apple-m4-adopts-armv9-run-complex-workl...

If true, I'd much rather buy an M4 and use Asashi linux (once it supports the M4) than Oryon, but I enjoy low level coding. Being NEON-only doesn't really bring anything new or interesting to the table from that perspective.

celrod··on GPUs Go Brrr
Bill Dally from Nvidia argues that there is "no gain in building a specialized accelerator", in part because current overhead on top of the arithmetic is in the ballpark of 20% (16% of IMMA and 22% for HMMA units) https://www.youtube.com/watch?v=gofI47kfD28
celrod··on The Emacs Window Management Almanac
EXWM is ;). I used it for about a year, but ultimately found the performance, and in particular a low language server locking up the window manager, unacceptable.
celrod··on Speeding up C++ build times
I'm still waiting for clangd support, e.g. [0] before trying modules. But maybe I should just try it, as at least one person reports that it already works [1].

[0] https://github.com/clangd/clangd/issues/1293 [1] https://github.com/snu-sf-class/swpp202401/issues/21

celrod··on Speeding up C++ build times
What's the trick with explicit template instantiations? Including them in the precompiled header?
← PreviousPage 2 of 13Next →