HNHacker News
TopNewBestAskShowJobs

dzaima

1,741 karma · joined September 30, 2019

https://github.com/dzaima/

https://codeberg.org/dzaima/

submissionscomments
dzaima··on Fearless SIMD v1.0
Stop the presses! 5 instructions!

    vpor        xmm2, xmm0, 0x80000000
    vpcmpeqd    xmm2, xmm1, xmm2
    vcmpunordps xmm2, xmm2, xmm0
    vminps      xmm0, xmm1, xmm0
    vblendvps   xmm0, xmm0, xmm1, xmm2
yay for superoptimizers
dzaima··on Fearless SIMD v1.0
That vminpd+vminpd+vorpd actually does handle signed zero properly! Screws up NaN payloads to the max though. Can be easily extended to canonicalize the NaN with 2 instrs + constant though of course. (which ends up at the same number of instrs as your proper impl (albeit with worse port distribution and latency), but you get to have a canonical NaN!)

Hit upon https://github.com/llvm/llvm-project/issues/217376 while playing around with proper minimumnum, failing to SMT-verify whatever version of LLVM I had; did find a funky working 6-instr (+ constant) version though:

    vpandn      ymm2, ymm1, ymm0
    vpcmpeqd    ymm2, ymm2, 0x80000000 # whether ymm0 is -0 and ymm1 is +0 (or other cases that magically don't cause issues)
    vcmpltpd    ymm3, ymm0, ymm1
    vcmpunordpd ymm3, ymm3, ymm1 # regular NaN-is-larger ymm0<ymm1
    vpor        ymm2, ymm2, ymm3
    vpblendvb   ymm0, ymm1, ymm0, ymm2
dzaima··on Looking forward to Git 2.56 – and 3.0
> they're going through the pain of changing the hash function

They've already gone through the pain, deciding on it on 2018[0] (and functional & non-experimental 3 years ago per TFA). What's left is just changing the default (and some stragglers to complete support). Changing the function now would push back changing the default by a couple additional years until the new git version gets widespread deployment (incl. on LTS distros and whatnot).

[0]: https://github.com/git/git/commit/0ed8d8da374f648764758f1303...

dzaima··on Microcode in Intel's 8087 floating-point chip: the scale instruction
Except the ALU isn't actually 64-bit, it's 67-bit as per article, extra bits for rounding. I'd imagine it was just taken for "prettiness", with 15 bits for exponent being basically reasonable. (maybe some algorithms which double precision per iteration would like it being a power of two? but any such probably vary significantly on initial estimate precision anyway)
dzaima··on Working with Git Worktrees in Magit
Or "--ignore-other-worktrees" for "git switch".

This leaves things in a weird state though, as modifying the branch from one of the worktrees will effectively change the ref the other is pointing to, but won't update the other's working tree, making it look like it's preparing to revert all the changes that the first worktree made.

(obligatory mention of jj, whose workspaces track what state they were snapshotted at, thus never losing what the true diff of a workspace should be, allowing a "jj workspace update-stale" to safely update the working copy with whatever the other workspaces changed (if you change the working copy and in parallel from another workspace change what that workspace points to, you'll get a saved commit of the working copy changes before being updated, which you'll need to squash/rebase wherever those changes were needed))

dzaima··on Replacing a Rust Enum with a 64-Bit Word Made My Interpreter 17% Faster
How so? It's replacing multiple enum variants, but just one enum, "enum Value".

(also; if anything, the title is implying the exact opposite of "Rust compiler was able to optimize ...", "Replacing a Rust [...] with [...]" is clearly moving away from Rust-magic to something else)

dzaima··on Replacing a Rust Enum with a 64-Bit Word Made My Interpreter 17% Faster
Unless the highly-optimized parts are wrapped by an interface that looks similar to the non-optimized version.

In the case of a tagged object in Rust, depending on how well the compiler can wrangle through it, you might even be able to add a `.unpack()` method that returns a pretty enum from a packed value, that you can pattern-match on or whatever, and let the compiler remove all the code of unpacking unused cases.

(using that directly for the addition example would end up less efficient of course, but still most likely beneficial. It's after this when there's a potential true readability vs performance tradeoff)

dzaima··on bzip3
> That always seemed annoying to me. They couldn't allocate 5 more bits somewhere to let the decompressor autodetect longer window sizes?

I believe this is just to prevent the decompressor from arbitrarily blowing up memory usage based on the input; I think if you want to accept long windows you can just always decompress with --long=63 regardless of whether the input needs it? (you will run out of RAM decompressing a long=63 file though of course)

dzaima··on GrapheneOS says Pixel 11 has MTE support after all
Heh, you can kinda think of MTE as ASan except instead of a small range of a guaranteed redzone around heap pointers, it's a massive `2^56 * (random number, ≥0, on average 15)`-byte "redzone" (and some padding up to a multiple of 16 bytes which can predictably hide a bug, though at least such a bug won't corrupt unrelated heap). Stack handling is a significant difference though.
dzaima··on The creator of Jujutsu has joined ERSC
Cleaning it would change the commit ID, so a forge cannot clean it even if it wanted to (not without rewriting all descendant commits too and breaking signed commits, at least).

The change-id is exactly as much part of the commit object as the author/committer name/timestamp, description, parent commit ID, tree, and participates in the commit hashing as those do.

dzaima··on The creator of Jujutsu has joined ERSC
Technically, you can still tell actually - jj writes a "change-id xyz..." in the git commit object header, which remains there as it's pushed around. It's just typically not made visible by regular things. (I wonder what other random garbage has been hidden in git commit headers that noone has seen)
dzaima··on Luanti removed from Google Play due to baseless AI copyright notice
The things/concepts that those screenshots have that infiniminer (a voxel game made before minecraft) doesn't is... grass, trees, glass. I hate to bring it to you, but minecraft didn't invent those. And it certainly didn't invent the concept of a voxel world (not that it could even copyright that if it did).

Never mind that the things in those in-game screenshots aren't even in the play store app, they're separately downloadable things.

dzaima··on RISC-V: They Should Have Known Better
That is quite a good bit more evenly-spread (the "..." is 5159 instrs).

Wonder what's up with bit 21; if whatever uses it so much is repositionable (and not an aarch64-specific thing), could save like 2KB on x86-64 via putting it in the low 8 bits instead.

dzaima··on RISC-V: They Should Have Known Better
Some stats on an aarch64 binary of my current main project (1.6MB .text, 6600 symbols as per whatever "nm the-binary | wc -l" includes, from "objdump -d the-binary"):

    19546 /tbn?z/
    18029 /tbn?z.*, #0x0/  (but this includes boolean checks)
      224 /tbn?z.*, #0x1f/ (i.e. 32-bit x<0)
     1139 /tbn?z.*, #0x3f/ (i.e. 64-bit x<0)
      154 other immediates
Said project doesn't do fixed bitfields much (there are some, but a chunk of those test multiple bits) so unsurprisingly not much. (I could imagine that the kernel has significantly more, but it's an edge-case (though perhaps an important one) of being basically massive amounts of fixed configurable glue)
dzaima··on RISC-V: They Should Have Known Better
Both clang and gcc do actually generate TBZ/TBNZ for checking a bool: https://godbolt.org/z/K6evhaxGT
dzaima··on RISC-V: They Should Have Known Better
Some more:

> The spec says that bit must be zero, and yet no encoding uses the space opened up by that bit being one.

The spec says "the code points with shamt[5]=1 are designated for custom extensions.", so the space is specifically reserved for custom vendor extensions.

So, if I wanted to add a custom "dzaima.c.clear_top_n_bits rd, imm5" instruction, that's space I could safely put it in, knowing that no future standard instruction will be added there that I may regret overlapping. So while that space goes unused in the standard, its existence helps with the overlapping encoding problem!

> For I-type instructions, bit 1 [...], bit 11

Of course, that's cherry-picking two of the 25% of bits that have multiple positions they come from, and specifically 11 as it's the worst one. Full stats:

    1 position: 24 bits: (everything that's not listed below)
    2 positions: 7 bits: 0, 1, 2, 3, 4, 12, 20
    3 positions: 1 bits: 11 (the single worst case)
So that's like 9 muxes for merging all immediates to the same place (or less of course if the different encodings' immediates go to different places), the rest is just wires.

Obligatory note is that some of the funkiness is to place the sign-extended bit in the same bit position, so some saved muxes from that.

Now, I am a "software person who's never written verilog", but I highly doubt a 3:1 mux is as cheap as a 2:1 mux in silicon, so even if you always need to merge in the sign bit, reducing the number of cases is still beneficial.

Compressed does make it a ton more ugly though (combining both 32-bit and 16-bit instruction encodings, placing the 16-bit ones in the low 16 bits):

    1 position: 13 bits
    2 positions: 7 bits: 10, 13, 14, 15, 16, 17, 20
    3 positions: 4 bits: 3, 4, 9, 12
    4 positions: 5 bits: 0, 1, 2, 5, 11
    5 positions: 3 bits: 6, 7, 8
looking at aarch64 on https://asmjit.com/asmgrid/:

    tbz Xt, #imm, #relS*4   imm:1|0110110|imm:5 |    relS:14 |Rt
    lsl Xd, Xn, #n            1   1010011|01|immr:6|imms:6|Rn|Rd
Fun! (lsl being a subset of the bitfield extract instrs is neat; tbz's similar-functionality 6-bit field is just entirely-differently placed though. Also.. using the Rd slot for an input-only Rt? that's one thing RISC-V doesn't do, even across compressed and 32-bit instrs!)
dzaima··on RISC-V: They Should Have Known Better
Random minor-ish notes:

- A big problem with extension detection RISC-V has is that there's no central authority mandating vendors to not overlap things (obviously, given RISC-V being an open standard), so basic bitmasks for supported extensions is generally rather problematic (and of course even if you collected a standardized bitmask of all extensions from all vendors, it'd grow quite massive quite quickly); you'd at least want some grouping/marking by vendor, if not full extension strings. That said, it would be nice to at the very least have some standard in-memory blob format if nothing else, that you could query from any OS/libc. (which maybe somewhat-exists to some extent with a C API meant for libc, but as-is still doesn't attempt to figure out vendor extensions).

- many, if not the vast majority, of aarch64 TBZ/TBNZ are probably branching on a boolean; something RISC-V can also of course do in one instruction. Generally, comparing instruction frequencies across ISAs is messy if not approximately meaningless due to different sorts of things existing for solving the same tasks.

- "Having this happen means that instead of a clearly-understandable crash you get ... well ... anything." - RISC-V will do you one better - it doesn't even guarantee a crash when an instruction isn't defined at all! Overlapping extensions is definitely messy for disassembly, sure, but that's also just basically unavoidable as long as RISC-V is open (see my first point). (perhaps there could've been stricter rules for reserved-for-standard encodings than reserved-for-vendor ones? of course still doesn't help vendor encodings, nor non-compliant vendors)

dzaima··on Moving integer division to floating-point is trivial
Float divides are still pretty expensive; 8-10 cycles of latency on modern hardware, integer divides being 8-20 cycles. (on Apple M1 both are 8-10 cycles; int div is much worse on older x86 hw)

Integer multiply is also pretty universally 3 cycles of latency, i.e. basically the same as float multiply (or even add!).

What float div definitely has over int div is throughput, as float div comes in vectorized versions on x86 & ARM, and it usually is actually parallelized.

On top of generally fp div generally having higher throughput (M1 gets down to 1 instr/cycle! though int div isn't bad either at 0.5 instrs/cycle; x86 numbers are messy but even 32-bit int div is never better than f64 div, though they're close; also an annoying aspect is that x86 division instrs actually always take a 128-bit divisor, though hopefully a sign-/zero-extended 64-bit value skips the extra work)

dzaima··on Faster floating point math with Rust's new API
> Sure, and that's basically what I was wondering about with respect to "can't define it" being shorthand for something else

Eh, I'd say it's still the same thing; can't define a sanitizer for it if what you define isn't a sanitizer. Depends on a specific definition of "sanitizer" though.

> At least making your own isn't horrendously difficult.

In C++ perhaps, but impossible in C. (and there are still some funky edge-cases where multiplying two `uint16_t`s can overflow due to implicit promotion to signed int; C's _BitInt solves at least that)

Clang does actually have experimental support for these - https://clang.llvm.org/docs/OverflowBehaviorTypes.html

dzaima··on Faster floating point math with Rust's new API
> it's still a counterexample for "you can't define it because it means sanitizers can't warn for it"

Sure, technically you can write a sanitizer for anything. It just becomes less a "sanitizer" you can always recommend everyone everywhere use, and more of just a heuristic thing that only really works if you design your code for its arbitrary desires.

> and in the cases where it actually is intentional you can suppress the check.

imo it'd be nice to have separate types for wrapping and non-wrapping integers for that, so that you have actual language-level semantics and an easy way to mix things (e.g. wrapping arith for hashing, mixed with non-wrapping arith for loop index or whatever) instead of suppressions.

dzaima··on Faster floating point math with Rust's new API
> -fsanitize=unsigned-integer-overflow

Of course, -fsanitize=unsigned-integer-overflow isn't enabled by default, and few people use it (github code search gives 6K results for that, compared to 175K for "-fsanitize=undefined"; which to be fair is a lot higher than I expected, but still not a lot).

> Makes me wonder whether "sanitizers can't flag defined behavior" is meant to be shorthand for some more nuanced position

And signed overflow checking would have to be off-by-default too, if people were allowed to start relying on it. It'd be less "false positive rate too high", more "it disallows you to use a genuine language feature that is actually useful", defeating the point of defining signed overflow in the first place.

(imo defining signed overflow specifically for reducing attack surface from exploitable UB is a mostly-separate discussion, which should not affect core language semantics, and certainly not what users would be suggested to do)

dzaima··on Faster floating point math with Rust's new API
C23 specifies that signed integers must be two's complement, but still leaves signed arithmetic overflow as undefined behavior.
dzaima··on Branchless Rust: Making a Filter 4x Faster by Removing an If
There's no autovectorization there; scalar f64-s just are always stored in xmm registers. And the bounds check is still there.
dzaima··on Branchless Rust: Making a Filter 4x Faster by Removing an If
Compress patterns aren't recognized by any open-source compiler autovectorizer as far as I'm aware of. (I think intel's proprietary C/C++ compiler can?)
dzaima··on Nvidia’s Vera Whitepaper Has a Thread Loose
aarch64 has a CPU mode, DIT (Data Independent Timing), specifically for allowing software to request all fancy value prediction stuff to be disabled for the duration of processing of sensitive data.

(doesn't help when the attack target is general-purpose/user-controlled code leaking things, but if you're relying on a process not leaking memory plainly available to it without full careful control of what the process runs, you've already been fully-SOL on that for decades and nothing has nor will nor can change about that)

dzaima··on JEP 401: Value Objects (Preview) merged to OpenJDK master
While I agree with the comment being overly-praise-y, this is complexity that any multithreaded language with mutability and reasonable sanity/safety desires (but without intrusive compile-time rules a la Rust) has; others just might pretend some of the options don't exist / aren't desirable. (..and value objects, as landed in the PR, doesn't yet have tearable objects)
dzaima··on Zig's Incremental Compilation Internals
A quick test (C, clang) gives me that a binary depending on 1000 shared libraries, each containing a single function returning an integer, with a main function summing up the results of all those functions, takes ~270ms to run from the dynamic linker overhead. So you'd definitely want a good chunk below thousands.

(a process with 100 shared libraries takes 6ms to run, which is a lot better (0.9ms for 1 library, for reference), but, especially with in-place patching skipping work on unchanged values, static linking still has a good shot at beating dynamic linking, especially if you run the binary multiple times)

Incremental compilation generally already depends on its stored intermediate data not getting corrupted, and the final binary need not be any differently handled in that aspect.

dzaima··on SIMD for Collision
Unaligned loads/stores aren't super bad; if still within a cacheline, there's zero penalty, and on crossing cachelines it's alike two ops (except page crossing, which is more bad).

So, for 32B loads/stores and 64B cachelines, it's 1.5x more L1 cache ops (as half of the ops will cross a cacheline); perhaps bad if you're L1-cache-throughput-bound, but less so if you're at L2+ as the extra work sits in L1.

dzaima··on Codeberg: ToU extension to prohibit LLM-extrusions
Where I eat doesn't have a "WC" sign, walls separating it from public areas, and a toilet present. Codebergs private repos have a clear "private" label, infrastructure for making them, acknowledgement and rules for it in ToS, and they are kept private.

Anyway, originally I didn't quite agree with fs111 on straight up calling your post misinformation, but I guess it is indeed your intent to completely ignore facts and just say random garbage.

dzaima··on Codeberg: ToU extension to prohibit LLM-extrusions
...in what way whatsoever is this "They dont allow private repos"?

The size limit is........a size limit, something every host should have; there's one for public codeberg repos too.

And while there are restrictions on what private repos are allowed, there are also restrictions on what public repos are allowed too, and it's extremely clear that neither requirement is equivalent to "they don't allow [public / private] repos".

My reading of that section is that, besides the size limit, the rules on private repos are, to an extent, less strict than of public repos; anyone who falls under your second quoted sentence couldn't have any public repos either, by the nature of public repos needing to be FOSS.

Page 1 of 32Next →