HNHacker News
TopNewBestAskShowJobs

camel-cdr

2,238 karma · joined January 31, 2021

submissionscomments
camel-cdr··on Everybody’s home. No one’s coming over
I see that it's currently not culturally accepted, but fundamentally is also a case of entitlement from the hosts part.
camel-cdr··on Everybody’s home. No one’s coming over
> everyone demands they are catered to on their terms.

How much of a thing is this in practice? To me this also feels driven by the expectations of the host towards the guests beeing fully accommodated by them. It's probably a mixture of both.

Like, I'm happy to bring my own food, when invited. E.g. I would presume most vegans would be cool with bringing their own or ordering, as you are used to that anyways.

camel-cdr··on The state of SIMD in Rust in 2026
Also, rust could just be based and make -mno-strict-align the default.
camel-cdr··on Navier-Stokes – Tristan Buckmaster [pdf]
https://reddit.com/r/IndieDev/comments/1vfwuf5/a_player_foun...
camel-cdr··on On the Navier–Stokes Millennium Prize Problem
I found this post interesting in that reguard: https://www.lesswrong.com/posts/thXohzXrWCA2EhZCH/mateusz-ba...
camel-cdr··on Navier-Stokes – Tristan Buckmaster [pdf]
There also is an insentive to silently give prominent people (e.g. Linus) or reasearchers like this custom tuned system prompts or even more powerful models.
camel-cdr··on Discovery of a new OpenAI agent message board
I suppose nobody sane would give their AI internet access (even read) while training it. Though if they did, I don't think they'd want this to be public, because how can you even protect against this?
camel-cdr··on RISC-V is now officially supported by CPython
I wonder how much this matters for python. As long as the important dependencies like numpy runtime dispatch RVV, it should probably be fine.

Zba would probably give a small boost. Zbb gives a substantial boost to perf for applications that use clz/popc heavily, but I don't think that would apply to python.

camel-cdr··on SiFive's First Server Platform
https://camel-cdr.github.io/rvv-bench-results/sifive_p870/in...
camel-cdr··on Hot Chips 2026: CUDA Targets RISC-V
related: https://www.sifive.com/development-platforms/sifive-bigsky-s...
camel-cdr··on Everyone says assembly is untyped—everyone is wrong
I really like what you are doing here, the state of inline assembly is a similar travesty to the state of guided codegen/autovec.

On concern I have is how this maps to ARM64 syntax, because ARM64 is massively overloading all mnemonics.

For example:

    ld1d z0.h, p0/z, [x1, x2, lsl 3]
    ld1d z0.h, p0/z, [x1, z0.h, lsl 3]
Have extremely different performance characteristics, yet would map to the same code:

    ld1d dst, p0/z, [base + idx<<3]
Imo this makes reading the assembly quite bothersome. I'm already not a fan of ARM64 doing the mnemonic overloading, but at least you can figure out the operation by looking at the same line further to the right.

Also, maybe I missed it, but how are you dealing with things like the /z modifier, pre/post-increment load/store and load pair? Or things like TBL/ST4/LD4?

Oh and how are the types going to work for RVV, where the type can't be determined at compile-time in all situations?

camel-cdr··on Everyone says assembly is untyped—everyone is wrong
> I was aware that some architectures had distinct floating point registers

Most ISAs do.

On x86 and arm scalar and FP registers are separate, it's just that they overlap FP and SIMD registers.

On RISC-V there are three separate register files for scalar, FP and SIMD. Although you can overlap scalar and FP in some minimal embedded configurations.

camel-cdr··on Turbovec – Google's TurboQuant for vector search in Rust
Already in the works: https://lists.riscv.org/g/tech-announce/message/782
camel-cdr··on A third world engineer responds to “RISC-V: They should have known better”
The P870 is probably faster than the Cortex-X1 in my phone, which would put it on a similar per level to Zen1.

When I had brief access to a very early P870 devboard running at 2GHz it scored 2% worse than an 2.86GHz Cortex-X1 in the 7-zip benchmark compression (it was a 40% worse in decompression, but that has NEON optimized Arm assembly path and non for RISC-V). Because this was very early hardware still in bringup I expect them to get closer to the announced 2.7GHz.

But yeah, the RISC-V geekbench scores of publically available RISC-V processors are suprizingly low and idk why. Part of it is the lack of RVV support is certain workloads, but that doesn't explain the entire gap. When I looked at ST scalar performance my self it looked more comparible to similarly sized Arm CPUs.

I think I'll have to order a Pi5 (Cortex-A76) and investigate where it performes different to the SpacemiT K3 (SpacemiT X100).

On paper they should be very similar, both 2.4GHz, 4-wide decode, 3xALU, 1xBranch, (although the A76 has those on 4 execution units, while the X100 has 3 with one sharing ALU and Branch support), 2x load/store, 2xFP. The biggest obvious difference (appart from ISA), may be the cache setup after L1, with the K3 having a 4Mib L2 shared between 8 cores (no L3) and the Pi5 having a 512KiB L2 and a 2MB L3 shared between 8 cores.

camel-cdr··on A third world engineer responds to “RISC-V: They should have known better”
I think it will be closer to Zen 1/2
camel-cdr··on A third world engineer responds to “RISC-V: They should have known better”
Great news! How did I miss that?

I'll have to try out the branch.

camel-cdr··on A 3rd World Embedded Engineer Responds to "RISC-V They Should Have Known Better"
Which may actually be a good analogy. One thing I've seen rarely discussed is existing expertise/experiance, existing varification/tooling and existing quality reference implementations. Arm wins in all of the above over RISC-V currently, if you have the money to license from/work with Arm.
camel-cdr··on A third world engineer responds to “RISC-V: They should have known better”
I basically spend way to much time with RISC-V related things. But the easiest way to get more info about RISC-V developmemts is by watching the youtube uploads of the RISC-V summit talks.
camel-cdr··on A third world engineer responds to “RISC-V: They should have known better”
Here are a few random things I know of:

* Tenstorrent Ascalon has a neat optimization for certain LMUL>1 SIMD operations. LMUL=2 effectively unrolls the SIMD operation making it read two SIMD registers from every source and write two SIMD registers to the destination. There are however some instructions where LMUL=2 only needs to write to one registers, those are narrowing instructions (e.g. 64-bit to 32-bit truncation) and comparisons (which write to a LMUL=1 register with packed bits). When those SIMD instructions have to .vx form, which means one argument comes from a GPR, they now only need to write one SIMD register and need to read two SIMD registers. This matches what regular SIMD instructions need and because the silicon for the execution is much cheaper than register file ports, Ascalon can exexute these instructions in a single operation. So you can compare twice as many SIMD elements against a scalar, then you can against another SIMD register.

* Ventana (now under Qualcomm) talked a tiny bit about their fetch-block-optimizer and something that sounded like a L1i-trace cache. The fetch-block-optimizer would go to certain hot L1i entries and "optimize" them, with agressive instruction fusion including fusion of non-adjacent instructions.

* NextSilicon: Idk any details yet, but they said they handled RVC without increasing latency and that they've found a good solution for implement RVV and especially LMUL, which is a challange in out-of-order designs.

* OpenXiangShan: The fastes open-source CPU, is working on doing 2-ahead instruction fetch (the thing Zen5 added).

Now that being said, Ventana was bought by Qualcomm, we know the RISC-V team is still alive, but who knows if we'll ever see anything from that outside of Qualcomm?

The Tenstorrent Ascalon devboard is way behind schedule and on 12nm TSMC instead of a 4nm node the processor was designed for and is now supposed to clock at 1.38GHz. Though I think the delay has more to do with TT management problems then with the actual design.

While the scalar part of OpenXiangShan looks really good, the RVV imolementation is currently basically unusable. They want to have fix for the problems until the end of the year, but we'll have to see.

camel-cdr··on RISC-V: They Should Have Known Better
https://support.arm.com/documentation/109697/2026_06/Feature...

> In an Armv9.0 implementation, if FEAT_FP and FEAT_AdvSIMD are implemented, the following features are implemented: ...

This implies they don't have to be implemented.

https://support.arm.com/documentation/109697/2026_06/Feature...

> FEAT_FP is OPTIONAL from Armv8.0.

> FEAT_AdvSIMD is OPTIONAL from Armv8.0.

But also:

> All Armv8-A systems that support standard operating systems with rich application environments also provide hardware support for Advanced SIMD instructions.

and

> All Armv8-A systems that support standard operating systems with rich application environments provide hardware support for Advanced SIMD and floating-point instructions. All Armv9-A systems that support standard operating systems with rich application environments also provide hardware support for SVE2 instructions. It is a requirement of the ARM Procedure Call Standard for AArch64, see Procedure Call Standard for the Arm 64-bit Architecture

So, both FEAT_FP and FEAT_AdvSIMD are optional for Armv9-A and Armv8-A.

But both are mandated in cores for "rich operating systems", which basically means it's mandated by the OS.

Also, Armv9-A on OS level mandates SVE2.

camel-cdr··on RISC-V: They Should Have Known Better
Did a quick grep over object files of a half build defconfig kernel:

    total: 15510
    #0:  3720
    #31: 2247
    #1:  1349
    #21: 1208
    #2:  810
    #3:  524
    #8:  493
    ...
camel-cdr··on RISC-V: They Should Have Known Better
> Adding a flags register doesn't really add any more complexity, it's just a small bit of extra state attached to it.

I agree in general, we do however see that the cost of flags isn't free by the fact that most modern Arm processors only support ADCS on half of the ALUs supporting ADD. If it was free/negligible, you would see ADCS support on all ALUs.

camel-cdr··on RISC-V: They Should Have Known Better
No, debian requires ARMv8.0-A + FP + NEON, as those are optinal extensions (even optional in ARMv9.0-A)
camel-cdr··on RISC-V: They Should Have Known Better
> You shouldn't be sharing that

Ah, I suppose.

camel-cdr··on RISC-V: They Should Have Known Better
Edit: removed

Yeah, fusing is probably easier, if you already know what to fuse. On the other hand, if you want to fuse load pair on RISC-V you have the entire rename stage to figure out which uops can be fused independently of the rename stage, if fusion haopens after rename as well.

camel-cdr··on RISC-V: They Should Have Known Better
> While smaller cores have the option of cracking the multiple writeback instructions, many arm cores just pay the extra cost of having a 3 read, 2 write register file, so they aren’t actually cracking those instructions.

No, every high performance core I know of cracks them at decode, some re-fuse some of them after rename (Apple). Because otherwise you would need to rename up to 4 destinations per rename slot, effectively 4xing your already limiting rename stage.

Cracking other stuff later in the pipeline isn't expensive.

camel-cdr··on RISC-V: They Should Have Known Better
The decoder decodes them into two or more internal instructions (uops).

Take for example a post increment load, which does a=mem[b++], notice how this writes to two registers. Handeling two writes (up to 4) would explode the stage after decode (rename). So high performance arm implementations generate two uops for this. But since the number of decoders is fixed and the number of rename slots as well, you now have alnost the same problem as in RISC-V with compressed instructions: the nth input to the rename stage can come from a variaty of outputs of the decode stage, so you need a large shuffle network, and propagate the uop counts from start to end.

Cracking is a lot cheaper, if you can do it later in the pipeline. E.g. the cheapest is if you can simply "replay" the instruction. That is, instead of removing the entry from the issue queue, when it starts executing, you decrement a counter and keep the entry to do something else next. But as I mentioned that doesn't really work with multiple write back.

camel-cdr··on RISC-V: They Should Have Known Better
I want to try looking at codesize for -Os builds with the different ISAs including the compressed variants you mentioned. As well as dynamic icount with overflow checking.

Do you have any specific project in mind that I could use for testing?

camel-cdr··on RISC-V: They Should Have Known Better
https://github.com/riscv/riscv-isa-manual/pull/3269
camel-cdr··on RISC-V: They Should Have Known Better
The RVA point releases don't add new mandatory features, so every RVA23 complient board is also RVA23.1 complient. They only add new optional extensions.
Page 1 of 25Next →