The other way to look at it is that quite frankly most business services just don't hit a performance barrier where you're sitting and waiting on a string operation. You're usually waiting on some sort of I/O.
The other way to look at it is that quite frankly most business services just don't hit a performance barrier where you're sitting and waiting on a string operation. You're usually waiting on some sort of I/O.
I/O wait is difficult to model. On an NVME drive over PCIe 5 I doubt you'll find yourself waiting on I/O all that often, especially if you're reading a lot of data. Decompression tools, for example, are often written in such a way to keep the CPU pipeline fed. They're still doing lots of output but while some other core(s) deal(s) with that, the performance optimized core can still crunch numbers.
You're probably not storing hundreds of megabytes of JSON because there are much better formats for such data sets, but text algorithms like these make an impact regardless.
Loading CSV files, for example, can be a huge CPU bottleneck. I've had to deal with CSV datasets of over a gigabyte and all libraries were kind of terrible at it. The data source also had a JSON version that contained fewer fields for some reason, I guess to save space? Either way, an optimized escape code path would've probably saved seconds on each run of the data analysis, and more if the string split was optimized better as well. The data was loaded into memory already beforehand so the only I/O would've been exchanging data between the RAM and the CPU cache.
Anandtech:
"Critically, however, AMD is diverging from Intel in one important aspect: whereas Intel built a true, 512-bit wide SIMD machine for executing AVX-512 instructions, AMD did not. Instead, AMD will be executing these instructions over two cycles. This means AMD’s implementation still benefits from all of the additional instructions, register file space, and other technical improvements that came as part of AVX-512, but they won’t gain the innate doubling in SIMD throughput."
What is really known from the AMD disclosure is that Zen 4 has the same execution resources as Zen 3, only the load/store units are improved, perhaps their bandwidth has been increased to match that of all recent Intel CPUs.
Zen 3 has 4 AVX pipelines which can execute four 256-bit instructions per cycle, but some more complex instructions, e.g. multiplications or FMAs can be executed by at most 2 pipelines.
There are 2 possible ways to execute a 512-bit instruction in such pipelines, either the instruction can occupy a pipeline for 2 cycles, or it can occupy 2 pipelines for 1 cycle.
The throughput is the same, but the latency of the operation is different. The simpler and better way is to occupy 2 pipelines for 1 cycle. This is also how most AVX-512 instructions are implemented in all Intel CPUs, with the exception of FMA/FMUL, for which a few models of Intel CPUs have a second 512-bit pipeline and with the exception of some other instructions for which there is a 256-bit extension of one of the three 256-bit pipelines that exist in Intel CPUs, allowing the Intel CPU to do two 512-bit instructions per cycle, even if it can do only three 256-bit instructions per cycle.
The Intel CPUs can do two 512-bit instructions per cycle, except for a few instructions like FMA/FMUL that can be done only one per cycle in the cheaper CPUs, but two per cycle in most Xeon Gold, all Xeon Platinum and the Xeon W models with AVX-512.
The AMD Zen 4 is certain to have the same 512-bit throughput per clock cycle (two 512-bit instructions per cycle, of which only 1 can be FMA/FMUL) as all Intel CPUs with the exception of the models with two 512-bit FMA units, which will have double throughput only for FMA/FMUL.
Have you tried this one? https://lib.rs/crates/csvroll
It calls avx2 intrinsics directly for simd (unfortunately - it's an old library in Rust time, Rust nowadays has great support for portable simd)
Edit: After looking over the docs, it's not clear that Rust's portable SIMD exposes tbl or pshufb *at all*, which would basically prevent it from being used for string parsing.
If you care about raw JSON throughput, use Ragel or something similar to build a state machine that directly parses JSON into a flat, native data structure. Now you have zero allocations without even needing to shim malloc/free. AVX-512 would still be at least as useful, but it's a much more difficult problem to leverage SIMD in a parser generator than in a simple string escape routine or behind a more abstract interface like a regex library.
Quite a few language environments these days provide in-language JSON deserializers, but they're still significantly slower than they could be even when they deserialize to flat data structures. The macro languages and internal compiler intrinsics used to accomplish this are the worst possible environments for development. Lisp-like languages aren't really an exception as they tend to trade easier in-language transforms for a steeper climb when it comes to generating optimized native code for the transform.
For large files it is best to use a binary format that can be read quickly without parsing or allocation. https://rkyv.org/ is an example.
Being 'friendly' is not why JSON is popular. JSON is popular because the decoder is included the web browser.
I appreciate the thorough "shootout" benchmarks provided by the authors as well!
https://lemire.me/blog/2020/03/31/we-released-simdjson-0-3-t...
(Edit: just realized this is the same site as subj, heh)
Business logic is quite a bit more likely to just be waiting on IO, usually fetching from the DB or sending/receiving over the network. This is because there's a lot of business logic where those things are _everything it does_.
In a world with 400 Gb/s NICs, being IO bound takes some doing. How much SW manages that without SIMD, even on multiple cores?
depends on what you call I/O. From the perspective of the caller, that’s the entire stack starting at a http library call or sql driver call. If those can’t use the hardware to their fullest (or can’t, the way you’re using them) your process becomes I/O bound from its perspective.
Amen. CPU have stopped improving significantly a decade ago or more. IO never stopped.
Most apps and games could open instantly, unless you are doing something like heavy map generation or hitting the network. But it would require structuring the app specifically for it and there is often no incentive to do so.
Tiny files are fine if you fetch them from NVMe and keep the queues fed, i.e. you need to issue multiple concurrent reads, not do it sequentially. Not being on windows helps too.