JSON Parsing: Intel Sapphire Rapids versus AMD Zen 4
lemire.me
lemire.me
Can someone ELI5 how I can run these benchmarks?
I cloned the repository, and tried to compile the bench_ondemand file, but it complains about missing header files, which are in the repository, but under the include/ directory. And I couldn't find any documentation about this
EDIT: okay found some instructions for an old version, but these still work. if someone is interested:
https://openbenchmarking.org/innhold/a192889420a87740daf3c7c...
simdjson was not as performant as the author's intel or AMD results, but yyjson gave way better results than the author had.
I do wish there was a newer compiler in use. GCC 11 was from March 2021. Using a nearly 3 year old compiler feels like it's leaving hard work & tuning on the table, isnt indicative of what these professors might be capable of.
https://developer.arm.com/documentation/dui0801/h/A64-Floati...
I just did the usual git clone, mkdir build, and: cmake .. -DCMAKE_BUILD_TYPE=Release -DSIMDJSON_JUST_LIBRARY=OFF
To get the benchmark results:
~/git/simdjson/build/benchmark$ cat log | grep partial_tweets | grep "simdjson_ondemand" partial_tweets<simdjson_ondemand>/manual_time 66705 ns 75966 ns 10453 best_bytes_per_sec=9.59662G best_docs_per_sec=15.1962k best_items_per_sec=1.51962M bytes=631.515k bytes_per_second=8.81706G/s docs_per_sec=14.9913k/s items=100 items_per_second=1.49913M/s [BEST: throughput= 9.60 GB/s doc_throughput= 15196 docs/s items= 100 avg_time= 66705 ns]
partial_tweets<simdjson_ondemand>/manual_time 107280 ns 132468 ns 6469 best_branch_miss=376 best_bytes_per_sec=8.10705G best_cache_miss=1 best_cache_ref=553 best_cycles=342.737k best_cycles_per_byte=0.542722 best_docs_per_sec=12.8375k best_frequency=4.39987G best_instructions=1.42334M best_instructions_per_byte=2.25385 best_instructions_per_cycle=4.15285 best_items_per_sec=1.28375M branch_miss=378.313 bytes=631.515k bytes_per_second=5.48235G/s cache_miss=0.655434 cache_ref=609.834 cycles=352.768k cycles_per_byte=0.558605 docs_per_sec=9.32143k/s frequency=3.2883G/s instructions=1.42334M instructions_per_byte=2.25385 instructions_per_cycle=4.03477 items=100 items_per_second=932.143k/s [BEST: throughput= 8.11 GB/s doc_throughput= 12837 docs/s instructions= 1423337 cycles= 342737 branch_miss= 376 cache_miss= 1 cache_ref= 553 items= 100 avg_time= 107279 ns]
Given in the article Xeon 8488C at 3.4GHz got 6.83 GB/s, and my 3495X at 4.4 GHz got 8.11 GB/s (1.29 times CPU freq, 1.18 times throughput), it does look like the throughput scales to the CPU frequency (although not entirely linear)
For JSC at least the costs of allocating and constructing the JS object graph rapidly dwarf the cost of the parsing itself for any non-trivial JSON structure (though parsing JSON is still sufficiently fast relative to actual JS that there's still a significant win for all JS execution in JSC to start with an attempt to parse the input as a slightly extended JSON syntax, due to the wonders of JSONP and people still using eval for json parsing).