ojc_parse_str 1000000 entries in 709.300 msecs. ( 1409 iterations/msec)
simdjson_parse 1000000 entries in 450.724 msecs. ( 2218 iterations/msec)
In those results simdjson is roughly 60% faster.Telling simdjson we have aligned input gives us further improvement:
simdjson_parse 1000000 entries in 369.234 msecs. ( 2708 iterations/msec)
About 90% faster. Example changes: simd_parse(const char *str, int64_t iter) {
simdjson::dom::parser parser;
int64_t dt;
+ auto padded = simdjson::padded_string(std::string_view(str));
int64_t start = clock_micro();
for (int i = iter; 0 < i; i--) {
- simdjson::dom::element doc = parser.parse(str, strlen(str));
+ simdjson::dom::element doc = parser.parse(padded);
}
dt = clock_micro() - start; ojc_parse_str 1000000 entries in 3607.615 msecs. ( 277 iterations/msec)
simdjson_parse 1000000 entries in 418.997 msecs. ( 2386 iterations/msec)
and might as well throw in my own parser... uj_parse 1000000 entries in 1959.731 msecs. ( 510 iterations/msec)
The -O3 seems to make a large difference for simdjson.Thank you for being civil with your reply. Much appreciated.
simdjson relies deeply on inlining to let us write performant code that is also readable.
Sorry to have sent you down a blind alley!
One thing to note: if you want to get good numbers to chew on, we have a bunch of really good real world examples, of ALL sizes (simdjson is about big and small), in the jsonexamples/ directory of simdjson. And if you want to check ojc's validation, there are a number of pass.json and fail.json files in the jsonchecker/ directory.
UPDATE and your test data is quite short. You're maybe mostly measuring startup overhead?
As for the length of the test, I also tried with 10 times the number of iterations with the same results but then the simbbench takes almost a minute to complete. I figured 5 seconds was good enough to get the results and someone can bump up the iterations if they want to.
The test is not realistic. Measuring performance on small cases is a legitimate use case, of course, but using the same data every time means that both ojc and simdjson will train the branch predictor almost perfectly on any modern architecture. This will help either parser during its "branchy" code, which is all of ojc and a significant portion of simdjson (despite all our yelling about how great SIMD is, once we have tokens, we go to a fairly traditional state machine).
On a tiny message, startup costs are probably more significant. Fixing buffer alignment, -O3, and working with larger messages I would not be shocked to see a factor of 20x. There's nothing wrong with ojc but it is a typical parsing strategy that isn't all that different from the other 'normal' JSON parsing libraries that we evaluate (RapidJSON, sajson, dropbox, fastjson, gason, ultrajson, jsmn, cJSON, and JSON for Modern C++).
It's not clear to me why you were so convinced that your library, by extension, would run 10-25x faster than all these other libraries. With the exception of RapidJSON and sajson - which have some performance tricks to go faster - most of these libraries read input in the same way.
EDIT: Actually, there is a parameter 'realloc_if_needed' that defaults to true and that will reallocate the string because it assumes you haven't padded it. You're likely just measuring malloc performance.
Most real world applications like sockets read into a buffer, and can easily meet this requirement.
If you are interested to know what it's for, the place where it parses/unescapes a string is a good example. Instead of copying byte by byte, it generally copies 16-32 raw bytes at a time, and just sort of caps it off at the end quote, even though it might have copied a little extra. Here's some pseudocode (note this isn't the full algorithm, I left a out some error conditions and escape parsing for clarity):
// Check the quote
if (in[i] == '"') {
i++;
len = 0;
while (true) {
// Use simd to copy 32 bytes from input to output
chunk = in[i..i+32];
out[len..len+32] = chunk;
// Note we already wrote 32 bytes, and NOW check if there was a quote in there
if (int quote = chunk.find('"')) {
len += quote;
break;
}
len += 32; // No quote, so keep parsing the string
i += 32;
}
}I wish there was a standardized attribute that C++ knew about that pretty much just said "hey, we're not right next to some memory-managed disaster, and if you read off this buffer, you promise not to use the results".
It is awful practice to read off the end of a buffer and let those bytes affect your behavior, but it is almost always harmless to read extra bytes (and mask them off or ignore them) unless you're next to a page boundary or in some dangerous region of memory that's mappped to some device.
This attribute would also need to be understood by tools like Valgrind (to the extent that valgrind can/can't track whether you're feeding this nonsense into a computation, which it handles pretty well).
Also, number of iterations is not the issue. Length of the input is. What happens if you r json is a few kb, is bigger than l1/l2 cache, ....
I have no doubt a more complete set of benchmarks could be made with various sized JSON and from file as well as string. This was put together quickly to get an answer to the questions raised in this post. Since it does seem to be a topic of interest I'll expand and cleanup the tests in the future but for now it does give at least one data point.
You might notice that the somdjson::dom::parser is reused. Without that optimization to allow warming up the iteration/msec was only 160.
Given your response, I spent 10s extra reading the simdjson docs and noting you violate the SIMDJSON_PADDING requirement. So either your code is a crash waiting to happen, or you use a very non- optimal code path in simdjson that requires it to re-copy all data.
That's also the maximum amount of attention you'l get from this random stranger. my time is up ;-)
As for the path in simdjson being non-optimum and having to copy bytes, that should be expected if the string is to be modified. The buf argument type is a const after all so it should not be modified. In any case, glangdale's code is clean and solid so I doubt his code is anything but optimized and the examples correct.
Suppose there is a socket with json data. What you do is, at init you create 1 or more buffers of, say, 64kb+required padding. Then, when epoll or whatever says there is data, you call read on this buffer. This gives you a length.
At this point, you have a padded buffer and a length, so requirements for simdjson(reallocifneeded=false) are met, so you can now parse at full speed. When done, reuse the buffer for the next epoll/read cycle.
There are complications, of course. Data is not guaranteed to arrive all together in 1 read call. There might be a http library feeding you. All of this amounts to mostly buffering and chunking. When carefull, they can be solved in a mostly zero copy way, ready for optimal simdjson.
The example simdjson code you refer to is a kind of demo mode. It gets you of the ground quickly, but is far from optimal.
I assume simdjson does not write to the buffer. It just reads more than 1 byte at once, so if it reads say 8 bytes and there is only 1 left, it will read 7 bytes of random junk. And discard them when it notices its unheeded optimism. However, it needs to be allowed to read these extra bytes without causing a SISEGV, hence the padding.
UPDATE all of this an educated guess, the simdjson authors are welcome to fix/finetune whatever I said.
That being said, simdjson does take an interesting approach and for just pulling out specific parts of a JSON document it is very fast.
simdjson is designed to do the opposite of what you describe - nearly all the intellectual work in simdjson was focused on parsing a whole document in one hit, not on "pulling out specific parts of a JSON document".
So, I'm quite unclear on how you could derive a measurement that made us look fast on pulling out specific parts of a JSON document. That being said, I don't have a lot of experience with the current simdjson DOM traversal code and maybe there's some terrible "buried treasure" there (I mean sarcastically - i.e. we have some dumb decisions, maybe?) that you're tripping over. But the one-shot "parse" should be fast.
Anyhow, if you could try to code up something to show parsing speeds on workloads similar to what we describe at https://github.com/simdjson/simdjson it would be very interesting. We've made the effort to make all our benchmarks repeatable - if your approach is faster it would be cool and surprising.
That latter task is a valid use case, and interesting in itself, but not something we really knew how to measure and write about. The problem is with those sort of benchmarks is that they tend to resemble a "ask yourself a question then answer it" - I couldn't think of a good way of coming up with a set of "let's query a JSON file for a specific thing" benchmark that didn't seem ludicrously contrived.