I've seen parsers where the same real task, say parse, map, and reduce a 100MB of doubles encoded as JSON. In a decent library it is taking much less than half a second and very little memory, say the size of the result and the memory mapping of the JSON document. In many common libraries it takes multiple seconds while using hundreds/gigs of memory. That means one has to pay for bigger machines with more uptime per task. That is giving money away.
That sounds like a very bad use case for JSON. I would be surprised if your program wasn't more efficient with an ad-hoc binary format for that piece of data.
It was a bit contrived true, but when some are doing it in 0.1s and many are in the 1.5-2s range, and that was parse time, not loading the data or startup. If I would go binary it would probably be something like protobuf, not ad-hoc. Ad-hoc has issues with maintenance, interop, and tooling.
Yes, I agree with that. It's a neat format if you have small pieces of information to move around, and it's very easy to read for humans, but for large enough data, wouldn't it turn into a bottle neck?
> If I would go binary it would probably be something like protobuf, not ad-hoc. Ad-hoc has issues with maintenance, interop, and tooling.
Indeed, there's always these dimensions to take into consideration, as well as evolution. The main issue is to find a library/format which is well supported across all sort of languages, and JSON has that. I don't know if there are many binary formats with the same level of support.
I suggested ad-hoc because the format seemed simple enough to be mmapped directly, or something equivalent (not sure how scripting languages would do in that case).
Fortunately, this is really only true at the application level. Space/line delimited byte strings and binary formats are used frequently at lower levels.
A little ETL goes a long way.
JSON is just an higher abstraction layer which might introduce errors.