Fast columnar JSON decoding with arrow-rs
arroyo.dev
arroyo.dev
(Actually, the original version of Arroyo was purely based around ahead-of-time code generation, and used serde_json for deserialization. I wrote at length why we decided to move away from that approach here: https://www.arroyo.dev/blog/why-arrow-and-datafusion).
Arrow-json, by contrast, is doing columnar deserialization against a schema that's only known at runtime. To me, the interesting aspects of the design are how it does that performantly.
For the "Tweets" case, it reports a speedup of 229%. The old value is 11.73 and the new is 5.108. That is a speedup of 2.293 (i.e. the new measurement is 2.293 times faster), but that is a difference of -56%, not 229%, so it's 129% faster, if you really want to use a comparative percentage.
Because using percentages to express ratio of change can be confusing or misleading, I always recommend using speedup instead, which is a simple ratio. A speedup of 2 is twice as fast. A speedup of 1 is the same. 0.5 is half as fast.
Formulas:
speedup(old, new) = old / new
relativePercent(old, new) = ((new / old) - 1) * 100
differenceInPercent(old, new) = (new - old) / old * 100