We sped up time series by 20-30x
rerun.io
rerun.io
That being said, the engineering of Rerun fills me with joy. This thing is just crazy fast and really usable. Absolute gem.
Btw @Rerun folks you can update your Readme on Github and remove the time series shortcoming now :). Glad to see that landed!
On the topic of Readmes though, their Github also still says that their beta will be publicly available on February 15th (I guess they don't mean tomorrow). In any case I reckon they've got their priorities sorted.
Haha, yeah, we open sourced Rerun a year ago, minus one day I guess I’ll need to update that readme
There is another order of magnitude performance left to squeeze out in theory so I guess we can always leave it in until that’s done too
With a 20x-30x improvement in performance I really have to ask:
How did you manage to fuck up your code's design so badly ...
... that you've left so much performance on the table?
It seems like the modern database blog meta is to brag about the basics like they are a new discovery.
I was just about to write down that it would be too difficult to de RDP simplification in the (ClickHouse) database, but then I recalled PostGIS has it built in and low and behold, there is also something in ClickHouse for this [0]. Back to the drawing board.
[0] https://clickhouse.com/codebrowser/ClickHouse/contrib/boost/...
I'm not super-familiar with the Ramer–Douglas–Peucker algorithm itself but I've used implementations of it, and, from the looks of it, its CPU cost would largely be offset by the savings in triangulation done by egui's renderer (also done on CPU currently).
Given a first and last point, it finds the point furthest away from a straight line connection, then recursively divides down the pairs of (first, furthest) and (last, furthest) only if the furthest point is above a minimum threshold distance from a straight line connection.