HNHacker News
TopNewBestAskShowJobs

yoricksijsling

55 karma · joined January 5, 2023

submissionscomments
yoricksijsling··on Parallel streaming in Haskell: Part 4 – Conditionals and non-blocking evaluation
I never really used Arrows, let alone stream transformers with an Arrow interface, but yes that makes sense. With the conduit library you’d use ZipConduit which has the same characteristics.

On reddit, in reply to the first blog post, someone asked how ZipConduit compared to our parallel streams. Maybe that’s helpful for hosh as well. https://www.reddit.com/r/haskell/comments/10339fe/comment/j2...

yoricksijsling··on Parallel streaming in Haskell: Part 4 – Conditionals and non-blocking evaluation
Arrows don’t really _do_ anything. Just like monads, they’re just a generic interface and their behaviour depends entirely on the instance implementation.

You can have stream transformers that fit in de arrow class, and parallel composition might then mean that things are done concurrently. How that compares to our approach entirely depends on the exact stream transformer implementation.

Note that the ConduitT type that we use doesn’t really fit in the Arrow class, due to the additional parameter for the return type.

yoricksijsling··on Parallel streaming in Haskell: Part 4 – Conditionals and non-blocking evaluation
Oof, I don’t think that’s intentional. I’ve passed it on to the frontend team!
yoricksijsling··on Parallel streaming in Haskell: Part 1 – Fast, efficient, and fun
Precisely. Conduits basically do everything on demand.

For parallel processing the backpressure becomes a bit more complicated, in particular if you end with a sequential consumer. If you're just doing your parallel tasks when a sequential consumer demands it, you'll find yourself still doing one task at a time. The straightforward solution is to work ahead a bit, spawning just enough parallel work to fill a small buffer that the consumer can read from. You get backpressure this way, but may have done a little bit too much if the consumer doesn't want the rest of the data. We'll show this in the third blog post of the series.

yoricksijsling··on Parallel streaming in Haskell: Part 1 – Fast, efficient, and fun
> All the discussion of GC-friendliness of caches makes me think that there aren't any conduits that have more data than can fit in memory on a single machine.

Yep! That's currently still the case. We do have some ideas to put caches on disk using mmapped files, so that you have fast access when it all fits in memory but the OS can also drop them when it wants to.

For the moment we just use instances with 128GB memory, and those can still fit the datasets of the biggest customers that we have right now. Datasets go up to 10s of GBs, so it's not really 'big data' that needs to be distributed across machines. Due to the regular occurrence of aggregations (sort/group/window/deduplicate) that require at least _some_ synchronization, we only have small sections that can be completely parallelized. It's already a challenge to use all the cores on a single machine in an efficient manner, and I don't think we'll achieve much by using multiple machines for a single job. We've discussed a lot of this in an earlier blog post here: https://www.channable.com/tech/why-we-decided-to-go-for-the-...

> Separate question, do you have any actions that produce more output than their input?

Yep, we don't have cartesian products but we do have a 'split' action. Typical usage might be that a customer has a "sizes" field with values like "M,L,XL" and that they would split on that field so that a single item becomes three items for the separate sizes. The increase in the number of items is usually limited, and the increase in memory usage is even smaller because at most points during the data processing we only store the changed fields, and refer to the original item for the rest. In these cases multiple items will point to the same original item.