Ode to a Streaming ByteString
blog.sumtypeofway.com
blog.sumtypeofway.com
I had the impression a few years ago that Pipes was the most cleanly designed of these schemes, but haven't been keeping track.
But [Char] is quite nice for certain uses.
>ByteString represents a byte buffer and its associated length; in this it is similar to Go’s []byte or Rust’s &[u8] (though it has one fewer datum to track, as the Rust and Go types offer mutable access to the associated byte buffer and thus must keep track of its total capacity).
Rust's &[u8] and &mut [u8] don't track capacity. Mutable access only allows changing the data in the slice, not changing its length. It couldn't even if it wanted to, since it has no idea of what the underlying storage is. It's only Go's slices that do double-duty as a ranged view into an existing allocation but also with the freedom to become the owner of a new allocation.
That Rust and Go both have "slices" that are slightly different is, unfortunate, but that's just how it goes sometimes.
Would it be something simple as:
writeStdout :: ByteStream IO a -> IO a
?
writeStdout :: ByteStream a -> IO ()
Maybe you could even leave out the `a` type since arguably a `ByteStream` should only contain bytes anyway. When writing to stdout you don't really need a return value usually, so that could just be (). The IO monad can encapsulate any amount of side effects in the same function, so you could fit your additional side effects in besides writing to stdout.mapM_ :: (ByteString -> m a) -> ByteStream m b -> m b
And then you just have
writeStreamToStdout = mapM_ writeBytesToStdout
At the same time, it's not a big deal to handle potentially infinite streams with strict I/O and explicit chunks.
I imagine it is possible to build projects around such libraries, possibly wrapping all the other I/O into it. Similarly to how in principle it's possible to unify all the error-yielding functions, and those working with text and/or byte streams. But I guess it doesn't happen often.
Hmm, I find the opposite - iteratee-style streaming libraries are pretty much the only way you can implement something like "open this file and transform it in this way" as a library function and ensure it gets used safely, because the stream can be responsible for resource management.
(I'm opposed to Haskell-style implicit laziness though, I agree with strict and explicit chunking being the way to go).
> But once you have a bunch of dependencies with their own (byte)string/text types, errors/exceptions, and ways to do I/O, all of which should be combined and work together, those become rather annoying.
Sure. I think the answer to that is that iteratees are pretty foundational and should be built into the language, in the same way that e.g. monad is not just a library type but something that's fundamentally part of the "platform" and maybe even the standard library. Efforts like the "Haskell platform" help us move in that direction.
lazy I/O (such as file reads) is suspect and even bad, but laziness + IO (or something like it) is a great way to stream from an API: http://blog.vmchale.com/article/lazy-io
Makes it possible to transcode for one. Basically makes streaming "compose" nicely - bzip2 and lzip otherwise have very different ways of streaming! And one would have to line a lot of things up.