We had a numeric counter that was enumerating the frame number of a simulation. When we removed the debug statement, we didn’t realize that no other code ever evaluated the counter - so instead of resolving each X+1 into an integer, Haskell kept it as a ever-growing nested series of unevaluated thunks: 1+1+1+1+1+1+1+1+1+1+... , and our efficient little program started filling up all its available RAM and crashing.
A good rule of thumb is to by default mark fields for basic types like Int, Bool, etc. as strictly evaluated, and leave larger structures (trees, lists, etc) lazy. But you still need to be careful.
The compiler trying to silently fix things is probably a bad idea. The current behavior is at least easy to understand; I'd hate to have a program that's working because the compiler could figure out that it could make something strict, and then I bump something mostly unrelatrd such that the optimizer can't be sure anymore, so I get a space leak.
Also, unintended evaluation can potentially cause high memory use as well (e.g. [1..1000000]), so the compiler also has to be careful about introducing excessive memory use.
The compiler does do some strictness analysis, but it's a hard problem.
I like the way idris does things -- strict by default, laziness controlled by the type system, and some nice support for automatic coercions.
The way I think about it is that it's not so much that your program executes a lazy stream of side effects, it's that your program doesn't execute side effects at all. Instead, your program returns a pure structure which describes how side effects should be executed, and the runtime executes a program based on this description. By the time the runtime is executing that program, the pure structure which describes how the side effects should be executed has already been forced, so the side effects happen eagerly within that context.
That said, I'm not much of a Haskell-er, so I'd be interested to hear from more experienced Haskell programmers on whether there's a better way to think about this.
I didn't actually say this. I was suggesting a lazy chain of IO actions for such problems. For example, ListT a IO, instead of a regular lazy list with hidden side-effects. The Haskell types help make the distinction clear.
> Instead, your program returns a pure structure which describes how side effects should be executed
Yes exactly, the pure structure can still be computed lazily on demand, providing IO actions.
My favourite Hackage streaming library for Haskell is "streaming" and it can work exactly this way.
Also, by it's nature ListT can't stream IO unless you use lazy IO aka lazyRead aka unsafeInterleaveIO which deserves its name.
newtype ListT m a = ListT (m (Maybe (a, ListT m a)))
It is a nice simple example of an effectful streaming type.
> Also, by it's nature ListT can't stream IO
This is only true of the deprecated one, which is also misnamed. The community is rightly re-using the name for the behaviour shown above. See list-t on hackage, or the implementation that comes with Pipes.
For example, this code will fail because the file handle is closed prematurely by withFile as hGetContents returns immediately without actually reading the file, deferring that part until the data is forced.
main = do
contents <- withFile "foo.txt" ReadMode hGetContents
putStrLn $ length contentsNot unlike the Streams/Collectors framework in Java.
That is, the only reason what you said is accurate, is because you said "effectful stream of bytes" meaning "eagerly effectful." Right? If you do all of that lazily, you can just as easily shoot yourself in the foot here.
(Please nobody take this as a damning criticism of Haskell. I'm not intending it that way.)
Returning the stream itself from the "withFile" and trying to run it later still is still a bug.
I like this approach better that conventional Haskell lazy I/O, which I personally find quite confusing.
You could argue that libraries shouldn't contain unsafe code either and go with a substructural type system, use indexed monads, fake regions or something similar. Whether the extra complexity is worth it is arguable, though.
If anything, eager evaluation will reveal the bug faster.