To me, the biggest practical difference is whether your documentation is primarily written as a tutorial or for reference on demand. Presenting ideas in a natural order for tutorial purposes is useful to someone coming to the code for the first time, or if you’re coming back to look at a module a while after you first wrote it when all those little details are no longer so familiar.
In my case, the documentation is generated using LaTeX and so can also include maths, diagrams, tables and other illustrative material right there next to the associated code, as well as providing a natural place to put module- or program-wide summary information to give an overview of how everything fits together. Like a mathematical paper, it takes a little effort to present all this extra documentation well. However, it’s hard to overstate how much better it is if you’re trying to understand some intricate mathematical code that you wrote three months ago and the actual maths is right there and is then directly reflected in the shape of the code.
Others have mentioned that literate code might be harder to maintain in the long run. I’m not sure how realistic that really is, based on my experience so far. If you’re making code changes significant enough that you’d want to reorder the whole presentation, you’re probably rewriting significant chunks of that documentation anyway, and it’s not as if our editing tools can’t cope with moving code and/or text around.
What does suffer, significantly in my experience, is the scannability of the code. Those few thousand lines of Haskell I mentioned produce well over 100 pages of typeset documentation at this point. That’s partly because of the extensive textual notes and mathematics and diagrams and so on. It’s also partly because typeset documentation is naturally more spaced out because of things like headings and blank lines. But the fact remains, if I looked at just the source code in my usual editor, I’d probably have 50–100 lines visible in a single window, and I can open several of those windows at once on a big screen. If I’m looking through the literate documentation (or the source file from which it is generated) then I am probably only seeing one third to one half of that at most, and crucially, that code only appears a few related lines at a time, often a single function or a small family of related type definitions. Since Haskell itself is rather uniform in appearance and tends to be written by composing very short elements anyway, this makes finding and understanding individual fragments of code noticeably harder than regular coding when you want to refer back to something in isolation.
So far, I’m finding that a price worth paying, at least for this kind of heavily mathematical work with me as the sole developer. In practice, I don’t actually want to refer to a small code fragment on its own very often. I’m more likely to come back to a whole module, skim the entire literate documentation for it (probably just a few pages) to remind myself of how it all fits together, and then not need to jump around understanding small individual elements in isolation. Still, there is definitely a cost here, and it definitely affects how I read and understand the code as I’m working with it later. I suspect some of that cost would be incurred anyway by using Haskell, or any other language and programming style that emphasize composing many small elements, but using literate programming does exaggerate the effect, and so far I’ve found that to be its biggest drawback over more conventional styles.