HNHacker News
TopNewBestAskShowJobs

0xd34df00d

15 karma · joined January 9, 2020

submissionscomments
0xd34df00d··on Destroying C with 20 lines of Haskell: wc
I was mostly curious about how wc handles spaces and whether ignoring non-ascii spaces brings me closer or farther from what wc does. So I focused on that, and this specific printable characters handling didn't caught my eye.

On a meta level, I wasn't even considering that the notion of a word might be different from "a sequence of characters that aren't space characters".

Live and learn indeed.

0xd34df00d··on Destroying C with 20 lines of Haskell: wc
No worries, that's a natural reaction!

> In practice it's unlikely that you find yourself in a situation where you can write code in a high-level language that runs considerably faster than what you could realistically write in C.

I do way more C++ (in fact, I don't do pure C at all), and aliasing has bitten me and my code performance more often than I'd like. While there are workarounds, I'd probably consider spending time and effort on them as rather unrealistic in a sense. So it surely doesn't contradict my world model if a language with a stricter type system (Haskell? Rust? ATS anyone?) achieves better results on at least some of the tasks with less effort and less dependence on implementation details.

Although ironically I'm going to write something low-level for the Haskell bytestrings library today evening, in C with intrinsics (so almost assembly modulo stuff like register allocation).

0xd34df00d··on Destroying C with 20 lines of Haskell: wc
> For reference, I'm referring to the "oops I did it again" part. It's really hard to take that comment as "honours go to GHC authors".

Was overly excited when I created the repo after obtaining the first results. Childish indeed, thanks for reminding, fixed.

> Also, I suggest you try running the GNU wc with unicode turned off because unicode is computationally expensive and you're deliberately disabling unicode support in your own code anyway.

I tried running wc as `LC_ALL=C wc file.txt`, and it (surprisingly for me) resulted in worse run time for wc (my default locale is ru_RU.UTF-8 for comparison). This reproduced on two machines of mine and also on a machine of a friend of mine who also gave my code a shot.

My bad for omitting this in the post, I'll update it accordingly.

0xd34df00d··on Destroying C with 20 lines of Haskell: wc
This is a great point about handling printable vs non-printable characters that I originally missed when I read wc code. Thank you for pointing this out!
0xd34df00d··on Destroying C with 20 lines of Haskell: wc
> So until you put in locale handling, alternate line endings, option handling, and error handling, I don't see that your post is at all convincing.

That's precisely what the second part would be about.

And, if I succeed, IMO, that's where Haskell would really shine (because composability and local reasoning), and where I would be able to claim to achieve something — the stuff in the post we're discussing is indeed trivial (I didn't want to say that in the post itself though as I think it'll look like I'm belittling the guy who did the original post), while modularizing this is _fun_.

0xd34df00d··on Destroying C with 20 lines of Haskell: wc
Me neither. But experience shows that these titles lead to more folks looking at the post, leading to more feedback, leading to better writing/experimenting/etc in the long run.

Also, I was trying to make a reverence to the original post that was "beating C", with the connotation of further improving on that. I'm not a native speaker so my language model might be terribly flawed.

0xd34df00d··on Destroying C with 20 lines of Haskell: wc
wc was actually slower with LC_ALL=C as opposed to ru_RU.UTF-8 that my system normally runs with (about 10 s against 7.2 s).

Which actually raises a good question of whether I should have been comparing with that one — but that'd probably raise more questions and lead to more people accusing me of cheating in favour of Haskell.

0xd34df00d··on Destroying C with 20 lines of Haskell: wc
Those are fairly trivial and well-known optimizations that I did (and I by no means am an expert in writing high-performant code), so all the honors go to GHC authors.
0xd34df00d··on Destroying C with 20 lines of Haskell: wc
Sup, author here.

> Does the Haskell version really do the same thing as the C version?

It counts bytes, words and lines and, modulo intended Unicode space handling and unintended bugs, does the same thing.

Indeed, it does not count things like max line length or char count, but those can be plugged in without significant performance overhead (and that's what the second part is gonna be about).

> Does it handle all of the same error cases, providing the same quality of error messages if they occur?

There are no error messages at this point. Although I don't really see how this should affect performance.

> Does it handle localization?

If you mean counting multi-byte characters, then not yet. Although I'm pretty convinced it does not require doing something much more complicated — but maybe I'm wrong, we'll see in the next part.

Also, thanks for the feedback, those are important questions! Something to keep in mind when writing subsequent posts.