On a meta level, I wasn't even considering that the notion of a word might be different from "a sequence of characters that aren't space characters".
Live and learn indeed.
15 karma · joined January 9, 2020
On a meta level, I wasn't even considering that the notion of a word might be different from "a sequence of characters that aren't space characters".
Live and learn indeed.
> In practice it's unlikely that you find yourself in a situation where you can write code in a high-level language that runs considerably faster than what you could realistically write in C.
I do way more C++ (in fact, I don't do pure C at all), and aliasing has bitten me and my code performance more often than I'd like. While there are workarounds, I'd probably consider spending time and effort on them as rather unrealistic in a sense. So it surely doesn't contradict my world model if a language with a stricter type system (Haskell? Rust? ATS anyone?) achieves better results on at least some of the tasks with less effort and less dependence on implementation details.
Although ironically I'm going to write something low-level for the Haskell bytestrings library today evening, in C with intrinsics (so almost assembly modulo stuff like register allocation).
Was overly excited when I created the repo after obtaining the first results. Childish indeed, thanks for reminding, fixed.
> Also, I suggest you try running the GNU wc with unicode turned off because unicode is computationally expensive and you're deliberately disabling unicode support in your own code anyway.
I tried running wc as `LC_ALL=C wc file.txt`, and it (surprisingly for me) resulted in worse run time for wc (my default locale is ru_RU.UTF-8 for comparison). This reproduced on two machines of mine and also on a machine of a friend of mine who also gave my code a shot.
My bad for omitting this in the post, I'll update it accordingly.
That's precisely what the second part would be about.
And, if I succeed, IMO, that's where Haskell would really shine (because composability and local reasoning), and where I would be able to claim to achieve something — the stuff in the post we're discussing is indeed trivial (I didn't want to say that in the post itself though as I think it'll look like I'm belittling the guy who did the original post), while modularizing this is _fun_.
Also, I was trying to make a reverence to the original post that was "beating C", with the connotation of further improving on that. I'm not a native speaker so my language model might be terribly flawed.
Which actually raises a good question of whether I should have been comparing with that one — but that'd probably raise more questions and lead to more people accusing me of cheating in favour of Haskell.
> Does the Haskell version really do the same thing as the C version?
It counts bytes, words and lines and, modulo intended Unicode space handling and unintended bugs, does the same thing.
Indeed, it does not count things like max line length or char count, but those can be plugged in without significant performance overhead (and that's what the second part is gonna be about).
> Does it handle all of the same error cases, providing the same quality of error messages if they occur?
There are no error messages at this point. Although I don't really see how this should affect performance.
> Does it handle localization?
If you mean counting multi-byte characters, then not yet. Although I'm pretty convinced it does not require doing something much more complicated — but maybe I'm wrong, we'll see in the next part.
Also, thanks for the feedback, those are important questions! Something to keep in mind when writing subsequent posts.