HNHacker News
TopNewBestAskShowJobs

ajeetdsouza

254 karma · joined February 16, 2019

meet.hn/city/12.9767936,77.590082/Bengaluru

Socials: - github.com/ajeetdsouza - linkedin.com/in/ajeetdsouza

submissionscomments
ajeetdsouza··on Crafting Interpreters with Rust: On Garbage Collection
This is really well-written!

Shameless plug: you may want to check out Loxcraft: https://github.com/ajeetdsouza/loxcraft

I too followed the path of "ignore safe Rust for maximum performance". It got pretty close to the C version, even beating it on some benchmarks.

ajeetdsouza··on GitHub – nushell/nushell: A new type of shell
Shameless plug: zoxide recently added support for Nushell. It's a smarter cd command for your shell (similar to z or autojump).

https://github.com/ajeetdsouza/zoxide

ajeetdsouza··on Beating C with 70 lines of Go
Fixed, thank you!
ajeetdsouza··on Beating C with 70 lines of Go
The size of a compiled Hello World program is 2 MB on my machine, so you're probably right!
ajeetdsouza··on Beating C with 70 lines of Go
The tests were run 10 times, and I used the median value. There wasn't much variance between runs, so I don't think disk caching played much of a role here.

Being a garbage collected language with a runtime, Go certainly cannot match the performance of C, and it was never my point to prove otherwise - obviously, for the same algorithm, the C implementation would be faster. Instead, I was exploring Go to highlight how simple it is to write safe, concurrent code in it.

ajeetdsouza··on Beating C with 70 lines of Go
Author here. I think you should go through the article again. I think it's quite readable, and there are no "hand-optimizations" as you say. Also, the single-core implementation was already faster than the C version - the multithreaded version was only done to explore different methods of concurrency in Go.

Hope that clarifies things.

ajeetdsouza··on Beating C with 70 lines of Go
I tested this on a fresh install of Fedora 31, so I didn’t really see any benefit of running it on a LiveUSB. As I mentioned in the article, the wc implementation I used for comparison has been compiled locally with gcc 9.2.1 and -O3 optimizations. I’ve also listed my exact system specifications there. I’ve used the unprocessed enwik9 dataset (Wikipedia dump), truncated to 100 MB and 1 GB.

I understand your frustrations with the previous posts, but I’ve tried to make my article as unambiguous as possible. Do give it a read, if you have any futher suggestions or comments, I’d be happy to hear them!

ajeetdsouza··on Beating C with 70 lines of Go
Argument parsing is hardly the bottleneck when you're scanning 1 GB of data. The reason I didn't implement the rest of the features was to keep the implementation light and readable, and I don't think a fully compliant wc implemented in Go with these patterns would be significantly slower.

As for having a low-memory footprint, 2 of the 4 implementations I mentioned (one single core, one multi-core) consume less memory than wc, and still run much faster.

ajeetdsouza··on Beating C with 70 lines of Go
Author here. I have addressed this in the article. The bufio-based implementation was the first one, and it was actually slower.

In the second section, I was able to surpass the performance of the C implementation - by using a read call with a buffer. As I mentioned in the article, the C implementation does the same, and in the interest of a fair benchmark, I set equal buffer sizes for both.

ajeetdsouza··on Beating C with 70 lines of Go
This is not true, see my comment here: https://news.ycombinator.com/edit?id=21587907
ajeetdsouza··on Beating C with 70 lines of Go
Author here. This is not true - I included a link to the manpage (https://ss64.com/osx/wc.html) in the article to avoid this confusion. I did not use GNU wc; I used the OS X one, which, by default, counts single byte characters. From the manpage:

> The default action is equivalent to specifying the -c, -l and -w options.

> -c The number of bytes in each input file is written to the standard output.

> -m The number of characters in each input file is written to the standard output. If the current locale does not support multi-byte characters, this is equivalent to the -c option.

Moreover, I also mentioned in the article that I was using us-ascii encoded text, which means that even -m would have been treated as ASCII text.

Hope that clarifies your issue.

ajeetdsouza··on Beating C with 70 lines of Go
Author here. Why do you feel it is code golf? My primary focus when writing it was readability - I'm sure I could do it in much less than 70 lines if I had to. If you have any suggestions for improving the readability of my implementation, do let me know!
ajeetdsouza··on Beating C with 70 lines of Go
Author here. If you read the article carefully, you'll see that the 70 lines I used to outperform wc was a single-threaded implementation. I multi-threaded it later for overkill.
ajeetdsouza··on Beating C with 70 lines of Go
Hey, author here. You are absolutely right - Go would have a hard time outperforming finely tuned C. Rather, the article was more focused on exploring concurrency mechanisms in Go. The (admittedly clickbait) title was more of a reference to the articles before it than anything else.

That said, I think you'll find this SIMD-enhanced wc interesting: https://github.com/expr-fi/fastlwc/