One Program Written in Python, Go, and Rust
nicolas-hahn.com
nicolas-hahn.com
And,on top of all that, once you've written code satisfying the image interface, the standard library even includes functions for saving your image as one of several possible image formats. And, due to the fact that the interface itself is specified by the standard library, virtually all third-party image encoding or decoding libraries for Go use it, too. So, every image reader or writer I've seen for Go, even third-party ones, can be drop-in replacements for one another.
Anyway, it's not Go's standard use case, but as someone who loves fractals and fiddles with images all the time it's one of my favorite parts of the language.
If you need speed don't use the standard library, use something specialized.
This is a great example I didn't even know about that reinforces why I like Go. We don't have to over engineer everything.
http://gs.statcounter.com/screen-resolution-stats/desktop/wo...
In gaming they are just at 1.61% right below 1280x1024 and the increase is so low (+0.01) that might as well be zero (compare with 1080p's +2.04% which is the one increasing the most):
https://store.steampowered.com/hwsurvey/
Tech minded people are a bubble, gamers are a tiny bubble among those and /r/pcmasterrace 4K-or-die boasters are a tiny bubble among gamers. 4K, or even 1440p, matters way less in practice than tech minded people think.
Unless your language introduces an unreasonable overhead, a for loop over the pixels is perfectly appropriate and fast.
This sounds like a limitation of a particular optimizing compiler/interpreter rather than a problem of the language itself. For example, the plain lua interpreter incurs quite a lot of overhead for this, but the luajit interpreter is oblivious. The standard python interpreter definitely adds a lot of overhead.
But I'd also point out that the Go standard library does not necessarily claim to the "last word" for any give task; it's generally more an 80/20 sort of thing. If you've got a case where that's an unacceptable performance loss, go get or write a specialized library. There's nothing "un-Go-ic" about that.
All the examples, in each language, could be rewritten as a cache oblivious algorithm to optimise cache usage. This would speed them all up. See https://en.wikipedia.org/wiki/Cache-oblivious_algorithm
But I have actually benchmarked this before, and it is possible to have a function body so small (like, for example, a single slice array index lookup and return of a register-sized value like a machine word) that the function call overhead can dominate even so.
(Languages like Erlang and Go that emphasize concurrency have a constant low-level stream of posts on their forums from people who do an even more extreme version, when they try to "parallelize" the task of adding a list of integers together, and replace a highly-pipelineable int add operation that can actually come out to less than one cycle per add with spawning a new execution context, sending over the integers to add, adding them in the new context, and then synchronizing on sending them back. Then they wonder why Erlang/Go/threading in general sucks so much because this new program is literally hundreds of times slower than the original.)
But it is true this amortizes away fairly quickly, because the overhead isn't that large. Even the larger random number generators like the Mersenne Twister will be a long ways towards dominating the function call overhead. I don't even begin to worry about function call overhead unless I can see I'm trying to do several million per second, because generally, you can't do several million per second on a single core because the function bodies themselves are too large and doing too much stuff, such that even if function call overhead was 0 it would still be impossible in the particular program to do it.
Java does have similar things, but, as far as I know, they're limited to the UI libraries and far more complicated (e.g. https://docs.oracle.com/javase/7/docs/api/java/awt/Image.htm...). Imagine if everyone who read or wrote image files in Java were guaranteed to conform to a dirt-simple "Image" interface with only three methods--with optimizations available but strictly optional.
For someone who's just tinkering or just wants to dump some cool bitmaps, the ability to write an "At(...)" function in Go and be looking at a PNG within a few seconds is just great fun to have in the standard library.
And yet, in the author's example, all memory handling in Rust was completely automatic. And that includes, AFAICT, no Box'd pointers¹, no ref-counting, and certainly no raw/unsafe pointers.
IME, this seems to be a common response among people coming from GC'd languages; I think the expectation is that they're going to be doing C-style memory management (manually alloc/free pairs), when the truth is that 99.9% of allocations will just happen automatically and invisible thanks to RAII.
In the end, I really think it's resources, of which memory is just one type of, that matter. Python doesn't do anything for resources, and you have to know you're allocating something (e.g., files, locks, connections, connection leases, etc.) that will require a manual .close() or with statement s.t. it gets dealloc/cleaned/released; RAII will handle this just like memory, and automatically handle it, and because of that, I find myself doing less resource management in Rust than I do in Python.
¹There are conditions in which I would argue that some uses of Box are "automatic", depending on the reason it's pulled in. E.g., I've used it to reduce the size of an enum in the common case, but still allow it to store a heavy structure in the rare case. The handling of the Box itself is essentially still automatic.
The "with" blocks can help with certain kinds of resources, for example:
with open('foo.txt', 'r') as inf:
data = inf.readlines()[1] https://docs.python.org/3/reference/datamodel.html#with-stat...
The thing about `with` (and similar constructs in other languages) is that you must remember to use them at every single site of use. If you forget, while the resource usually gets cleaned up when the language's GC decides to finalize the object, it's non-deterministic and might be too late. The consequences of forgetting are significantly better vs. C, as the GC will usually get to it in time, but that's not something I'd like to depend on. But the amount of "effort" and possibilities for mistakes is, generally, the same as C (in number, not severity): each resource allocation and each resource free requires the programmer's attention — excepting the specific, and I grant, quite common, case of memory, which one can say the GC will handle for you. RAII makes it s.t. you don't have to remember except in the (hopefully rare) case of implementing a new resource, and you can treat variables holding resources at sites-of-use like any other variable.
For your specific example, if you can do w/o the list of lines (which I find is often possible) Python's actually got a somewhat nicer construct w/ Path:
pathlib.Path('foo.txt').read_text()
This only reads the data into a single string, whereas yours is a list of strings. Slight different, but usually it is acceptable. But it moves the required `with` into Path's helper method, so then it's harder for you to forget it. But this is a specific case (reading all data in a file), and doesn't generalize.I really appreciate the fact that Rust does it automatically (and that it's not easy to turn that automatic management off).
I'd also be curious to know is pillow-simd [2] gets the Python performance closer to Go/Rust, and if using Rayon [3] and changing your .iter()'s in your Rust code to .par_iter()'s will yield an improvement there.
[1] https://github.com/sharkdp/hyperfine
image1
but instead from image1.raw_pixels().iter().zip(image2.raw_pixels().iter()
I would have created a local for that before the loop for clarity. And I would not have declared the ratio variable, that is surely unidiomatic as the last expression provides the result.But apart from those stylistic quibbles it seems a model of clarity, the opposite of murky, whereas the Python version simply glues together opaque library calls.
let diffsum: u64 = image1.raw_pixels().iter()
.zip(image2.raw_pixels().iter())
.map(|(&p1, &p2)| u64::from(abs_diff(p1, p2)))
.sum();
I find this a lot easier to read because there is no control flow in this snippet. I can see that it is just summing the diffs. I know some people prefer iterators and some prefer loops, however I think mixing them like was done often leads to less readable code.What I like to do is to make the last arg take a struct such that you indirectly have named (and optional) arguments:
type DownloadOpts struct {
UserAgent string
TimeoutSeconds int
//...
}
func Download(url string, opts DownloadOpts) (io.Reader, error) {
...
}
//use like:
contents, err := Download("https://news.ycombinator.com", DownloadOpts {
UserAgent: "example-snippet/1.0",
})
> Never needing to pause for garbage collection could also be a factor [in Rust's greater speed compared to Go].Would be nice if the author had checked
var ms runtime.MemStats
runtime.ReadMemStats(&ms)
print(ms.NumGC)
to see if there was actually any garbage collection performed.Rob Pike, and Dave Cheney both posted about this [1][2]. They summarised that using self-referential functions were a more optimal way for handling options to a function. This gives the benefit of allowing the options to be easily extensible by yourself, and any users that would be interacting with your API.
[1] - https://commandcenter.blogspot.com/2014/01/self-referential-...
[2] - https://dave.cheney.net/2014/10/17/functional-options-for-fr...
for x in 0..w {
for y in 0..h {
let mut rgba = [0; 4];
[...]
to let mut rgba = [0; 4];
for y in 0..h {
for x in 0..w {
[...]
cuts down runtime by between 35 and 50% on my side.EDIT : moving the call to get_pixel outside the inner loop take off another ~10%, bringing it to sub 0.145 from ~0.290.
let mut rgba = [0; 4];
for y in 0..h {
for x in 0..w {
let pix1 = image1.get_pixel(x, y);
let pix2 = image2.get_pixel(x, y);
for c in 0..4 {
rgba[c] = abs_diff(
pix1.data[c],
pix2.data[c],
);
}
let new_pix = image::Pixel::from_slice(&rgba);
diff.put_pixel(x, y, *new_pix);
}
}At least that is my understanding of how to implement a cache oblivious algorithm. See https://en.wikipedia.org/wiki/Cache-oblivious_algorithm
Used that approach once when computing a large cross product and it gave a good speed up.
- allocating/deallocating memory
- reading and writing values from/to memory locations
- performing arithmetic and logical operations on values
However, what the machine is actually doing can differ substantially from these naive models. An algorithm that should be slower according to the C model can in reality be fast. Registers, cache levels, vectorization, etc are all concepts that C doesn't teach. Compilers often have extensions to C that enable more direct access (and also do things like adjust memory layout) but at that point you're no longer using the C model.
Optimizing compilers (and even cpus themselves) also do a good job of trying to uphold the abstraction but at the end of the day it's still leaky.
So knowing what the machine is actually doing is helpful.
This is one of the biggest confusion points when comparing Python and Go.
Go's static type system is weak, the weakest of any widely-used static-typed languages except C/C++. Lack of interfaces means you have to cast in/out of object, so you're constantly sidestepping the supposed benefits of the static type system. Type errors will, of course, be caught at runtime, but that's completely identical to Python. You can, of course, use go generate to avoid a lot of the object casting, but then you get all the wonderful problems of macros.
Python doesn't enforce types at compile time (because there isn't a compile time) but it does enforce types--it does not "nod and seem to understand you", if, for example you decide to do `42 + "Hello, world"`. It very much tells you that this is not valid.
This confusion comes from the common misconception that static types = strong types, and dynamic types = weak types. The truth is that there are a good number of static languages (C being the most obvious) which have much weaker type-checking that some dynamic languages like Python.
Maybe one could argue that Python has a robust but hard-to-use interface system called "documentation".
EDIT: For example, implementing a merge sort (without using any .sort() functions) would be interesting to compare between the three languages. Though to be honest, I wouldn't expect major differences aside from basic syntax.
Not really, you should be comparing the idiomatic ways in each language.
And if Python is more of a "batteries included / easily installable" language, then that's how you should use it.
>EDIT: For example, implementing a merge sort (without using any .sort() functions) would be interesting to compare between the three languages.
Only if you want to see how each language feels in terms of primitives and syntax and so on. Not if you want to see how you'd actually use the language, and what facilities and ecosystem you can leverage, which is more important.
If I'm going to compare Python for scientific computing to Java, for example, of course I'll consider that Python has Nympy, Pandas, Anaconda, and so on, and similar for what Java offers, not just try to e.g. write my own math code in pure Python and Java. Same for most domains...
> And if Python is more of a "batteries included / easily installable" language, then that's how you should use it.
This isn't python specific. "Idiomatic" Rust would also be to import the crate that does image comparison.
It just happens to be that for this particular problem a python library was found, but no rust crate. This isn't all that common.
But in this case, it's comparing two languages generally to each other. And the example chosen biases Python in terms of brevity and complexity because Python has a library for the use case chosen for the comparison. If you want to do a general comparison you need to be general in your approach.
Which would also be true for many other examples they might have picked, which makes it a reflection of the ecosystem, not a distortion.
It's a "subjective, primarily developer-ergonomics based" comparison. Seems like fair game to me.
That is useful when considering languages for a project, but not if you want to compare how good the languages are.
Like any comparisons there are a lot of different ways to do thing (even python...), so they'll always be "we didn't they use this or that". And frankly there are very few of these comparisons written by experts in all the languages (usually its someone writing up there experience).
For example's one might complain, he's comparing speeds using "Pillow" which seems to be a c -library wrapped for python, but honestly that's how one might do it, ands its pretty easy, so its legit to me.
If you truly don't care about handling the different variants and want them all to execute the same code you could have just as easily done:
let w = image1.width();
let h = image1.height();
let mut diff = image::DynamicImage::new_rgb8(w, h);
If you don't care about image1.color() 's variants, why even have the match at all? I guess I don't understand why it 'rubbed you the wrong way'. You don't have to match anything if you don't want to, and if you do want to match variants, you have all the tools necessary to reduce duplication & only handle the variants you want to. use image::ColorType::*;
match ... {
RGB(_) => DynamicImage::new_rgb8(w, h),
RGBA(_) => DynamicImage::new_rgb8a(w, h),
}
If you don't necessarily want the variants in the broader scope, it keeps them out of it, but within the local context of the match, the reader often will know what's being referred to. (And if they don't, the use statement will tell them.)Also, the author already has DynamicImage available in the scope (there's a use for it at the top) so the image:: prefix isn't needed in that section.
There's also what looks like a manual expansion of try!() in run(), that could just be create_diff_image(...)? which would be less verbose.
Besides that, cpython is not the only Python runtime or interpreter. The author should have tried pypy runtime which has mature JIT.
Also, this blew my mind: https://lobste.rs/s/xz5l8t/one_program_written_python_go_rus...
Just odd to read noun-verb combinations like:
> "If you’re comfortable with Python, you can go through the Tour of Go in a day or two..."
> "I would go as far as saying that Go’s strength is that it’s not clever."
> "I decided to give an honest go at learning Rust."
> "Go propagates errors by returning tuples: value, error from functions wherever something may go wrong."
I actually appreciate this less for other developers, but even for my own code. Let’s be honest, It’s difficult to come back to a code base a year or two after it went into maintenance even if you were the writer.
(declaim (ftype (function (string string) double-float) img-diff))
(defun img-diff (first-file-name second-file-name)
(declare (optimize (speed 3) (safety 0) (debug 0) (space 0)))
(let ((im1 (png:decode-file first-file-name))
(im2 (png:decode-file second-file-name)))
(declare ((SIMPLE-ARRAY (UNSIGNED-BYTE 8) (* * *)) im1 im2))
(/ (loop
for i fixnum below (array-total-size im1)
summing (abs (- (row-major-aref im1 i)
(row-major-aref im2 i))) fixnum)
(* (ash 1 (png:image-bit-depth im1)) (array-total-size im1))
1.0))) ; Convert rational to float
(img-diff "file1.png" "file2.png")What does declaim mean? What about ftype? What is the (* * *) in the array declaration? What is ash? I've read through some of CLtL and the rust book both, but none of those are constructs I've come across.
Also (this doesn't matter in practice due to rainbow parens for almost every editor), it's really hard to read lisp code without syntax highlighting. Rust isn't super easy, but I don't have to count braces to see what lines up.
Also a genuine dollars and cents factor: Development Time.
Do not just measure for cpu cycles and bytes. Development time is money, and that is a measurable factor too that must be considered as a language tradeoff.