I'm not sure. The following claim is being made on Nim's website:
> Modern concepts like zero-overhead iterators and compile-time evaluation of user-defined functions, in combination with the preference of value-based datatypes allocated on the stack, lead to extremely performant code.
> Efficient, expressive, elegant
> Nim is a statically typed compiled systems programming language
A systems programming language need to have exceptionally good performance, so not sure it's to be expected that the performance is so different than Rust.
Apparently the Rust versions have received more optimization efforts compared to the Nim code as well, and it seems some people are working on optimizing the Nim versions so we'll see where it lands in the future I guess: https://news.ycombinator.com/item?id=30244188
Having programmed in both I think Rust often makes it easier to write fast code by making all the copying vs borrowing more explicit.
Nim is obviously still very fast and much easier to pick up, so there are trade-offs here.
As an aside, I'm not expecting a lot of languages to live up to the "C like performance" claim when they make it, as C has been receiving optimizations for over 50 years by now, so it makes sense it's still the fastest around, at least until there is some major breakthrough in programming language design.
EDIT: Downvoters - is that not the case?
C isn't inherently fast, commonly used constructs in C just tend to match up well with semantics that can be turned into high-performance machine code. For a language that compiles to C to be fast, its commonly used constructs would need to have semantics that match up well with what can be turned into high-performance C, which is often not the case. There are many examples of languages that compile to C, but are still very slow.
Maybe a better way to frame it is that most other languages are inherently slow. Garabge collection will always be slower than non-GC. JIT and interpreted languages will always be slower than ahead-of-time compiled languages. etc etc.
C has a runtime, it's called the CRT (C Runtime).
http://www.vishalchovatiya.com/crt-run-time-before-starting-...
It's quite small, but it's there.
I am well aware of the CRT. I don't think its worth mentioning because it is a "run time" in name only. It literally doesn't do anything at run time. The only thing crt0 does is call your main(). From that point on, only your actual code is executing.
Contrast with Java or Python, where your code is executing inside a runtime which eats up cpu cycles at runtime.
Look, I get what you're saying, but it's not exactly relevant here. Yes, after your C program is loaded, execution follows by just moving the PC to the next instruction in the loaded binary. In C, adding two integers can be compiled down to a concise, machine supported operation; whereas in a language like Python, adding two numbers is a branch-heavy operation that cannot be compiled to a concise instruction ahead of time, so it must be interpreted at runtime. That's true. But that's really more about C being a compiled language than whether or not it has a runtime.
For an example of a runtime-heavy language that can compile highly performant code, see Julia. Thanks to dynamic dispatch, you get the best of both worlds. Yes you pay a heavy upfront cost (TTFP), but it's the same deal with compiling code -- pay once upfront, benefit forever with improved runtime performance.
You used Java as an example of a runtime-heavy language, but I think Java demonstrates why C should be considered as having a runtime, regardless of how small it is. Imagine a hardware Java Processing Unit (JPU), which is a CPU that has Java byte code for machine code [1]. I don't see why it wouldn't be possible to handle all of the stuff the JRE does in the OS, and everything the JVM does in the JPU. Then we could say that Java has no runtime. But really all we've done is taken the runtime and moved it to another place.
You could also imagine a version of C which compiles to CBC (C byte code) and runs on a CVM (C virtual machine), which itself runs on the JPU. You could write C code as usual, but the resulting binary would be compiled to CBC, and would have to run through a number of layers before it is executed on the JPU. Then could we say Java is an inherently fast language, and C is necessarily slower because it has a runtime? Maybe in this strange reality, but it's not necessarily the case. And it certainly has nothing to do with the languages per se, but their implementations on specific hardware/software platforms.
Or to get even crazier with it, we could imagine a version of Java which compiles directly to x86 machine code, because why not? There's no reason that should be slower than C.
Anyway, I think the point is that whether or not one considers a language to have a runtime depends very much on how a language is compiled, what hardware it's targeting, and how much you're able to leverage the OS (there's usually a lot of leverage available when the OS itself is written in the language you're using). So really, all languages have a runtime, it's just a matter of where it's hiding.
Your definitions are meaningless. You may as well also say that all languages are garbage collected, since the C programmer must call `free` one way or another (or not, and let the OS be the garbage collector at cleanup).
By everyone's understanding of the words, C does not have a runtime nor garbage collection. Java does have a runtime and garbage collection. You can play with definitions and hypotheticals; it does not change the facts.
What? But we both just agreed it did.
The way I understand these words are that C is a language, and a language is different from its implementation, and implementation is dependent on hardware. I understand that a runtime is code that runs when a program starts, in order to enable its execution. Is this not your understanding?
Let me put it this way: if C is inherently fast as you, why can you take the same C source, put it into two different compilers, and get two different binaries with different performance profiles?
The answer is because the language itself has no bearing on speed, and it all comes down to how a language is compiled. So I still wouldn't say C is inherently fast, but I would say that C's design lends itself to compiling fast binaries for Von Neumann style machines. That of course doesn't preclude compiling just-as-fast or even faster binaries from other languages, e.g. Rust or C++.
> You can play with definitions and hypotheticals; it does not change the facts.
I'm not playing with hypotheticals. JPUs are a thing; I linked to the Wiki page. Natively compiled Java is basically GraalVM: https://www.graalvm.org, which I think pretty much proves the point -- it all comes down to the compiler. Garbage checking is also an implementation detail. You can even run a garbage collector for C if you want.
Moreover, it is also untrue on the other side of this axis - you can often go much faster. Even with -Ofast or -Omax or whatever there are still (often) opportunities for SIMD vectorization on modern machines that autovec cannot find. Most C compilers provide (not really standardized) escape hatches to assembly in various ways for that last 2..64X speed-up (single thread). (Yes, avx512/8-bits=64 wide vectors these days.)
Finally, whether it "has a runtime" in the sense of automatic memory management, well, that just depends upon how you link programs. Link with -lboehm and you can never free (and, yes, allocations become potentially expensive). These Rust conversations are unfortunately prone to absolutist statements that indicate very narrow exposure.
If your language has semantics that are difficult to translate into fast C, then your language isn't going to be fast just by going through C.
Even if you do use the GC, the compiler elides GC work if the type doesn't escape the scope.
But no, Fortran is not an alternative to C (und vice versa).
Languages like Nim/Rust/D/etc. should not have significantly different speeds for equivalent implementations for these kinds of benchmarks, as they ultimately can be written to compile to equivalent LLVM IR and thus the same machine code.
What you're usually seeing when you see such large differences is that there are differences in the implementation of the algorithm or of a standard library function.
There's nothing about these kinds of languages that should result in inherently significant performance differences on such synthetic microbenchmarks.
Allocation-heavy or write-barrier-heavy code might make a difference, but you literally have performance differences while iterating over arrays of scalars. And even then, last I benchmarked it, Nim actually compared favorably with e.g. jemalloc.
And in the case of Swift the atomic "release" calls at ends of scopes will run regardless whether there are hot loops or not.
So it all depends on the language.
Also, you pointed to 3 languages with 3 different memory management strategies and 3 different toolchains (Rust is LLVM-based; Nim compiles to C and then calls a C compiler, GCC recommended; D has its own non-GCC non-LLVM compiler). It feels weird to put them in the same bag implementation-wise.
[1]: <https://golangbyexample.com/goroutines-golang/> (potentially stale information)
> And in the case of Swift the atomic "release" calls at ends of scopes will run regardless whether there are hot loops or not.
Yes, but Nim does neither. Nim does insert write barriers, but that's only if you (1) write pointers to heap-allocated memory and (2) don't use the mark-and-sweep GC, so that can't explain big performance differences for those benchmarks where you primarily iterate over arrays.
> Also, you pointed to 3 languages with 3 different memory management strategies and 3 different toolchains (Rust is LLVM-based; Nim compiles to C and then calls a C compiler, GCC recommended; D has its own non-GCC non-LLVM compiler). It feels weird to put them in the same bag implementation-wise.
You don't seem to be aware of it, but D actually has an LLVM-based compiler (LDC). Nim can also utilize LLVM via choosing clang as its backend. You can definitely use them to generate equivalent LLVM IR and compare that. There's nothing weird here, this allows us to rule out differences related to the backend.
I guess it depends on what significant is but I certainly hope Rust can be faster than C in many scenarios due to optimizations llvm can make with the added type information such as inferring aliasing rules.
- multithreading runtime (i.e Rayon vs Weave https://github.com/mratsim/weave)
- Cryptography: https://hackmd.io/@gnark/eccbench#Pairing
- Scientific computing / matrix multiplication: https://github.com/bluss/matrixmultiply/issues/34#issuecomme...
There is no inherent reason why a Nim program would be slower than Rust.
https://programming-language-benchmarks.vercel.app/rust-vs-c
C has nothing to do with performance, runtimes and libraries do. Nim has runtime overhead whereas C and Rust do not.
Not expecting it to match Rust or C in performance, just surprised at the size of the gap.
Would be interesting to see what the generated C is like from the Nim code.
Just did a few quick tests. A hello world in nim is a one liner:
echo "hello world"
It compiles to a 102136 byte binary on Linux. A hello world c program consisting of a single printf is 16696 bytes.If you don't use automatically managed types like seq/string and ref, there is no GC underneath. And even then, the GC is just reference counting so no difference from Rust Rc/Arc.
The issue with targeting C is you ended up losing a lot of information about intent. Compilers looking at C have to put a lot of trust that what the programmer wrote was intentional (and thus, unoptimizable).
Rust and C both end up faster than Nim because they are talking directly to the compiler.
IDK enough about Nim to say if it could hit Rust/C speeds if it targeted the LLVM instead, but it would be a lot faster than it currently is targeting C.
> The issue with targeting C is you ended up losing a lot of information about intent. Compilers looking at C have to put a lot of trust that what the programmer wrote was intentional (and thus, unoptimizable).
The same issue exists in C-sourced compiler backends, which is both GCC and LLVM. It’s not exactly a secret that anything which is not strongly exercised by standard C(++) codebases tends to be rather broken for any non-trivial use and takes years to fix (restrict/noalias being the poster child for this, at least in the context of Rust, I think it hasn’t gotten re-disabled yet so it might make a year).
The JVM, for example, doesn't see the same issue so much with non-Java languages. Mainly because (IMO) it's been hosting non-java languages for several years now.