Aliasing rules in C make some optimizations not possible in C. Good luck putting that `restrict` everywhere to compensate. Generics with static dispatch makes it possible to get a code that is way better than what a sane person can write with C's preprocessor. Practical developer will very often sacrifice performance for sanity in C codebase. Memory safety makes it easy to write correct, reliable multithreaded programs in Rust and actual package manager makes it practical to reuse well-written and optimized code.
Even if there are some microbenchmarks where C gives better results, it's because Rust version doesn't "cheat" by using `unsafe` constructs. Essentially in Rust one can write functionally exactly same code that C uses, including inline asm etc. which is useful for performance-critical libraries that wrap it in safe abstractions, but would be frowned upon in a tiny synthetic benchmark program that is supposed to benchmark idiomatic code.
Rust type system prevents same for Rust. Those __m128 values are just raw bytes, different SIMD instructions view them as different types, and Rust's type system ain't OK with that.
To write efficient SSE code, you need to understand __m128i is a just a hardware register, and use whatever representation you need for particular instruction.
Example: http://stackoverflow.com/a/17355341/126995 The SSE2 code is impossible to translate to Rust, because _mm_sub_epi8 views the registers as std::simd::i8x16, while the subsequent _mm_srli_epi64 views the same registers as std::simd::i64x2
> The SSE2 code is impossible to translate to Rust, because _mm_sub_epi8 views the registers as std::simd::i8x16, while the subsequent _mm_srli_epi64 views the same registers as std::simd::i64x2
That's false. It might be impossible to implement that code safely using the current simd crate as is (which is not in the standard library), but the simd crate is an abstraction over the raw intrinsics and isn't blessed by the language. You can reach in and use the raw intrinsics (or unsafely cast between SIMD types). If you look more closely at the SIMD algorithm in the regex crate I wrote, you can see that I'm already doing this (to convert from an `&[u8]` to an `u8x16` without bounds checks).
Popping up a level, I can understand why this is not obvious from an outside observer. The SIMD story on Rust is very much still evolving, and we don't have it all completely worked out yet. When we do, I predict it will be awesome. :-) I have a whole bunch of ideas on how to use it to improve text search even more!
I'm curious if there is a way to limit the entries in the Benchmark Game to just those that actually use only the language standards, where those exist...
SIMD will get you into platform specific woes unfortunately, but that shouldn't stop us. :-)
http://www.iso-9899.info/n1570.html
Intrinsics are not library functions - they would hardly provide the desired performance benefit if they were.
Intrinsics are "magic" functions that the compiler has built-in support for, and thus I would call them language extensions.
For example from GCC header avxintrin.h:
extern __inline __m256d __attribute__((__gnu_inline__, __always_inline__, __artificial__))
_mm256_permute2f128_pd (__m256d __X, __m256d __Y, const int __C)
{
return (__m256d) __builtin_ia32_vperm2f128_pd256 ((__v4df)__X, (__v4df)__Y, __C);
}
Over here the word "__builtin_ia32_vperm2f128_pd256" occurs in the "cc1" and "cc1plus" binaries and "avxintrin.h", but nowhere else. Notably there is no externally visible definition.In Visual C++, the prototype of that intrinsic is valid C99 code:
extern __m256d __cdecl _mm256_permute2f128_pd(__m256d, __m256d, int);The likely answer is "none", as these are built-in to Visual C++ in a similar way to GCC, with only superficial differences such as an apparent lack of indirection to some __builtin_foo.
https://msdn.microsoft.com/en-us/library/26td21ds(v=vs.100)....
"An intrinsic is a function known by the compiler that directly maps to a sequence of one or more assembly language instructions. Intrinsic functions are inherently more efficient than called functions because no calling linkage is required."
https://msdn.microsoft.com/en-us/library/y0dh78ez(v=vs.100)....
"The use of intrinsics affects the portability of code, because intrinsics that are available in Visual C++ might not be available if the code is compiled with other compilers and some intrinsics that might be available for some target architectures are not available for all architectures."
C99 language spec says nothing about the things you’re talking about. It only specifies the language, not how C compilers implement various features.
From the programmer’s point of view (except when the programmer works on the compiler itself), intrinsics are just C library functions, only inline, and very performant if done right. Fully compliant with C99 and C++ language specs.
(Rust is currently beating C on one benchmark, within a few tenths of a second on a few more, and then way slower on the SIMD-reliant ones)
Btw, the last time this topic was discussed we measured the performance using a microbenchmark. Here is the C version: https://gist.github.com/bjourne/4599a387d24c80906475b26b8ac9... None of the Rust aficionados were able to produce a Rust program (using the nightly build of Rust even) that came close.
I actually don't think microbenchmarks are a good way to think about performance anyway. There's two reasons:
First, Rust's safety guarantees let you get away with more dangerous things. Consider scoped threads, or non-atomic reference counting.[1] You _could_ write this in C, and you'd do so for a microbenchmark, but not for a real codebase, as it's far too dangerous. Or, you might do it, but end up with bugs that you don't detect, that aren't there in the Rust version.
1: http://blog.faraday.io/saved-by-the-compiler-parallelizing-a...
Secondly, microbenchmarks are often "can the best person at language X write something faster than the best person at language Y"? I don't think that's nearly as interesting as "When an average programmer of languages X and Y write a program, which is faster?" Rust's default patterns and style is already extremely fast. It's not a guarantee, mind you, but for real-world usage, it's the average case that matters more.
The reason we have micro benchmarks is because it is so hard to reason about the average case. Yes, code written in Rust could on average be faster than C due to its safety features. But so could Python, Java or Haskell so we're back were we started and can't say anything about relative language performance.
It's a small niche, yes, but some people absolutely need to squeeze out every gram of performance, and they should continue to prefer C over Rust.
That was last year, before rustc 1.0.0-alpha
> There's been a lot of variance on the benchmark game lately
With Swift, because programs written for Swift 2.0 failed with Swift 3.0
Rust's promise is if it compiles it's probably good code, as that compiler is exceptionally picky. Most common C++ errors that cause "undefined behaviour" are not possible due to the strict checking.
I guess you missed this in the sibling, but ripgrep does this today, and it works on Windows, Mac and Linux, and should be as fast or faster than GNU grep. See benchmarks: http://blog.burntsushi.net/ripgrep/#single-file-benchmarks See binaries: https://github.com/BurntSushi/ripgrep/releases