Announcing Rust 1.27
blog.rust-lang.org
blog.rust-lang.org
SIMD is another notch towards never needing to use nightly Rust. What this really looks like in practice is that less frequently will a user encounter a fast-mode and a slow-mode for a crate depending on whether they’re using nightly, with all the unstable features enabled, or not.
I get your point that there are energy efficiency tradeoffs to consider. However for massively data parallel tasks like image operations, 3d graphics, machine learning, vision etc, I haven't seen SIMD implementation come close to GPU compute in terms of efficiency. Perhaps short lived text pressing and file compression where data access is more random could benefit from SIMD.
Your GPU is actually running much less than you'd expect. Unless pixels are changing there's a 95% chance that everything on the system is sleeping and just the display controller is being kept spun up. Battery requirements mean that DVCS on any mobile chipset is going to be really aggressive.
Either way GPU compute has a pretty large overhead both in terms of latency and scheduling. Earlier GPUs didn't schedule nice(hello waiting 16.6ms for your next compute request) and generally the type of places where you use SIMD(3D transforms, audio processing, etc) are so tightly coupled with other CPU operations that even the act of moving them to the SIMD registers is something you need to consider before diving into it. A lot of times waiting for some work queue to complete(or adding pipeline latency by waiting for next frame) just isn't feasible.
True regarding Android's APIs being awful, but taken literally, your statement implies that even Core Animation on-GPU compositing from 2007 is bad, which I'm sure you didn't mean. :)
Like I mentioned above, earlier GPUs are pretty poor at fine grained scheduling + latency and NEON has pretty ubiquitous support on the target devices we were shipping.
Also notice that the execution model of GPUs and CPUs is quite different. You need a far larger "breadth" of execution to efficiently use a GPU, compared to a CPU.
More generally SIMD is useful when you are repetitively performing the same instruction on a long stream of data but you don't want to send it over to the GPU because you don't want to incur the many order of magnitude slowdown of sending stuff back and forth to another chip or you can't rely on a GPU being there (e.g. embedded), or it's just overkill.
Also, most servers have only very basic GPUs (if at all). Unless you're on a dedicated GPU server for machine learning etc., you're going to work with the CPU only.
Everything from counting the length of a string to rasterizing the text you're reading right now.
But seriously, if SIMD would be useful there, why doesn't it use it?
It also makes for much easier googling. ;)
Would someone please explain the problem (and the solution) for someone who doesn't know Rust yet?
Box<Foo>
If Foo is a struct, this is a single pointer to the heap. If Foo is a trait, this is a double pointer: a pointer to the data on the heap, and a pointer to a vtable for its methods. These two things are very different, but look similar. That's the mistake.This change separates the two. Now:
Box<dyn Foo>
is always the latter, a "trait object". We have to keep the old syntax working for stability reasons, but can lint against it, so that it nudges you into the better syntax. So in this world, // always a single pointer to a struct (or enum)
Box<Foo>
// always a double pointer
Box<dyn Foo>
This is more clear for humans. The compiler can tell with 100% fidelity, of course. But humans aren't compilers. Different things with different costs look different.As an addendum, the parent mentions another feature, "impl Trait", which can be used with traits. Their point was, originally, Box<Trait> was the only thing that existed like this. With "impl Trait", it was "impl Trait" vs "Trait", and is now "impl Trait" vs "dyn Trait", which is much more clear, at least in many people's opinions. :)
Will Rust 2018 enable you to remove the old syntax? Or is that a bigger change than the opt-in is allowed to cover?
That being said, this particular fix is 100% automatable, and rustfix can already do it, so the time to update is “a few seconds”.
#![deny(bare_trait_objects)]
To the top of your crate and cause a compiler warning if someone uses the old syntax on your project.EDIT: just realized I mis-parsed "you" in parent. Oh well.
Edit: I didn't see the answer from steveklabnik below. Apparently, it's not clearly decided yet.
If you want to express something like "this is a variable that holds something that fulfils this trait", without knowing the _actual_ type it is, that variable effectively has an unknown runtime size. std::io::Read is an interface for reading bytes of some source, like a file or a socket.
This matters because we're talking about a stack frame. So the size needs to be known at compile time.
let a: u64 = 42; // ok, because well known size.
let b: Read = ...; // illegal, because unknown size.
A "trait object" places the object on the heap and has a pointer in its place. let b: Box<Read> = ... // legal, because pointer is a known size
However it's a bit more complicated, because this syntax allows for dynamic dispatch at runtime using a vtable. So there's a quite big difference between Box<u64> (a 64 bit unsigned integer on the heap) vs Box<Read> (a runtime dispatched lookup via a vtable).This difference is not obvious at a glance though. Hence the new syntax: Box<dyn Read>.
(I think I got that right)
Which “data” is the vtable stored with in C++? An object contains only a pointer to its vtable, not the vtable itself...
In other words, in Rust, a pointer to a shape is
(pointer to vtable, pointer to data)
whereas, in C++, a pointer to a shape is (pointer to shape)
where shape is (pointer to vtable, data)
If that's incorrect, I'm quite happy to be corrected! This isn't an area I'm an expert in.I get it now, that you meant the “pointer to the vtable” under “vtable” in your original post.
struct ClassVTable {
void (*firstMethod)();
void (*secondMethod)();
}
struct Class {
struct ClassVTable *virt;
int firstMember;
int secondMember;
};
and a pointer to it looks like: struct Class *objectRef;
A language that uses fat pointers, like rust, has this separated completely: struct Class {
int firstMember;
int secondMember;
};
struct FatPointerToClass {
struct ClassVTable *virt;
struct Class *data;
}
FatPointerToClass objectRef; // note lack of pointer, it'd be stored on the stack directly for eg.
This means a much simpler object layout in exchange for passing around larger pointers, and it's a good fit for trait-based typing since it doesn't require you to know all possible interface subtypes of the object in order to describe or use its layout.(please note that I'm using C struct to be precise about the concept, and this should not be taken as a perfectly verbatim description of the actual object layouts in memory)
Essentially Box<Foo> is static dispatch. You know at compile time which exact methods you are going to call. Box<dyn Foo> is dynamic dispatch. Since Foo is a trait, at compile time you won't know which methods are called, it depends on the type of the object passed in (as long is it implements the Foo trait).
Illustrating the need to dump the old unqualified syntax.
Is this right:
If Foo is a struct then Box<Foo> is static dispatch. If Foo is a trait Box<dyn Foo> is dynamic dispatch, but then so is Box<Foo> - but that is the regrettable point. So from now on, if Foo is a trait we should be writing Box<dyn Foo>, which is largely syntactic sugar to make the situation more clear?
I don't fully understand what the release notes are saying about using impl Trait. When would you use Box<impl Foo> vs Box<dyn Foo>?
It’s not that you would use Box<impl Foo>, but when taking about the two features, it’s easier to compare them, as they’re actually distinct syntactically.
Duck typing is implicit interface.
Foo().bar() is an implicit Foo.bar(Bar()) as self is only explicit on declaration, not call.
__init__ is called and passed parameters implicitly.
__new__ is an implicit class method.
Some time the magic is nice to have. Sometime it's like yield.
Yield implicitly turns your function into a very different object, which makes everybody goes wtf the few first times.
They fixed that with async / await, but I still wish we had a "gen def" to allow yield in bodies, like we have "async def".
$ rg-with-simd --version
ripgrep 0.8.1 (rev 223d7d9846)
+SIMD +AVX
$ rg-without-simd --version
ripgrep 0.8.1
-SIMD -AVX
$ time cat OpenSubtitles2016.raw.en > /dev/null
real 0m1.280s
user 0m0.020s
sys 0m1.257s
$ time wc -l OpenSubtitles2016.raw.en
336602465 OpenSubtitles2016.raw.en
real 0m4.303s
user 0m3.132s
sys 0m1.167s
$ time rg-with-simd -c 'Sherlock Holmes|John Watson|Professor Moriarty' OpenSubtitles2016.raw.en
6033
real 0m2.099s
user 0m1.750s
sys 0m0.347s
$ time rg-without-simd -c 'Sherlock Holmes|John Watson|Professor Moriarty' OpenSubtitles2016.raw.en
6033
real 0m4.128s
user 0m3.781s
sys 0m0.343s
$ time rg-with-simd -c 'Sherlock Holmes|John Watson|Irene Adler|Inspector Lestrade|Professor Moriarty' OpenSubtitles2016.raw.en
6731
real 0m1.989s
user 0m1.621s
sys 0m0.366s
$ time rg-without-simd -c 'Sherlock Holmes|John Watson|Irene Adler|Inspector Lestrade|Professor Moriarty' OpenSubtitles2016.raw.en
6731
real 0m18.417s
user 0m18.000s
sys 0m0.403s
Looks like `cat` is still faster, so there's some room for improvement. ;-) With a single pattern, we're almost there: $ time rg -c 'Sherlock Holmes' OpenSubtitles2016.raw.en
5107
real 0m1.333s
user 0m0.974s
sys 0m0.357s
This one is mostly thanks to glibc's memchr implementation (which uses SIMD of course), and the regex crate's frequency based searcher.Of course, I'm presenting best cases here. Plenty of inputs can make ripgrep run quite a bit more slowly than what's shown here!
The crazy thing is that we're still only barely scratching the surface. Check out Intel's Hyperscan project for some truly next level SIMD use in regex searching!
BTW as a happy daily user of rg thanks for all the work you put into it, definitely shows.
Furthermore, there are worries that depending on the llvm behavior might impede work on alternative rust backends, such as cretonne.
Stabilizing inline asm in its current form would be a mistake. It's simply not ready.
[0]: https://github.com/rust-lang/rust/issues?q=is%3Aopen+is%3Ais...
Do you see a future where wasm has its own std::arch module?