Announcing Rust 1.19
blog.rust-lang.org
blog.rust-lang.org
Just started using Rust in a serious capacity this month to secure some C++ functions that are called by our Erlang apps, with great assistance from Rustler [1]. Several people have complained to me about the decision to remove async IO from Rust, but I'm really grateful that it happened, because it lets Rust focus on being the best at what it is. Erlang's concurrency primitives and Rust's performance & security are a match made in heaven.
> the decision to remove async IO from Rust
I'd be interested in hearing more about this, did they maybe mean green threads? https://tokio.rs/ is the big async-io project, it's certainly not removed in any sense!
Metal IO is an awesome library as well:
The release notes mention my RFC, but a huge thanks also to Vadim Petrochenkov for the implementation, and all the myriad RFC contributors.
struct foo {
int num_elements;
char data[0];
}
struct foo *d = malloc(sizeof(foo) + num);
d->num_elements = num;
/* Now d->data is effectively an array char[num_elements] */The way custom DSTs work in Rust is super annoying at the moment, but slightly less annoying for generics.
You are free to define a custom DST that is like:
struct Foo {
header: u32,
flexible: [SomeType]
}
However, there is no way to construct this type. The most you can do is calculate field offsets and write a ton of unsafe code.If the type was instead
struct Foo<T: ?Sized> {
header: u32,
flexible: T
}
you would be able to construct a `&Foo<[SomeType]>` from a `&Foo<[SomeType; N]>` via DST coercions.This only works if you know the size of the array at compile time.
For a runtime sized thing you have to implement a bespoke vector-like thing. You can find an example of that code in https://github.com/servo/servo/blob/e19fefcb474ea6593a684a1c... where we have a generic "HeaderWithSlice<H, [T]>" type which can be heap allocated as a header followed by the flexible DST [T].
This could be improved. The HeaderWithSlice thing could probably be a useful crate for implementing completely-heap-allocated vectors (where the len/cap are on the heap) or shared reference counted types with flexible members. We haven't really split it out as a crate because we don't actually ever mutate it so the amount of code we need is significantly less (but it's not as useful as it could be).
So yeah, flexible array members can be implemented in Rust, but it's a lot of hacky work. It's an equivalent amount of work and unsafety to write a custom Index impl on a repr(C) type with a zero-length array at the end. The DST doesn't actually get you much here.
For example-- what happens if I create an array of foos and then malloc data of each element to some arbitrary size?
In terms of things you can express now but that could still use some work: anonymous structs and unions. You currently still have to name all the intermediate aggregates, even when the C API doesn't. See https://internals.rust-lang.org/t/pre-rfc-anonymous-struct-a... .
A kind of similar issue appears even with enums, because it is hard to guess the size of the integer that a given enum is implemented with. This is why it some not to use C enums in public interfaces to libraries.
True; fortunately, Rust has a means of declaring the size of an enum, with #[repr(u8)] and similar. But you do have to figure out the size of the C enum.
But yes, I can imagine that for Mozilla there is a need to mix Rust and C++ right inside an application at boundaries that are not usually considered external, in which case all kinds of non-portable C constructs are acceptable.
myregister.tx_baudrate_bitfield = 1; // Writes the whole myregister
myregister.tx_now_bitfield = 1; // Reads the whole myregister, then bitmasks and finally writes.
The same mistake can of course be done with manual masking also but then it's more obvious that you are reading the registers also. When the code looks like above it's very easy to not realize that it also requires a read.I read that they're essentially the same as C unions (~untagged enums), but I have no clue how `match` works in that case.
Some consider it a counterpattern or bad memory to be able to not fill out all parameters or reference them by name. However some interfaces easily require like 50 different function parameters that cannot be removed in a simple way and they all make sense in different configurations. Without default values and named parameters you're lost there. I don't get this design decision on Rust side at all.
[1] https://doc.rust-lang.org/std/default/trait.Default.html
https://matplotlib.org/api/pyplot_api.html#matplotlib.pyplot...
I count about 40 plus more which are not listed.
You surely could first make a dictionary or structure and fill it with these arguments. But then "kwargs" are just that, a dictionary.
I find the interface of the plot function pretty straightforward. There are a lot of options but it's pretty clear what they do and when to use them. Most often you use only some of them - but often different combinations. Splitting this up into multiple helper objects that need to be constructed and filled beforehand would turn the default one-liner into a ten-liner which is not better.
Maybe Ruby on Rails is an exception, but while I used to be a fan of default/keyword arguments, especially in combination, seeing how they were used there made me very much not like them any more. It's impossible to tell what's going on.
It's very clear there...
Similarly, in Swift, the order of defaulted arguments is significant. So you wouldn't see anyone do something like that crazy matplotlib method in Swift, because remembering the order of all the arguments in that function is impractical.
#[derive(Default)]
struct Config {
interesting_thing: bool,
a: i32,
b: i32,
c: i32,
}
fn main() {
Config {
interesting_thing: true,
..Config::default()
};
}
In addition, the builder pattern is especially useful for larger configuration objects like this.1) More readable code. The classic is foo.Bar(true). It breaks the flow of reading to have to hover and see what that 'true' means. Much nicer to see foo.bar(launchMissiles: true).
2) Protects from a particular class of dumb mistakes. You have a function foo(x, y) where x and y have the same type. You refactor it so that one of the parameters isn't needed anymore. It's surprisingly easy (read: I've seen it, and I've done it), when you clean up the function calls, to accidentally delete x instead of y or vice-versa. Named parameters prevent that.
This is a case where I've started using Enums in Java. For example:
foo.bar(LaunchMissiles.YES);
In Rust, you could have a macro to make defining these types easier (and have it automatically generate to_bool and from_bool methods): named_bool!(LaunchMissiles);
foo.bar(LaunchMissiles::Yes);With my language-nerd hat on...
While they may seem basic as a user, designing a language means you need to think about edge-cases. Methods in Rust have some special rules around dispatch, and in order to implement one or both of these features, all of that stuff needs to be considered and designed.
In other words, a lot of work goes into new features, even ones that are easy to use.
In Rust's case, we haven't ruled out adding these features completely, but nobody has put in that work to come up with a proposal. Part of that is that while people tend to see the lack of these features as a mild annoyance, it's not enough to prioritize over other work. We'll see how it all shakes out.
> nobody has put in that work to come up with a proposal.
Features like this mainly need a champion who really cares about getting it into the language & can work with the language & compiler teams to complete that process.
I'd probably be in favor of a hard cap on # of function arguments. :)
(There's also another another option, which is to have a different function for each combination of parameters. This obviously doesn't scale in the large, but it's perfectly acceptable for functions that take only a single optional parameter, which IME is a plurality of APIs that want optional parameters).
There's already such a language, it's called Haskell and it only allows one parameter per function.
Please tell me you're not responsible for this monstrosity https://salilab.org/modeller/9.18/manual/node315.html
These interfaces are poorly designed; as in natural language, you almost never need more than a subject, direct object, and one or two indirect objects, any of which may themselves be compound entities.
A call with more than about four parameters has usually decomposed at least one thing that should be composite in the argument list.
They mention the case where the type can be distinguished by the lest significant bit, but wouldn't it be better to handle that case as an enum? That is, the least significant bits define the enum tag, while the remaining bits define the associated value.
(By the way, I really mean this as a straight question, not a criticism in the form of a rhetorical question. I really don't know enough about it to be criticizing it.)
TL;DR: FFI with C is much harder without unions. There are smaller reasons as well.
Besides, enums aren't always optimal for some things. Let's say that you have an array of 32-bit values whose prime-numbered indexes contain either a signed or unsigned 32-bit integer, depending on a single global boolean tag. You can't encode that invariant into a (non-dependent) type system. If you want to do it safely, using enums, you are going to have a tag for every value, and that's going to cost you. (A bit crazy example, but bear with me.)
Unions are a way to circumvent the type system a bit. They allow you to keep track of the stuff inside memory slots using the way you see the best. But there's a reason they're unsafe - the responsibility is on you!
Rust doesn't let the programmer specify the layout of enums in detail right now, so you can't specify where the compiler should place the discriminator.
The example I care about is a word which is a pointer if the bottom bit(s) are 0, but otherwise contains a bunch of packed fields, where the least significant bits are a nonzero value I can do something useful with.
https://github.com/servo/rust-bindgen/blob/master/tests/expe...
Basically, if you have a union of A, B, C, you create a struct with three zero-sized fields using BindgenUnionField<A>, and then add a field after that containing enough bits to actually fill out the size. Because the BindgenUnionField is zero sized, a pointer to it is a pointer to the beginning of the struct, and it has an accessor that treats the pointer as the contained type.
This makes the API for field access `union.field.as_ref()` instead of `union.field`, but that's still pretty clean.
It's still a hack, and I'll be happy to see it go, but it's a really fun hack.
Is that a stable assumption?
This is zero cost stack allocation of data in a way which avoids destructors.
Basically, currently, in Rust, if you want to allocate a type and avoid destructors being run, you have to write a wrapper around `Option<YourType>` that nulls the option in its destructors. Or you heap allocate it and turn the box into a raw pointer after allocation. Both have overhead. The zero-overhead way of doing it is to stack allocate an array and cast pointers, but then you need to know the size beforehand.
With unions, you can stack allocate a `union Foo {x: YourType}` and just use that. Unions don't have destructors, so this stack allocates enough space for your type, and lets you unsafely but conveniently access it as your type (no ugly pointer casting hacks), but you can guarantee that destructors won't be run.
The obvious question here is -- why is this even necessary? Surely you can just call mem::forget to avoid destructors right before the function returns.
However, destructors also get run whilst panicking, so if your function accepts a callback, and that callback panics, you can't avoid destructors without the overhead mentioned before.
For a concrete example of this use case, check out ArcBorrow::with_arc() (https://doc.servo.org/servo_arc/struct.ArcBorrow.html). ArcBorrow<'a, T> is basically a borrowed reference to a T that is known to be backed by an Arc (atomic reference counted type). You can obtain borrowed references to an Arc normally, which is great -- lets you share the data without bumping atomic reference counts all the time and paying the atomic overhead. But if you have an &T -- a borrowed reference to a T -- there's no way to bump the reference count on that if you need to escape the borrow; since there's no guarantee that the &T borrows from an Arc allocation and not something else. So you must pass down an &Arc<T>, and that has double indirection. There are other reasons (pertaining to the existence of RawOffsetArc, which has to do with FFI constraints) as to why &Arc<T> won't work for us there, but I won't get into those. ArcBorrow<'a, T> lets us freely pass around borrows of &T which can be cloned as an Arc if necessary.
Anyway, ArcBorrow has a with_arc method (https://doc.servo.org/servo_arc/struct.ArcBorrow.html#method...). This takes a closure and passes an &Arc<T> to it.
But we don't have an Arc<T>, we have an ArcBorrow<T>, which has a different representation (in particular, ArcBorrow contains a pointer to the data, whereas Arc contains a pointer to the allocation, which starts earlier because of the refcount).
So we construct a fake Arc<T> on the stack (https://doc.servo.org/src/servo_arc/lib.rs.html#884-907), and share it with the closure. Because it's a fake Arc<T> we can't actually let its destructors be run, so we put it inside NoDrop, which on nightly uses unions (but on stable uses the non-zero-cost methods I mentioned above).
I use this same trick in array-init (https://github.com/Manishearth/array-init/blob/a0cb08928b42d...), where I stack allocate an uninitialized array and let you fill in the elements with a closure. If the closure panics the destructor of the _partially_ initialized array should not run, so again, it's in a NoDrop.
In general when writing unsafe abstractions you often need escape hatches like these.
Great effin post btw, i spent about 40 minutes readin that shit
> why doesn't ArcBorrow just store a reference to the Arc and also a reference to the underlying T
That's two words you're copying around on the stack.
Admittedly, that's a negligible cost. We don't care about that cost. I bet that cost never shows up in profiles. It would be premature to optimize for that cost :)
The real reason is within the "There are other reasons" I mentioned above ;)
These other reasons have to do with RawOffsetArc; it's a long story. The short version is that you may not always have an Arc<T> that you're creating an ArcBorrow from, it may be something else.
So basically this code is Servo's style system, and it is being used by Gecko (Firefox's browser engine). Gecko is in C++.
Servo's style system is quite parallel. So certain things are shared via Arc<T>. Pretty normal.
However, some of these things are shared with C++ code too! We've taught Gecko's refcounting setup about what an Arc is, and it does the appropriate FFI call when it needs to addref/decref it. This is all great. It works. These types are otherwise opaque to Gecko, and it does FFI to get to each.
However, we have one struct, ComputedValues, which stores all the "style structs" (where computed CSS styles go). It's basically a bunch of Arc<T>s of these style structs. ComputedValues is a Rust-side thing, and it's stored in a heap allocation dangling off a "style context" in the C++ code. It's otherwise opaque to C++.
The main operation Gecko does with ComputedValues is fetch a style struct. The style structs are C++ structs which both C++ and Rust can understand. So these getters are a bunch of FFI calls that take ComputedValues. This FFI call turns out to have an overhead that turns up in profiles, and there's an extra cache miss involved in hopping to the ComputedValues allocation (which also turns up in profiles). Both are major.
The fix is to store ComputedValues inline in the style contexts, and make it non-opaque so that C++ can actually read the types. Basically, C++ should see some regular pointers to the style structs. Rust will see Arc<T>.
But Arc<T> is a pointer to the allocation of the Arc. Arc is allocated with the refcount first, and the type T next. And the Rust struct layout isn't something C++ can understand, so code that assumes the offsets and does pointer arithmetic will be brittle and can change in a Rust upgrade. Thus arises RawOffsetArc<T> (https://doc.servo.org/servo_arc/struct.RawOffsetArc.html), which is represented as a pointer to the T, but it has the foreknowledge that T is arc-allocated and has a refcount preceding it. RawOffsetArc<T> is the same as an Arc<T> in all other aspects.
So these structs are now stored in a RawOffsetArc<T>, to make the pointers match with the C++ side representation.
However, there's also pure servo code that uses Arc<T> for this. So we can't just pass around &RawOffsetArc<T> because the servo code doesn't have that. It's not easy to migrate, nor do we really want to (Arc<T> has some more APIs and I don't want to add support for all that to RawOffsetArc). So it becomes easier to create ArcBorrow<T> as something that is guaranteed to have come from either a RawOffsetArc<T> or an Arc<T> (both are the same in behavior and heap representation, just that their stack pointer representation is offset. Converting between the two is a simple pointer bump on the stack). Because they're the same, ArcBorrow<T> can just be a pointer to the T, and the rest works out.
This is one of the reasons -- the other reason is that unlike Rust, where the refcounting is done by the wrapper (you can stick anything in Arc<T> and Arc<T> will handle the refcount), Gecko puts the burden of refcounting on the inner type. This means that if you use RefPtr<Foo> in Gecko, RefPtr will not create a refcount for you, Foo is expected to have AddRef()/Release() methods, which usually bump a refcount field it defines. Furthermore, it's taken as a given that if you have a `Foo`, it is heap allocated (and thus can be refcounted).
This means that having Foo instead of RefPtr<Foo>* is pretty common in Gecko. And it gets passed over FFI a lot to Servo, which again has to either construct transient Arcs, or treat it as an ArcBorrow. We currently do both, but I'm planning on migrating stuff to be more reliant on the ArcBorrow model since it leads to cleaner code.
(A lot of this complexity stems from the fact that browser engines are pretty tightly coupled codebases, and thus the "style system" doesn't have a clean API boundary. There's a lot of reaching into each others' guts that is necessary to make this work)
(block outer
(loop for i from 0 return 100) ; 100 returned from LOOP
(print "This will print")
200) ==> 200
(block outer
(loop for i from 0 do (return-from outer 100)) ; 100 returned from BLOCK
(print "This won't print")
200) ==> 100
[0] - http://www.gigamonkeys.com/book/loop-for-black-belts.htmlEDIT: I always do formatting wrong on here.
CL-USER> (loop for i from 1 below 10
sum i
when (= i 6)
do (loop-finish))
21 x = (1..100).each do |i|
if i > 9
break i
end
end
Evaluating that gives > x
=> 10
You can also supply a value to next within a loop, which is occasionally usefulhave code around my app doing things like:
loop do
code = random_string(6)
break code unless code_already_used?(code)
end