Four years with Rust
words.steveklabnik.com
words.steveklabnik.com
Not having exceptions tends to generate workarounds which are uglier than having exceptions. Rust seems to be digging itself out of that hole successfully. Early error handling required extremely verbose code. The "try!()" thing was a hack, because a expression-valued macro that sometimes does an invisible external return is kind of strange. Any macro can potentially return, which is troublesome. The "?" operator looks cleaner; it's known to affect control flow, and it's part of the language, so you know it does that. Once "?" is in, it's probably bad form for an expression-valued macro to do a return.
There's still too much that has to be done with unsafe code. But the unsafe situations are starting to form patterns. Two known unsafe patterns are backpointers and partially initialized arrays. (The latter comes up with collections that can grow.) Those are situations where there's an invariant, and the invariant is momentarily broken, then restored. There's no way to talk about that in the language. Maybe there should be. More study of what really needs to be unsafe is needed.
[1]: https://internals.rust-lang.org/t/all-the-rust-features/4322
(1.14 is tomorrow, so even if that was true, I'm off by a day.)
Could you say a bit more about what you mean by "now"? You write like it's a recent realisation but (as linked in the parent article) it was "discovered"/described/promoted more than 4 years ago: http://smallcultfollowing.com/babysteps/blog/2012/10/01/move... .
---
Are your other two paragraphs related to the article, or are they just your general observations about Rust?
The "implicitly copyable" problem is amusing. That was dealt with by Wirth in Modula 1 with the rule "if the programmer can't tell, it's up to the compiler". Thus, non-writable objects could be passed either by reference or by copy, depending on object size. This was up to the compiler. The usual rule was that anything up to 2 words in size was copied. Since the called function couldn't change the value, it didn't matter.
This avoids philosophical gyrations. You want a rule that says you can copy ints and floats, for performance reasons. Trying to reach that via type theory makes it harder. It's an optimization.
I think as long as those patterns can be abstracted out and moved into thoroughly-vetted libraries with a safe interface, unsafe code isn't really a problem. I expect that getting Rust's standard libraries to a place where regular applications very rarely need to create their own unsafe code blocks is going to be a major long-term effort.
(Of course, if there were some simple language feature that would make some common use case of unsafe code blocks unnecessary, by all means we should do that as well.)
(My current project I'm using to learn Rust is to port a ray-tracer I wrote in Haskell. I wouldn't expect functional-style code to require a lot of unsafe blocks and I haven't needed any yet, but who knows?)
Like you said all the other cases are nicely wrapped into a library but I'm not quite sure how you'd abstract the first one above. It's small enough to do the unsafe block inline that it's not annoying but definitely shows up(esp when you hit APIs that are copy of C interfaces).
Like I said, not really an issue, just a common pattern that I see.
Ironically, this is also a good example of how unsafe is complicated; "hey make an array (not vector) of copies of this thing" has lots of edge cases!
- https://github.com/andelf/rust-adivon/blob/master/src/deque....
- https://github.com/andelf/rust-adivon/blob/master/src/queue....
These are all the doubly-linked list problem:
struct Node<T> {
item: T,
next: Option<Box<Node<T>>>,
prev: Rawlink<Node<T>>
}
Since this is templated code, it might be possible to break it by instantiating it on a type with unusual semantics.- https://github.com/BurntSushi/aho-corasick/blob/master/src/f...
Looks like unsafe code for "performance reasons". But there are no comments near "unsafe", so it's hard to tell.
- https://github.com/SiegeLord/RustAlgebloat/blob/master/algeb...
Matrix math. "Unsafe" all over the place, and unsafeness is exported, allowing the caller to do unsafe things. This is an example of why I occasionally stress the need for multidimensional array support at the language level. If the compiler knew about multidimensional arrays, it could optimize the subscript checks for them, avoiding code such as this.
The unsafe version. The caller can store anywhere in memory.
fn unsafe_set_idx(&self, $mat: &T, v: f64)
{
let $self_ = self;
let (r, c) = $rc_expr;
unsafe
{ $mat.raw_set(r, c, v) }
}
Safe, but inefficient version. The compiler can't hoist those checks out of loops. fn set_idx(&self, $mat: &T, v: f64)
{ let $self_ = self;
let (r, c) = $rc_expr;
assert!(r < $mat.nrow());
assert!(c < $mat.ncol());
unsafe
{ $mat.raw_set(r, c, v) }
}
This is the sort of thing that leads to exploits in code that reads things like JPEG files. Yet you can't do much better in Rust. That's why doing multidimensional arrays in macros and templates isn't good enough.But if those checks are in asserts, it's tougher. Is LLVM allowed to fail an assert early? If the array is 0..999, and the index is 0..1000, a subscript out of range condition will occur on iteration 1001. For best performance, you want to detect the subscript out of range condition at the point it becomes inevitable, rather than checking on every iteration and failing on iteration 1001. (Although technically you could generate a special case.)
But that requires special treatment of "assert". In Rust, "assert!" is just a macro. The compiler can't optimize it that aggressively and fail early. Especially since you can now catch assertion failures during unwinding.
If all those optimizations really exist, why is there code like this (at https://github.com/SiegeLord/RustAlgebloat/blob/master/algeb...)?
MatrixMul<LHS, RHS>
{ unsafe fn raw_get(&self, r: usize, c: usize) -> f64
{ let mut ret = 0.0;
for z in 0..self.lhs.ncol()
{ ret += self.lhs.raw_get(r, z) * self.rhs.raw_get(z, c); }
ret
}
}
If what you say is true, all that unsafe stuff is unnecessary.Or you could just use a little unsafe code to implement iterators and matrix multiplication on top of the array, and then never use unsafe with the matrices anyway. This is what gets done with regular arrays and vectors. Directly indexing these types is pretty rare. You would do the same with matrices. "Put it in the language" is not a solution for everything, especially when the tools are there to build it with almost the same level of ergonomics.
Requiring unsafe for designing some kinds of abstractions is not a bad thing. There's no need to shove everything into the language if it can be implemented as a library with a smattering of unsafe code.
It's almost equivalent, really. Making a mistake when writing this unsafe code and making a mistake in the builtin optimization are mostly equivalent risks. There's nothing inherently worse about doing it as a library, aside from the minor issue that the indexing syntax isn't so great.
for i in 0..X {
if !(i < X) { panic!() }
...
}
It is capable of doing this with the various induction variable passes it has, showing that i < X is always true.There's also the IRCE (inductive range check elimination) [0] pass—which isn't listed in that document—that should help even more, if/when it is enabled.
Hopefully you agree at this point that LLVM is perfectly capable of handling multidimensional arrays and matrices even if it doesn't have optimisation passes with names that obviously apply to them specifically. Most of the optimisations you keep talking about are basic consequences of other standard passes, a fact that has been pointed out many times to you before.
> Consider a matrix multiply, the most common operation in number-crunching
Also an operation you won't be implementing manually if you actually care about performance (better implementation: call a BLAS library). And, any high performance implementation will be doing more than the naive triple-nested loop (blocking, SIMD, etc.). Focusing on this operation is somewhat missing the forest for the trees.
> You're indexing through three 2D matrices along both axes. The indices are usually controlled by FOR statements, so the compiler knows the range the indices can take.
Yes exactly, the compiler knows the range of the indices.
> If the compiler knows about multidimensional arrays, it's easy to make those checks once at FOR loop entry.
This is a non-sequitur: the compiler can still make/move the checks based on what the code looks like, without having to have a hardcoded concept of arrays. Ensuring optimisation passes are powerful enough to handle these sort of relatively straight-forward cases is more general than relying on a language-level concept of multidimensional arrays, as it allows them to apply to cases which can't quite be implemented with an array directly.
> Although technically you could generate a special case
This is exactly what compilers do, see the IRCE documentation.
> The compiler can't optimize it that aggressively and fail early
It certainly can: these sort of assertions shouldn't ever trigger, and with appropriate top level assertions (e.g. asserting that the dimensions of the incoming matrices work for multiplication) the compiler can easily see this. Also, the compiler can hoist assertions early if this cannot be observed externally, e.g. the loop kernel only mutates locals.
> If all those optimizations really exist, why is there code like this (at https://github.com/SiegeLord/RustAlgebloat/blob/master/algeb...)?
That code is almost 2 years old, and so the compiler will have improved since then. Of course, the compiler may not have improved enough (e.g. IRCE still may not be enabled). Additionally, there's a lot of uncertainity/lack of clarity (as demonstrated by your own comments) about exactly what the compiler can do, and so people may use `unsafe` unnecessarily. I've certainly been guilty of this myself, and been glad when people have pushed back in code-review, making me work a bit harder to get equal (or better) performance in safe code.
[0]: http://llvm.org/docs/doxygen/html/InductiveRangeCheckElimina...
This is the other problem with not having built-in multidimensional arrays. There are so many implementations to choose from. Here's the list of all 34 Rust matrix math packages.[1]
Looking at crate "matrixmultiply", it's all unsafe code.[2] That's because it's C written in Rust:
pub unsafe fn sgemm(
m: usize, k: usize, n: usize,
alpha: f32,
a: *const f32, rsa: isize, csa: isize,
b: *const f32, rsb: isize, csb: isize,
beta: f32,
c: *mut f32, rsc: isize, csc: isize)
Arrays? What arrays? Raw pointers are good enough, right? What could possibly go wrong?Still trying to find a package with a matrix multiply in safe Rust code with the subscript checks optimized out.
Update: checked crate "ndarray". Indexing is unsafe.[3]
Update: checked crate "matrices". Empty project.
Update: Checked "scirust" - more raw pointer manipulation.[4]
Not finding real-world matrix libraries in which all those fantastic checking optimizations are used and working. I'd like to see that stuff in action.
[1] https://libraries.io/search?keywords=matrix&languages=Rust [2] https://docs.rs/crate/matrixmultiply/0.1.13/source/src/gemm.... [3] https://github.com/bluss/rust-ndarray/blob/master/src/dimens... [4] https://github.com/indigits/scirust/blob/master/src/matrix/m...
Rust doesn't know about multidimensional arrays, so they have to be supported in a library. This matrix representation was extracted from the "algebloat" crate:
pub struct Matrix
{ data: Vec<f64>,
nrow: usize,
ncol: usize
}
Access functions, get and set written in the obvious way: #[inline]
pub fn get(&self, r: usize, c: usize) -> f64
{ assert!(r < self.nrow);
assert!(c < self.ncol);
self.data[c + r * self.ncol] // index
}
#[inline]
pub fn set(&mut self, r: usize, c: usize, v: f64)
{ assert!(r < self.nrow);
assert!(c < self.ncol);
self.data[c + r & self.ncol] = v; // set value
}
Matrix multiply, written in the obvious way: #[inline]
pub fn mult(&self, other: &Matrix, result: &mut Matrix)
{ assert!(self.ncol == result.ncol); // out of the loop checks
assert!(self.nrow == result.nrow);
assert!(self.ncol == other.nrow);
assert!(self.nrow == other.ncol);
for r in 0..self.nrow
{ for c in 0..self.ncol
{ let mut tot = 0.0;
for rr in 0..self.nrow
{ tot += self.get(rr,c) * other.get(r,rr); }
result.set(r, c, tot);
}
}
}
Generated code for the inner loop: ///
/// mult - matrix multiply, straightforward approach
///
#[inline]
pub fn mult(&self, other: &Matrix, result: &mut Matrix)
{ assert!(self.ncol == result.ncol); // out of the loop checks
assert!(self.nrow == result.nrow);
assert!(self.ncol == other.nrow);
assert!(self.nrow == other.ncol);
for r in 0..self.nrow
{ for c in 0..self.ncol
{ let mut tot = 0.0;
for rr in 0..self.nrow
{ tot += self.get(rr,c) * other.get(r,rr); }
result.set(r, c, tot);
}
}
}
Generated code for the inner loop. "rustc 1.14.0", optimization level "opt-level = 3", AMD64 instruction set. Debug mode, so asserts should be checked. (Not sure about this; it is
possible that opt-level=3, "Aggressive" disables some checking. Documentation is unclear on this.) .LBB8_50:
.Ltmp254: ; in Matrix::get()
.loc 1 122 0 ; self.data[c + r * self.ncol]
movq %rdi, %rax
mulq %r11 ; doing the multiply for the subscript every time
jo .LBB8_74 ; and checking it for overflow
addq %rsi, %rax ; doing the add. No strength reduction
jb .LBB8_76 ; another check
.Ltmp255:
.loc 17 1362 0
cmpq %rax, %r9
jbe .LBB8_72 ; and another check
.Ltmp256:
.loc 17 1362 0 is_stmt 0
cmpq %rcx, %r15
jbe .LBB8_78 ; array overflow check
.Ltmp257:
.loc 1 188 0 is_stmt 1
incq %rdi
.Ltmp258:
.loc 1 122 0
movsd (%r12,%rax,8), %xmm1
.Ltmp259:
.loc 1 147 0
mulsd (%rbx,%rcx,8), %xmm1 ; The real work: floating multiply
addsd %xmm1, %xmm0 ; and the add
.Ltmp260:
.loc 18 746 0
incq %rcx
cmpq %r14, %rdi
jb .LBB8_50 ; loop counter check - required
This is relatively decent code. the compiler got rid of multiple checks on the same values. There was some strength reduction of indices, too; only the subscript that's traversing the "wrong way" generated a multiply. The ones that are advancing one element at a time along the underlying vector are just adds. About five instructions could come out, but it's not bad code.I've seen FORTRAN compilers do this kind of matrix multiply with a five instruction inner loop on a mainframe, incrementing pointers in registers for both dimensions. That's the advantage of multidimensional array support.
The point I made about multidimensional arrays stands - the compiler didn't strength reduce the multiply and eliminate the rest of the checks. That requires inferring too much from the user's code.
On the other hand, the code is good enough that using "unsafe" for performance reasons is very seldom justified.
However, there's a trick. Take a look at https://github.com/rust-lang/rust/commit/6a7bc47a8f1bf4441ca... where the Rust developers had the same issue, and managed to avoid the bound checks without having to use unsafe code. You have to somehow make the optimizer see that both sides have the same length, so it'll elide both bounds checks. A slice is actually a pair of a pointer to the first element, and the length. Their trick copies the length from one slice to the other (after a bound check, obviously), so the compiler is sure that the lengths are the same.
It's a mistake in the library to export that as safe, yes (filed an issue). But that unsafe code being unsafe has nothing to do with multidimensional arrays. It has to do with arrays in general.
Implementing multidimensional arrays in the language would not change the fact that they would have `get_unchecked()` and `set_unchecked()` methods. The only thing that would change would be that you might have slightly nicer syntax for them, and "one way to do it". It doesn't change the scope for optimizations, either. Subscript checks do get optimized out in 1D arrays and they should get optimized out in this case too. The way indexing works for 1D arrays is basically exactly the same; there's a pair of methods for checked and unchecked; the subscript operator is specified to do checked indexing via an Index impl, and the optimizer usually gets rid of the checks. This would not change if indexing for the 1D array type was not implemented by the language itself.
Sometimes when the invariants aren't easily seen by the optimizer it won't get optimized out, and that's when folks use unchecked indexing.
Now, there is a problem here, and that is that unchecked indexing is an unsafe operation in Rust, whether with 1-D or 2-D arrays. Could be fixed with dependent types, but that's a lot of complexity and 99% of the cases where dependent types would work would have been optimized anyway.
But this has nothing whatsoever to do with multidimensional arrays, and would not be helped at all by multidimensional arrays being in the language.
The compiler already knows enough about multidimensional arrays to be able to optimize things. Rust supports multidimensional arrays in the language. It just doesn't have syntax sugar for it; and syntax sugar can't really affect optimization.
> it might be possible to break it by instantiating it on a type with unusual semantics.
Can you give an example? A lot of the semantics are well-encoded in the marker traits, so as long as you correctly specify the right Sized/Copy/Send/Sync bounds these problems should go away.
(Panic safety could be an issue if you were calling methods on T, but you're not)
Rolling your own collections is typically not that great anyway; how much better is your linked list/hash table/skip list than everyone elses?
Reducing this unsafe surface area should result in outsized gains overall. I'm not familiar with the unsafe patterns, and to what extent they could be addressed/reduced through Rust language/compiler design, but that'd be one area where research focus might produce even more benefit.
The other weak point may be the LLVM itself, which I guess would be sort of a background hum of risk to Rust.
Outsider view: not sure of the relative levels of risk in each layer, and to what extent they can be addressed.
Can (or will) these unsafe patterns be capable of being addressed by future changes in Rust?
Just my 2p for others learning: for me, Rc::RefCell was what I was missing, even after I thought I was up to speed. I was fine using Channels for inter-thread communication and I never needed Arc, but use of Rc is common in the Rust ecosystem and a lot of my early fights with the borrow checker weren't fights I needed to have. In situations where it wasn't a trivial change, and I found myself banging my head against the wall, it was often because I was in a situation that called for a reference counted cell.
I started (and then stopped) learning Rust a few months ago (before their docs rewrite) and the borrow checker and concepts weren't explain in a way that I could understand. Programs wouldn't work at all, and when I read about Rc::RefCell I was scared because I wasn't sure how much garbage collection rust would do and whether or not it would actually be safe.
So yeah. Rust definitely has problems for beginners.
I should take this opportunity to mention that the book has been stagnating because I (along with Carol Nichols || Goulding) have been re-writing it out of tree: https://doc.rust-lang.org/stable/book/ The ownership/borrow checker stuff has been completely re-done. The existing book chapters explain it in the abstract; in the new book, we use String/&str as a motivating example, since that's an area that often trips up new Rust programmers.
IMO Rust has the challenges that are very similar to all other languages have. But if you are trying to make the leap from a GC'd language like Java or Python to Rust without ever having written C/C++, you should expect to have to learn not only new language concepts but new programming concepts.
I do not agree that Rust's challenges are the same challenges with other programming languages. Because Rust will not let you even play with the langauge unless you understand how to write safe programs (which is a concept that is defined in Rust). So there's a whole bootstrapping problem of "how the hell do I play with this thing to understand it if I can't play with it until I understand it fully". C, C++, Python, Go -- none of them have this issue. They will let you write bad code and won't stop you from running it (which I admit is not a good thing, I'm just saying that Rust's strengths are not without their downfalls).
For me the key was just doing toy problems over and over again. Without understanding the idioms, it's hard to dive in and write something real. I think Rust is approachable if you take it like that (or at least it was when I looked at it about 2 years ago).
Rust doesn't do magical garbage collection.
Rc<T> does reference counting, which is a form of garbage collection, but you get to choose where it gets applied, so it's a linear cost with no magical GC pauses or whatever. Rc isn't unique to Rust, it exists in C++ too.
RefCell makes mutation within an Rc safe. It panics if you misuse it.
http://manishearth.github.io/blog/2015/05/27/wrapper-types-i... has more on the Rc<RefCell<T>> pattern
In academia reference counting generally falls under the umbrella of GC. In the industry "GC" usually means "tracing GC".
The terminology is irrelevant to the point I'm making. I did make a distinction between RC and the "regular magical kind" of GC.
From a 1976 paper on automatic memory management https://www.cs.purdue.edu/homes/hosking/690M/deutsch.pdf
"Automatic reclamation of storage no longer in use is done by the following two techniques:
* Garbage collection * Reference counting"
Note that these are two separate items and that one is not a subset of the other (and vice versa). It is simply incorrect to call reference counting "garbage collection" when the academic and industrial practices already have explicit meanings to these two terms.
I'm not saying the term always is used that way in academia. I'm saying that it's sometimes used that way, and that it's valid to call RC a form of GC.
Ironically I said that RC was a form of GC to avoid precisely this argument, because if I say "Rust doesn't have GC" I'll have a bunch of folks telling me that RC is a form of GC. This is largely irrelevant to the point I was making, which explicitly distinguished between RC and tracing ("magic") GC -- which I suspected (but wasn't sure) was what was being asked about.
One angle that we've been focusing a lot on lately is productivity. Think about some feature of Rust that provides safety, like out borrow checker, which ensures that pointers don't do bad things. One way to think about this is safety, but another way is productivity: if your application segfaults, you have to track down what actually caused it, and then what caused the cause: "oh this pointer was null because I did something incorrect in this other part of my code." While having the compiler do these checks can sometimes feel like a productivity _loss_ at the beginning, you do a lot less debugging later, so it's a productivity _gain_ overall. (Or at least, we feel that way.)
Other features of Rust make it feel productive as well: a focus on iterators and composable iterator adapters is often more productive than manual for loops, Cargo and crates.io enable wide-spread code re-use[1], we've been working on good tooling for IDE integration for those that use IDEs, and "zero-cost abstractions" help you have nicer interfaces while not having to pay the cost for them. Like this: https://news.ycombinator.com/item?id=13117608
One interesting area where we've seen lots of production Rust usage is embedding Rust in other languages. Since Rust can expose functions that look like C to the outside world, you could write a Ruby, Python, Javascript, or whatever extension in Rust instead of C. And Rust's safety features are appealing here, since you may not be the kind of person who writes C all the time: that's why you write Ruby/etc in the first place.
Soon, we expect to see more usage of Rust on the server: the "tokio" project is gearing up for an initial release, which provides a foundation for asynchronous IO. Even before tokio is released, people are playing with this: crates.io is Rust on the backend, and "npm recently began replacing C and rewriting performance-critical bottlenecks in our registry service architecture with Rust": https://medium.com/npm-inc/npm-weekly-73-no-love-for-http-ur...
Basically, lots of places. We'll see how things go into the future!
1: Anecdote time: my favorite package is https://crates.io/crates/x86, which provides low-level bindings to various x86 platform details. Writing an operating system? No need to define your own IDT entries, just grab the library and http://gz.github.io/rust-x86/x86/irq/struct.IdtEntry.html has you covered. Re-usable packages in operating systems is super cool. On a higher level, this kind of thing is enabling Firefox to re-use Servo components more easily; Firefox can just pull chunks of servo with (relative) ease.
So someone over there is paying some degree of attention, at least :)
FWIW C# <-> Rust integration is really straightforward. You can actually pass delegates as C fn pointers and then treat them as a closures in Rust. Much less painful that I initially thought.
Yes, I do use VSCode, but only for dabbling on Rust during plane/train travels. The language is not yet at a level it just fits on MS stack and is requested by our customers on their Requests For Proposals.
Regarding C# <-> Rust interoperability, it is very badly documented. I gave up on searching for it, and just used C# <-> C++/CX <-> Rust instead.
Or I am very bad searching for it.
extern/dllimport[1] Covers most of it. There's automatic conversion for CString/string and delegates as function pointers. If you need to go deeper than that there's the marshal namespace[2]. Going through C++/CX sounds really painful.
[1] - https://msdn.microsoft.com/en-us/library/e59b22c5.aspx
[2] - https://msdn.microsoft.com/en-us/library/system.runtime.inte...
Not for someone that knows C++ since C++ARM. :)
I am pretty comfortable with C# and native interop, my issue was trying to map Rust strings with .NET UTF-16 ones, including passing ownership from Rust to .NET side.
Previously I was using JVM languages for this purpose, but grew weary of the resource footprint, and especially the unpredictable GC pauses. I am aware of the Azul JVM which removes GC pauses and of various Java techniques to avoid GC altogether, but switching to Rust provided a GC-less model from the ground-up, a powerful type system, and familiar functional programming facilities at no cost.
I haven't had the chance to try https://github.com/fede1024/rust-rdkafka yet, but it looks promising and partially wraps the C/C++ library https://github.com/edenhill/librdkafka.
I think it's great for any problem area where you'd instinctively reach for C. System utilities, bare metal development, etc. Stuff where you care about the precise layout of memory but would prefer that a simple but non-obvious mistake didn't end up as a high-profile CVE.
> What languages is it largely meant to replace/improve on.
C. It has C's straightforward machine model in mind, like C its memory behaviours are entirely predictable, it's entirely explicit about error handling (no hidden paths of errors exiting functions as in C++).
It improves on C by adding strict checking to make managing memory safely and avoiding race conditions tractable problems.
I'm really surprised you think that's the case. Most rust codebases I've worked with use Rc very sparingly, if at all. If a codebase does use Rc there's usually one central thing that is Rc'd, with everything else using regular memory management.
I have noticed that beginners coming from GCd languages often tend to structure their code in such a way that paints them into a corner where they must use Rc. This might be what hit you. I'm not really sure how to teach idiomatic Rust though.
Learners should read and write lots of Rust. That's a great way to learn what is idiomatic. There will be early missteps, but thankfully the language makes it more comfortable when you do things the right way.
But I would love to figure out how to teach this in a way that doesn't involve random codebases. Some folks learn by doing (I do!), but others fare better with tutorials.
Right.
In my case, I was using the nphysics library, which for unsurprising reasons uses Rc for entity handles.
Similarly, in my app's code, I used nphysics' proximity and contact handlers to watch for events that indicated a change in status. Those handlers needed to initialize/set values in a HashMap accessible (and periodically cleared by) the main sim loop -- 'one central thing' that is Rc'd.
It's not sprinkled everywhere in your codebase, and for obvious reasons most libraries don't need to use it. But in applications, a central loop + library-provided callbacks + some Rc'd state of doesn't seem too uncommon.
At least a carefully written garbage collector can free objects incrementally, and concurrently.
So memory management in Rust is certainly not a solved problem.
EDIT: Removed mention of RefCell.
Well, a different kind of unpredictability from other GCs.
And Rc is rare enough (IME) that this doesn't usually matter.
http://manishearth.github.io/blog/2016/08/18/gc-support-in-r...
Allowing Rust to generically adapt to an arbitrary Gc that it is nested inside of would be awesome.
Then again so does manual memory management if you're freeing a large object graph at once.
> At least a carefully written garbage collector can free objects incrementally, and concurrently.
You could get that with refcounting, stashing Rc0 objects in a list of items to free incrementally rather than freeing it all at once and synchronously.
This looks like nay-saying just for the sake of it, but in case you're actually just confused:
Rust isn't reference counted unless you yourself add reference counting. Rust tracks ownership in the compiler so it knows statically, at compile time, when an object needs to be freed, and the compiler emits the code to do the freeing in that spot. Rc is a utility function: you can choose to opt into reference counting on a per object basis. But if you're not actually typing Rc yourself then it's not happening. Rust is no more reference counted than the ability to build the same thing in C makes C a reference counted language.
How is it compared to other forms of GC?
It wouldn't be too hard to replace that Rc with an Rc that performs deferred drops by putting them in a queue and performing a number of drops at an oppurtune time (eg. using an idle-hook in a event loop).
It might even be possible to send all drops to a thread dedicated to dropping, but that would probably only work if the contents are Sync or Send. This form of async drops should always be possible for Arc.
But big pauses should rarely occur since refcounting is something that has to be added explicitly, Rust is a language where compound types are used extensively instead of building dynamic structures using collections for trivial things and the use of Drop trait is not something that can be fully relief upon (making that most solutions try to go without using it). This should make most drops pretty lightweight, even when dropping large numbers of items.
That's still a predictable pause: it's the number of references contained within the array. If you know your array will hold at most ten references, it will take at most the time to free the ten references. Also, you know this will happen when the last reference is dropped - no earlier, no later. If your program can't handle the pause at that moment, you hold the reference until it can. If you know another component (perhaps in your own call stack) holds a reference, you know it won't be freed. And so on.
True. But if those references themselves contain references, it may become tricky to manage all of this. You basically don't want to think about it.
But, like others mentioned, you can put the objects in a queue, and free them in the background. It would be interesting to know how such a solution stacks up, performancewise, against an incremental, concurrent garbage collector.
A lot of folks think that optimizations like having allocation arenas and spreading out deallocation pauses are unique to garbage collectors, but they're completely orthogonal to GC. You can have GCs with these optimizations, and GCs without. You can have regular allocators with these optimizations, and regular allocators without. Jemalloc does have arenas and stuff for allocation (I'm unsure if it spreads out deallocation loads). Of course, with a GC you can also defer the cost of iterating through large vectors and calling destructors.
But anyway, this only becomes a problem when you have large complicated Rc-trees in your application, which tends to not be the case in Rust.
In a GC'ed language such as Go, this would not be a problem. The flip-side is that I can safely and efficiently do updates in place, because Rust's reference counted type has a way to determine when you're holding the only reference.
The one time I've done this in C++, delayed deallocation actually made major loads/unloads worse for us on mass batch operations - more cache thrashing perhaps? Since this involved GPU resources, perhaps some bad interaction with the driver? I did keep around the "optimization" conditionally for smaller operations as this let us get rid of some hitching.
> In a GC'ed language such as Go, this would not be a problem.
In C# this traditionally manifested as lengthy GC pauses you had to jump through hoops to workaround. Modern GCs are much better these days, but it's not 100% solved.
But after 4-5 years of ruby the dynamism which initially was super cool, has grown a little frustrating and I long for a more sophisticated type system and a compile step, since frustrating bugs crop up from time to time that would be caught by that, and which slip through a hole in our test suite.
I've dabbled in Haskell and Rust and the whole "if it compiles it's likely to work, or at least the bugs will be significant logic ones" aspect of them is very cool.
I wonder, though, if the flipside is true? Maybe people who have a compiler and strong type system get frustrated with its rigidity over a number of years, and when they see ruby for the first time are blown away with what it can accomplish. After all, it took a number of years day in and day out with ruby to start seeing its blemishes, I wouldn't be surprised if the reverse were true.
So as someone who's now spent 4 years in the other grass, do you still find it greener? (Or am I wrong from the start, and maybe you like Rust for other reasons and its type system wasn't something that attracted you to it over ruby?)
* C
* Scheme (Gambit)
* Python 3.4
* Pony
I find all of them frustrating at times.
Python's dynamism is nice, but it's so damn inflexible, requiring my to follow the One True Way.
That can be good, and makes it easier to eliminate bad code in reviewing.
But it also means that you can fight with the interpreter to do what you want.
Scheme gives me both dynamism and flexibility, with less speed tradeoffs. Yay!
However, there was someone on my team obsessed with turning everything into a macro.
What's the point of first-class functions if you just macro everything?
Also, Scheme's stdlib is purposefully small, so you sometimes need to reinvent the wheel. Thankfully Scheme makes it both easy and pleasurable to do so.
The rundown is, though fast and dynamic, code review can be painful unless you follow standards, and you might run up against, "Oh... I need to build my own FTP library", though SLIB (depending on your circumstances), can eliminate some of that.
C's compiler feels like a breath of fresh air after that.
Unless you hit a runtime error, it makes writing code a breeze, quickly and efficiently.
Unfortunately, it doesn't protect you against yourself, or lazy people on the team trying to use void pointers for everything.
So code review can be harder, and catching edge case segmentation faults can be quite difficult.
Enter, Pony.
Pony deals with the same areas Rust does, but for reasons I won't go into, when both Rust and Pony were young, my team started using Pony for a few little things.
It has been... An experience.
Pony is very type safe, it is exception safe, and data race free, all enforced by the compiler.
Which sometimes means the compiler will sit there bashing your code for a full day before you realise that you weren't writing it safe enough, and you can't just tweak it, the whole thing needs to be rewritten.
Also, the Actor Model being central to everything can be quite annoying, when you just need to fit a couple extra functions in somewhere, but you aren't sure where's best.
However:
The compiler doesn't let you screw up.
What you do write can become massively concurrent fast programs, easily.
Pony also has some fairly good documents, for such a young language that really doesn't have the backing of Rust.
So...
I get frustrated with any language I am forced to deal with over time.
I wish Python had compile-time contracts, but there is mypy to ease the pain now.
I wish Scheme had a better stdlib, but it goes against the ethos. (See RSR6 community breakdown).
I wish C was safer, but it's lack of safety makes it easier to do things like JIT.
I wish Pony was more flexible, but it'll never let me point a shotgun at my own foot.
If I spend too long in any world... It's time for a breath of fresh air.
I find the language fascinating but I'm learning Erlang and I don't really want to get started with Pony at the same time. Do you have experience with Erlang?
what made you choose Pony?
Sorry for the bombardment of questions but I don't hear much about Pony. As a rust user, and a burgeoning fan of Erlang, I'm fascinated.
> What made you choose Pony?
We wanted a Type Safe language to run a REST API frontend. That is to say, we wanted to have something that could redirect requests to the appropriate servers, at scale, whilst maintaining Type Safety in the server itself. We got hit by so many issues from JSON's weak/absent typing causing runtime errors, we wanted something that could sanely prove that we wouldn't crash.
Whilst we were at it, the same language seemed great for the backend for a couple languages we develop in-house. For example, Owlang is a language developed for teaching with a group of teachers. The first iteration was written in Scheme, today, it runs on top of Pony.
We considered three languages:
* Rust
* Pony
* Erlang
Erlang was a bit odd to throw in the mix, but we couldn't ignore how amazing BEAM is, especially with recovery by dropping and creating thousands of processes without effort.
However, we found Rust was making too many breaking changes, and there was sort of a culture of using Rust Nightly, which doesn't give off a nice solid feel, or didn't back then.
Erlang has this sort of huge cognitive overload with it's syntax, making it take longer to learn, for no obvious benefits.
Pony's Philosophy[0] however, felt damn good, and we investigated their mathematical proof of safety and the like.
> Can you talk more about using Pony in a production environment?
Pony's got a few gotchas. Thankfully, they spell them out. [1]
We have been burned from long running procedures when we've had to tap out to an in-house C library. We went from using around 400mb of RAM for a group of processes, to running out of memory on a 16GB RAM machine.
That was a stupid mistake. However, using a Timer, like the docs tell you to, obliterated that, and we dropped back down to 4-500mb.
The non-preemptive nature of Pony's scheduler has taken a few people, myself included, some time to get used to.
> How do you find writing 'non-actor' code - like just doing some string manipulation?
If you try and avoid the fact you're using Actors, it'll bite you. Pony is designed for concurrency.
However, if we're just talking about methods and classes and so on, and how they feel... Pony feels simple, and easy to use.
Traits and interfaces make subtyping a breeze, especially interfaces.
But, being interested in Erlang, how about some pattern matching?
fun f(x: (String | None), y: U32): String =>
match (x, y)
| (None, _) => "none"
| (let s: String, 2) => s + " two"
| (let s: String, 3) => s + " three"
| (let s: String, let u: U32) if u > 14 => s + " other big integer"
| (let s: String, _) => s + " other small integer"
else
"something else"
end
One more bonus for something Pony has, that proved unexpectedly useful, is the ability to have multiple, or generate, Environments. In terms of Pony, that means argc, argv, argp. Being able to isolate environment variables between processes has been useful for making some configuration easier.[0] https://tutorial.ponylang.org/#the-pony-philosophy-get-stuff...
That's just weird. How much of an effort did you even put into it? Erlang may have ugly syntax but it's the simplest syntax of all popular languages and the cognitive load while writing is pretty much the lowest I've come across. The language itself is tiny. Or did yo mean you weren't used to a functional language and that was the cognitive overload?
Erlang seems simple on the surface, but each type seems to have it's own DSL, allowing things that seem the same, to not be, which means you have to hold more context in your head.
e.g. Are these two the same?
ensure(A, B) ->
if A == B ->
Does -> always mean function? Or does it mean something similar to progn or begin from CL and Scheme? Or does it vary with context?I had issues teaching where to use ; or . or end, and having people actually comprehend the reasons enough to do it without thinking.
At least with Scheme it's always )
We also investigated using Elixir, which has less overhead, and more people find easier to get up and running with. However, at the time Mix wasn't stable yet, which counted it out.
Never? I can't say it happens very often but occasionally the type system get's in the way, usually when you want the function to be type T1 but sometimes it'd be convenient to do a little bit more if the type is also T2.
T1 arg;
var foo = arg as T2;
if (foo != null)
....
Also, how do we format code here? fun f(x: (T1 | T2 | None)): ReturnType =>
match x
| T2 => /* do something */
| T1 => /* do something else */
| None => /* Handle badly behaved. Optional to have, compiler can enforce good behaviour if you want. */
Note: You can see formatting at https://news.ycombinator.com/formatdoc enum Either {
Type1(T1),
Type2(T2),
}
let's say you get a value of type Eitherthen you can match on them:
match value {
Type1(v) => func(v),
Type2(v) => func2(v),
} function(T1 thing) {
//do normal T1 stuff here
var foo = arg as T2;
if (foo != null)
foo.T2Stuff();
//more normal T1 stuff
}
If I'm understanding your example I would have to wrap all the T1 stuff in a match.Rust doesn't have type inheritance, so the only way you can define relationships between types is to have something external tell you how to package them together, which is what the enum would do here. So it's not a parent-child relationship between the types, it's a new thing which says "here is a thing which can be a T1 or a T2 but not both at once".
Not having OO does require you to think somewhat differently about how to build your data structures in Rust. I've found my intuition from Haskell is a stronger guide when working with Rust, but since Rust's type system is much more like that in an ML-family language this makes quite a bit of sense! Sadly this does increase the learning curve for more mainstream languages.
(Rust does have trait inheritance, but all that says is that to implement trait T2 you have to also implement trait T1, and thus anything working with a T2 can assume that thing is also a T1).
The above example was actually using two entirely different interfaces. It seems like rust is quite similar to the OO subset I prefer to use. In this particular case it would have been a better match because traits can implement methods.
Rust is on my "to learn" list over the holidays.
If T1 is the normal case, then there's a type in Rust called Result that handles this kind of pattern.
fn potentially_failing(thing: Result<Type1, Type2>) {
thing.map(|foo| success_case(foo));
}
this will only execute in the success case and ignore the error casemap_err does the same, but only touching the error case
I haven't had these frustrations in application code in Rust.
Although, the thing that bothers me about Python (haven't used Ruby) is the fact that it is interpreted. It's really annoying to have to run my program to discover syntax errors and typos, when a compiled language would do that for me. With interpreted and dynamically typed languages, you have to have 100% test coverage, because you have no idea if your code is even syntactically valid until you run it, and of course if it isn't valid, then you throw an exception. (Hopefully that crashes the program, unless someone was brilliant enough to catch all exceptions and not print any error messages, then it looks like everything worked fine until you discover it didn't.) So every change is a potential for an exception, until you execute that line of code to make sure there wasn't a misspelled variable name or something.
So, yeah, I love the compiler. (I also love Python for scripty stuff.)
I think you mean _semantically_ valid, e.g. no undefined variables. Most dynamic languages provide file-level syntax checking, rather than line-level. But I completely agree that this is the easiest argument against dynamic languages. Computers are better at bookkeeping than humans, and we should use them as such.
But isn't it the purpose of tools such as IDEs? PyCharm is pretty damn good at it. Of course, it's not always possible or convenient firing up an IDE. But I suppose editors such as vim should be fully capable of doing this via plugins as well.
So I _do_ like Rust for other reasons, or at least, I did initially. My opinion has changed over time. I used to say "I'd never write web application in Rust," but soon, I'd prefer it slightly. Needs some more libraries to come out, and they're close...
There's two sides to "is the flipside true"? One is, there's a huge variety of static type systems. Some of them are more flexible than others. So for example, if I had to use Java 1.4, or even Java 1.5 (which are very old, but were two releases I am very familiar with), I'd wish I was back in Ruby. This is because it's _too_ rigid; I can't express enough things in the type system to make up for the lack of flexibility. In a language with a better type system (and Java itself has a better one today, even), it changes the value of the equation.
In Rust though, there's another axis: speed, memory usage, stuff like that. I don't miss Ruby's flexibility here, because even basic operations in Ruby are _so so so_ much more expensive than in Rust. So when I think "oh yeah, I could do that, but if I think about how Ruby does it under the hood... it's not worth the cost." This makes me miss Ruby less as well.
Ruby is an excellent language, and I will always love working with it. But there's only so many hours in the day, and I've spent so much time with it already...
Also, retrospectives and posts where a core contributor summarizes the achievements and accomplishments of the previous few years are extremely valuable. These days software and tools evolve really fast and old resources fall off the face of the earth for a various reasons, and it's important to understand how things used to be and how we got to where we are now.
PS: I do not understand what irrefutable patterns means, so I google.
https://www.haskell.org/tutorial/patterns.html
in case anyone do not know yet.
https://mail.mozilla.org/pipermail/rust-dev/2014-June/010139...
The RustDT plugin for Eclipse makes the edit/save/compile/flag errors loop a lot tighter for me as a beginner. The autocomplete seems incomplete though and there's no hover docs or ctrl-click through to source that I remember.
I look forward to the day I can replace JS with Rust. It would be nice to have a quick way to start doing Rust in the browser. I tried installing emscripten to give it a go, but the version in Ubuntu's deb repo was too old. I don't remember why, but building it from source didn't work out.
I also got a bit lost trying to get boilerplate together for that purpose. Something like maven archetypes would be really great for Rust... Project templates that get you started with a working hello world for whatever you're trying to do.
Just some suggestions. I basically like where Rust is going.
Internationalization / Localization / Unicode (ICU): yup, no real progress. Needs some domain experts to drive it.
Date/Time : chrono is the most popular.
HTTP: hyper has been good for a few years now, tokio will make it async soon.
Crypto: there's lots of interesting work in this space, see ring and rustls.
SQL: Diesel is the gold standard here.
So, some progress! You're absolutely right that growing the crates ecosystem is vital; we've doubled our numbers in the last year, but there's always more to do.
> project templates
Thanks, for all the hard work, Steve. You're devotion helped turn me into someone who loves Rust.
It's taken lots of work. I think I'm a nicer and better person in Rust than I was in Ruby, but that's all I'll say about that. I am very much not a perfect human being.
> Honestly, a talk on how he stays so connected without being awake 24/7 would be interesting ;)
I wrote a blog post on that a while back: http://words.steveklabnik.com/how-do-you-find-the-time it's still mostly accurate kinda.
"For example, the biggest time I scrubbed out in my entire life was a choice that I made almost a decade ago: going to college."
A refreshing insight from a very intelligent individual!
This is a bummer. I remember when I first tried out rust the repl was still a thing. As someone who writes mostly Haskell and Python, interactivity with a language is a huge plus. I hope that eventually another repl comes around.
I have yet to be convinced that this is a good move. The attempted justifications I've seen hand-wave about problems with OOP, "composition over inheritance," etc. Meanwhile, in the real world, 99% of production code is OO and the paradigm continues to do its basic job--helping programmers model domain objects. Older languages like JavaScript and PHP have made their OO constructs more robust over time, which suggests to me that in the end a critical mass of developers will always clamor for this easily understandable and productivity-enhancing paradigm over some theoretically pure alternative.
It is unclear if there ever will be a Rust 2.0. Even those that want it agree that unless it's incredibly easy to upgrade to from Rust 1.x, it's a non-starter.
We still have some desire to indicate "epochs" of Rust development, as undoubtedly, things like idioms will change over time, new libraries will replace older ones, etc. But I'd prefer something like "modern C++", which could signify this kind of change, without giving up on backwards compatibility.
The idea is to come up with a set of processes for 2.0. If the community decides to do a breaking 2.0 for some reason, we should:
- Document exactly what has changed.
- Write extensive docs on upgrading
- Write good tools that do the upgrade for you when possible, and point out areas where they can't help with links to docs.
- Make it so that the cases where stuff isn't machine-upgradeable are minimal
- Be wary of actually removing deprecated APIs.
- Be wary of unnecessary extra breakage.
Python 3 took the attitude of "Okay, we're going to be breaking some things anyway, so let's break more!". I think that they had good reasons for doing that; and it makes sense in a way. But we may not want to follow the same philosophy.
Changing the behavior and representation of stdlib APIs. You would end up with e.g. two incompatible representations of String being used across the crate boundary.
Python 3 broke Python because it completely changed how strings work. Nothing of that magnitude will ever come to Rust.
I'm mostly just used to sublime and vim with minimal tooling so it doesn't matter as much for me. But I'll probably eventually write a plugin for RLS for sublime.
The only reason I have to install it.
cargo build --color=always
ls src/*.rs | entr -c cargo build
(entr is pretty awesome :)
If the programmer knows what the error is, and how to correct it, then their program should automatically do so, rather than molesting the user, as the purpose of computers is to speed up and automate. Why should the user have to foot the designers' bills?
Another cardinal sin and a sign of lack of thinking things through is breaking backward compatibility, like the example above: imagine you were an early Rust adopter, and wrote a non-trivial application in Rust; now imagine that the newest Rust compiler has several major performance and efficiency gains. It is understandable why one would want to re-compile one's application with the newest compiler, only to be thwarted by the authors' lack of understanding of just how important not breaking users' applications is. That is one of the core differences between engineering and hacking. What other land mines await potential Rust adopters from programmers who do not have any respect for their users' time?
But then the meaning of that syntactic construct differs contextually. Even if it could be parsed efficiently, it still goes against Rust's principal of favoring explicitness over convenience.
> Another cardinal sin and a sign of lack of thinking things through is breaking backward compatibility,
The examples Steve cited span throughout Rust's history, way before any commitment was made to backwards-compatibility. These sort of changes won't happen anymore (it'd require a "Rust 2.0").
Besides, the change we are talking about happened 4 years ago, when Rust emphatically did NOT have any compatibility guarantees and was in a process of rapid experimentation. This has changed a whole lot – it's a stable language now, and has been for a year and a half.
Calling changing the syntax on unsuspecting users _editing_code_ is really a stretch of gargantuan proportions. And I don't really care whether it's the compiler, the linker, or mega_mojo-2.546-alfa5-preview18, what ever it is, it runs on a computer so that whatever is being done would be fast, automated, and therefore efficient. If the programmer knew what to do with the input, the computer should churn through it, instead of chastizing the user and making them pay for the authors' oversight.
Otherwise, a computer makes no more sense than a toy does.
If I want to use a computer as a toy, I'll go play a video game. The rest of the time, it's a tool, a tool to do something an order of magnitude faster than a human could. If it slows me down because it requires me to babysit it, I have a problem.
I don't know if it's possible, though (considering libraries).
My idea is that if one could write in a language like Go (which has some kind of GC), but when one wants, say "I'll take care of this memory".
D has optional GC but their standard library is kind of fucked up.
Did you ever use Rust back when you had to write "move"? I did. When you wrote stuff like:
let (move x, move y) = (move z.a, (move z.b).append(move z.c));
It got old fast.Statements such as "But as of Rust 1.15, this restriction will be lifted, and one of the largest blockers of people using stable Rust will be eliminated!" and "It is unclear if there ever will be a Rust 2.0." (to avoid the "Python 2 / Python 3 effect") don't help.
I'm comparing this to Go (yes, yes, both Rustaceans and Gophers will yell "it's not the same!" - in vain), where the Go 1.0 syntax was set in stone in 2012, and remained mostly compatible now, 5 years since, and still it feels that the community is somewhat sparse and reinventing wheels on regular basis. The reality is that for complex projects and complex development environments, once "Rust 1.0" syntax gets stabilized, it will probably take 5+ years for the community to feel comfortable with it, and trust the language developers not to break things at a whim.
Remember that the still widely-popular Python 2.7 was released in 2010 (and it also sort-of set in stone the "Python 2 syntax"), and plan accordingly.
> It is unclear if there ever will be a Rust 2.0.
I'm not sure how this implies that Rust is changing too fast to be useful, but it's meant to state the opposite: we do not plan on making gratuitous backwards changes in the future.
> Remember that the still widely-popular Python 2.7 was released in 2010
I mostly remember that Python 2.7 is a straight descendent from Python 2.0[0], and as far as I'm concerned a much better language for all that was added in the meantime (though technically I only took up Python circa 2.3, a fair number of syntactic and semantic additions had already been performed).
[0] and actually older than that, while 2.0 added major features the main change was a switch towards a much more open and community-driven environment with less single-organisation control over the project, it introduced PEPs, sourceforge hosting of the tracker and source and a very large expansion in commit bits from ~7 at CNRI to ~27 people around the time 2.0 itself was released, technically 2.0 is a minor update to the 1.6 which had been released a few months earlier for contractual reasons.