Rust for C++ Programmers Part 7: Data Types
featherweightmusings.blogspot.com
featherweightmusings.blogspot.com
fn can_fail(arg: bool) -> Result<(), StrBuf> {
if arg {
Ok(())
} else {
Err(StrBuf::from_str("Oops! Something went wrong."))
}
}
Here, there's no value when the function succeeds `()`. When it fails (I don't mean fail as in a `fail!()` or a panic or anything), you get a string back.You can pattern match the return value:
match can_fail(true) {
Ok(_) => {},
Err(err) => {}
}
This can get quite cumbersome, however. That's why there's a `try!` macro that adds composability. The idea is that if you have a function returning a `Result`, wherever that function is being called could also return a `Result`. fn higher_up() -> Result<(), StrBuf> {
try!(can_fail(true));
}
`try!` is simply: match $e { Ok(e) => e, Err(e) => return Err(e) }
This allows error to propagate up the chain.When you're working in big-ish projects, it'd be best to have something better than a simple `StrBuf` for an error. You'd probably want a struct:
pub struct LibError {
message: StrBuf,
error: Error
}
pub enum Error {
One,
Two,
Three
}
Where `Lib` is the library/project name.You can then create a new result type based on your new error type:
type LibResult<T> = Result<T, LibError>;
Then you can use `LibResult<T>` everywhere in your app.You can view this example done in a few of my own libraries (https://github.com/TheHydroImpulse/gossip.rs/blob/master/src...) and cargo (https://github.com/carlhuda/cargo/blob/master/src/cargo/util...).
That's a simple overview of it. Error handling is super simple, not verbose (thanks to try!) and in your control. Because of Rust's type system, things like `Result<T, E>` is available and are so much better than simple return values (like integers: -1 vs 0 uhhh)
enum FooError {
XWasFalse,
XWasUnknown
}
enum BarError {
VectorTooLarge,
FooErr(FooError)
}
fn foo(x: Option<bool>) -> Result<uint, FooError> {
match x {
Some(true) => Ok(42),
Some(false) => Err(XWasFalse),
_ => Err(XWasUnknown)
}
}
fn bar() -> Result<Vec<uint>, BarError> {
foo(None).or_else(|e| {
// We can recover from an XWasUnknown error returned by
// Foo, but not from a XWasFalse, so we return the error wrapped
// in bar's error type.
match e {
XWasFalse => Err(FooErr(e)),
XWasUnknown => Ok(99)
}
}).and_then(|n| {
if n < 100 {
let vec: Vec<()> = Vec::with_capacity(n);
Ok(vec)
} else {
Err(VectorTooLarge)
}
}).map(|vec| {
vec.iter().map(|_| { 42 }).collect()
})
}
edit: added a better example fn main() {
let mut out = std::io::stdout();
out.write([0x48, 0x65, 0x6c, 0x6c, 0x6f, 0x21]);
}
The `.write()` method returns a Result type. The output of compiling this program: $ rustc pxtl.rs
pxtl.rs:3:5: 3:53 warning: unused result which must be used, #[warn(unused_must_use)] on by default
pxtl.rs:3 out.write([0x48, 0x65, 0x6c, 0x6c, 0x6f, 0x21]);
^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Again, it's just a warning, so the program will compile and run as expected: $ ./pxtl
Hello!
If you really don't care about the return value here, the simplest (and probably best) way of appeasing this warning is to explicitly ignore the return type by making use of pattern matching: fn main() {
let mut out = std::io::stdout();
let _ = out.write([0x48, 0x65, 0x6c, 0x6c, 0x6f, 0x21]);
}
In Rust, the underscore is a pattern that means "I don't care about this thing, completely ignore it". The advantage of using the underscore here rather than an actual variable, e.g. `let x = out.write(...)`, is that it will be impossible to refer to the return value later on and thus explicitly expresses your intent to ignore it. (Furthermore, if you assigned the return value to a variable and then didn't use it later on, Rust would emit yet another warning, this time for having an unused variable.)The warning message alludes to a second way of silencing this error, which is by sticking the `#[allow(unused_must_use)]` attribute on top of your function. This will silence any warnings that arise from that function. If you wanted to disable this warning for your entire program, you could instead stick the `#![allow(unused_must_use)]` global attribute at the top of your program. Alternatively, you could compile the program with the `--allow unused_must_use` flag to completely silence all warnings of this type.
(One final note: in all cases where you see word "allow" used above, if you replace it with "deny" it will turn the warning into a compile-time error, thus enabling you to enforce a more rigorous error-handling strategy if you so choose.)
Anything else just leads to buggy software that has a try/catch block at the top level of the event loop/main/thread start function to deal with all the errors that leak out of its implementation and leave the process in an undefined state.
Exceptions are simply broken and awful. Java does them sorta right with checked exceptions, but the only safe thing is to not do them at all.
I just helped a teammate work through a bug the other day where somebody decided to "handle" a case-sensitivity problem in their home-brewed SqlLite data-access code by simply returning null for the data member if you got the wrong case. This resulted in improperly-cased column names producing objects with null members - no error happened because they were valid SQL queries, but the dictionary-reading code was silently failing when it was reading the result-set. If the program had just blown up when there was a miss on the dictionary of column names? We would've quickly found out about that stupid case-sensitivity.
Defensive coding just means your bugs go non-local and become data problems instead of exceptions.
As for rewrapping errors, yes, each subsystem should have its own error space. You don't lose data by nesting errors; on the contrary, each level can add additional context to an error result that makes debugging an unexpected issue far easier.
Their failure model lies in proper task supervision, coding for the expected case, and letting errors propagate up to the task level, where you can either kill a task, log and handle, propagate, or do whatever you wish.
That's not "catch-what-you-can-handle", that's "use functional programming and pervasive consideration of fault handling to ensure that you can handle faults at any layer".
Thank Guava for Throwables.propagate().
I don't use Java APIs that don't throw checked exceptions; if your code does that, I won't even consider working at your place of business, because that means you don't understand that you've written a massive pile of ill-defined failure-prone code.
Unchecked exceptions are GOTO on steroids, and those GOTOs are part of the API contract. Java makes exception handling explicit and compiler checked -- hacking around checked exceptions makes exception handling implicit and human-checked, meaning that there's absolutely no static verification of a critical component of your API contract.
The problem isn't checked exceptions, the problem is that exceptions suck, and the only way to use them in a way that doesn't expose your code base to implicit GOTO failure modes is to use checked exceptions.
On our production software, we don't use exceptions at all, except where required by an API; instead, we always use monadic error handling. We have an uncaught exception handler for threads/thread pools/etc that does one thing: log the exception, and terminate the running Java process via System.exit(), allowing the process's watchdog to restart the failed process.
By its very nature, an uncaught exception is unexpected and places the process in an unknown state; the only safe thing to do is exit. Since the throwing of an uncaught exception triggers full process failure, it very much encourages defensive, safe practices that ensure that all error cases are handled and compiler-checked.
The result: our code is far more stable and reliable than any other project I've worked on, especially projects that have made use of runtime exceptions.
Praise be! The feeling is mutual. I agree to disagree.
As far as APIs are concerned, the important thing is that the API is documented to throw something. It's not at all important that the compiler forces you to pollute either the immediate method's body or its signature and the body of the calling method, etc.
This is not unique to C#; if you review coding standards for C++, you'll see plenty of people who have adopted a no-exceptions approach, Google included. Simply put, exceptions are a failed experiment, because checked exceptions are the only mechanism by which the type of your methods is fully defined.
As far as API documentation, that something gets thrown is part of the return signature, and it's no more pollution than expressing the return type is.
Your willingness to employ ambiguity as a means to avoid having to do the work necessary to fully specify your system's behavior is a lazy and logically flawed position; it creates a cognitive load for all consumers of your APIs, and breaks the utility of the compiler that we rely on to write and maintain reliable software more easily.
> This seems like a step back from Exceptions to me. I
> want to be convinced otherwise, but I'm struggling to
> see how this is better than other mechanisms.
In a low-level language, guaranteeing memory safety in the face of resumable exceptions would be a nightmare. See Graydon's original post on the choice to avoid exceptions:https://mail.mozilla.org/pipermail/rust-dev/2013-April/00381...
Selected quote:
> In particular, to summarize for the impatient: once you get resumable
> exceptions, your code can only be correct if it leaves every data
> structure that might persist through an unwind-and-catch (that it
> acquired through &mut or @mut or the like, and live in an outer frame)
> in an internally-consistent state, at every possible exception-point.
> I.e. you have to write in transactional/atomic-writes style in order to
> be correct. This is both performance-punitive and very hard to get
> right. Most C++ code simply isn't correct in this sense. Convince
> yourself via a quick read through the GotWs strcat linked to:
> http://www.gotw.ca/gotw/059.htm
> http://www.gotw.ca/gotw/008.htm
For more on the topic of exception-safety in C++, see the following paper by Bjarne Stroustrup:http://www.stroustrup.com/except.pdf
I don't think that Rust's error handling solution is ideal, but I think that it might be approaching the best possible solution for its chosen context. Error handling is a hard problem!
The word resumable is important to that quote, which isn't arguing against exceptions in general.
That said, i am skeptical that there is a significant safety practical difference between the use of checked exceptions, and the use of return values with a try! macro. In both cases, you are forced to acknowledge in the code that an exception can be thrown, which means that you have a chance to do the right thing about consistency.
We could imagine a version of checked exceptions where individual throw sites have to be tagged. A parallel universe version of Java [1] might look like:
InputStream in = whatever();
int b = in.read() throw IOException;
Wouldn't that be exactly isomorphic to Rust's use of try! ?[1] No, not that Parallel Universe version of Java: http://blog.paralleluniverse.co/2014/05/01/modern-java/
struct IntPoint (int, int);
fn foo(x: IntPoint) {
let IntPoint(a, b) = x; // Note that we need the name of the tuple
// struct to destructure. struct Foo(int, int);
fn bar(Foo(a, b): Foo) {
println!("a: {}, b: {}", a, b);
}
fn main() {
let qux = Foo(1, 2);
bar(qux); // a: 1, b: 2
} let (a, b) = Foo(a,b);
is a fine destructuring. It would be a special case for let, since the pattern would have to be more explicit in function arguments and in match, but I think they have a good point. struct Meters(f64);
struct Miles(f64);
let meters = Meters(10.4);
let miles = metric_to_imperial(meters);
let Miles(raw_miles) = miles;
The single-arity case above constitutes the vast majority of tuple struct usage. And, as you may expect, in any other context besides tuple structs a single-arity tuple is completely silly (the only reason that we have syntax for single-arity tuples at all is to make writing macros easier).Ultimately it's just not a feature that would be pulling its weight. If you want a structure with multiple fields where destructuring is not necessary, just use a struct in the first place. Honestly, if we found a better way to support newtyping then I wouldn't be sad if we got rid of tuple structs entirely.
Hm.
For example, say you have two functions, where each function takes a single tuple of two floating point numbers:
// Converts a Cartesian coordinate to a polar coordinate
fn to_polar(coord: (f64, f64)) -> (f64, f64) { ... }
// Calculates the area of a rectangle
fn area(rect: (f64, f64)) -> f64 { ... }
Now, if you have a tuple that represents a coordinate, perhaps you don't want to feed it to the `area` function. Likewise for feeding a rectangle to the `to_polar` function. But because tuples are just structural types, something like `area(to_polar((2.5, 3.7))` is completely legal.If you didn't want to allow this, or even if you just wanted to have greater control over all these anonymous tuples floating around, you'd use tuple structs to give them names:
struct CarteCoord(f64, f64);
struct PolarCoord(f64, f64);
struct Rectangle(f64, f64);
struct Area(f64); // bonus round!
fn to_polar(coord: CarteCoord) -> PolarCoord { ... }
fn area(rect: Rectangle) -> Area { ... }
Taking the above steps makes `area(to_polar(CarteCoord(2.5, 3.7)))` a compile-time error. It's all about how strict you want to be with your types.To start you off, the best and most major difference to C is that Rust structs can have destructors.