Typestates in Rust (2018)
yoric.github.io
yoric.github.io
enum SenderState {
ReadyToSendHello,
HasSentHello,
HasSentNumber,
HasReceivedNumber
}
Sender<const S: SenderState> {
...
}
impl Sender<{SenderState::ReadyToSendHello}>{
...
}
One can then give the individual states extra parameters: HasSentNumber {
number: u32
}
(Note that in this case that doesn't make much sense, the number that is sent is more associated data then an actual type parameter. There is no real difference, at the type level, between HasSentNumber { number: 3 } and HasSentNumber { number: 6 }, and the compiler generating two types for this would be unnecessary. It is only an example of the syntax.) trait StateData {
type Data;
}
struct Sender<const S: SenderState> {
... // sender implementation
state_data: <Self as StateData>::Data
}
impl<const S: SenderState> StateData for Sender<{S}> {
// default: no associated data
default type Data = ();
}
struct HasSentNumberData {
sent: u32
}
// using a specialized impl (also an unstable feature).
impl StateData for Sender<{SenderState::HasSentNumber}> {
// instead of a complete struct we could also have done Data = u32,
// but a separate struct would enable a nicer constructor, default values, helper methods and so on.
// (ommitted in this example though)
type Data = HasSentNumberData;
}
struct HasReceivedNumberData {
sent: u32,
received: u32
}
impl StateData for Sender<{SenderState::HasReceivedNumber}> {
type Data = HasReceivedNumberData;
}
Example usage: impl Sender<{SenderState::HasReceivedNumber}> {
...
fn validate_received_number(&self) -> bool {
self.state_data.received == self.state_data.sent
}
}
let sender = Sender::<{SenderState::HasReceivedNumber}> { state_data: HasReceivedNumberData { sent: 3, received: 28 } };
assert_eq!(sender.validate_received_number() == false) // client sent wrong number!
Playground link to see this in action: https://play.rust-lang.org/?version=nightly&mode=release&edi...We've used them in our Bulletproofs MPC (multi-party computation) where cryptographic requirement is not to replay the protocol from the mid-point. Since the user is supposed to perform these transitions within their application on top of network messages, such strongly-typed API guarantees that, if their program has compiled successfully, then:
1) Steps are performed in correct order. 2) The protocol cannot be replayed from the intermediate state.
We have wrote about it here: https://medium.com/interstellar/bulletproofs-pre-release-fcb... - scroll to "strongly-typed multiparty computation".
I was merely offering a perspective on how the const generics might offer an alternative way to express the same ideas, and possibly change the idiomatic way to write them. Type level values, especially enums, IMHO offer an intuitive way to encode what is going on.
Also what's the deal with the fi ligature in the code? It throws off the kerning and holds no purpose.
I'm not seeing a ligature there, or at least not an obvious one. What browser/platform are you using?
Click the gear if it's too small for you.
Looks like I'm falling all the way back to consolas. I'm pretty sure SFMono comes with macOS, which explains the platform differences between us.
That might explain why the fi ligature has been a common complaint lately, as more people start using San Francisco in their sites.
Not in a way that would be relevant here. I’d expect this font stack to fall back to Menlo on macOS.
This kind of design pattern measurably saves us time; it reduces the volume of unit tests we need to write/update when we make changes (we still test, just... with less fear), and it prevents newer developers from making mistakes. I haven't played with Rust (beyond a couple of toy projects) yet, and articles like this remind me I'm looking forward to sinking my teeth in over the Christmas break.
Nice to see something like this in Rust! One thing that's a bit of a bummer, and I'm sure there are very good reasons for this, is that we HAVE to use every type argument of struct in its definition. If this restriction were to be relaxed, we wouldn't need the "state" field in the struct at all, and we could make the state type variable truly "phantom".
[1] https://www.schoolofhaskell.com/user/konn/prove-your-haskell...
EDIT: Typos.
https://doc.rust-lang.org/nomicon/phantom-data.html
https://doc.rust-lang.org/std/marker/struct.PhantomData.html
So
struct Sender<S> {
/// Actual implementation of network I/O.
inner: SenderImpl;
/// 0-sized field, doesn't exist at runtime.
state: S;
}
I believe could become struct Sender<S> {
/// Actual implementation of network I/O.
inner: SenderImpl;
/// 0-sized field, doesn't exist at runtime.
_marker: PhantomData<S>;
}
I've never used this PhantomData personally, so this might be wrong. Cheers!In Haskell, you can omit these fields entirely, and achieve the same thing just by annotating the function.
For example, in Haskell, we can have
data Const a b = Const a
whereas in Rust, it would be: struct Const<A, B> {
konst: A,
// does not exist at run-time
discard: PhantomData<B>
}Type parameter variance is inferred from usage (e.g. covariant for normal fields, contravariant for function arguments) and without a usage there's no way to infer it.
Indexed types have their uses but just think about this: for your length-indexed vector example, what if you want to read a file and put its lines into such a vector? Your return type needs the length but it's not known until runtime. Okay you use existential quantification, but you now also need a runtime representation of your types. That's done through singletons. But that destroys Haskell's phase separation and means the inefficient ways you use to prove things and encode naturals now leak to the runtime! Have fun with your unary numbers then.
Although the language itself doesn't guarantee that the value is not used again after a move, good static analyzers will provide a warning in that case, so it can still be safely used.
Not soundly. Such static analysis can be trivially defeated using e.g. virtual methods.
Really neat techniques!
Here another article, but with a more classical state machine impl.
https://insanitybit.github.io/2016/05/30/beyond-memory-safet...
https://github.com/jonhoo/rust-imap
I did not contribute to that library, but it's nice to see a practical implementation. They use the same approach, essentially, as I describe in my blog.
[0]: http://www.cs.cmu.edu/~aldrich/papers/classic/tse12-typestat...
[1]: https://stackoverflow.com/questions/3210025/what-is-typestat...
trait IndexedStateT[A, B, C]
Which signifies a typelevel state machine moving from state A to state B emitting a value of type C.I can only speak for scala, but i'm assuming haskell has singleton and literal types as well. Meaning that code like this works great.
object DoorOpen
object DoorClosed
class Door {
def open: IndexedState[DoorClosed.type, DoorOpen.type, Unit]
def close: IndexedState[DoorOpen.type, DoorClosed.type, Unit]
}
val d = Door()
for {
_ <- d.open() //works
_ <- d.close() // works
// _ <- d.close() //compile error
}
By my understanding of the article, it uses the borrow/move state to implement the state transistion. Is this generalizable to arbitrary state machines, or only a simple 2-state one?Each state is a type, methods consume their receiver (the "move state") and return the new state. You can't accidentally keep around a copy of the old state.
It's pretty awful to deal with when you're unsure of what the state machine should look like, or if there needs to be a lot more flexibility in how the data is accessed. Maintainability nightmare.
An example of this I ran into is a data processing pipeline architecture where each vertex of the processing graph had a processing function called in a loop on its own dedicated thread. Using the type state pattern helped clearly define the "life cycle" of each vertex and enforce it, which provided for some powerful synchronization guarantees (e.g. we could provide some elements of memory safety even when loading things through shared libraries). If you dug into it you could break things, but that would be more work than just following the pattern.
that is pretty neat.
fn close(self) { }
...which is—not incidentally—the exact implementation of `std::mem::drop`(https://doc.rust-lang.org/beta/std/mem/fn.drop.html).Haskell can do it with some effort, maybe for simple stuff phantom types and GADTs are enough. With linear types the explicit "consumption" can be modeled.
None are really as powerful but they come close and cover many use cases.
> The second error, however, is much harder to catch. Most programming languages support the necessary features to make this error hard, typically by closing the file upon destruction or at the end of a higher-order function call, but the only non-academic language that I know of that can actually entirely prevent that error is Rust.
Why not simply add an `is_closed` flag and throw an error if it is?
That's a way of doing it at runtime. The article describes catching the error at compile time.
I do wish that we had reliable RVO so that this could come at zero cost.
my_file.open(); // Error: this may fail.
1) If this can fail, then it should be a compile error to not test the result code.
2) IMHO it would be nice if there was something like Python's with statement to correctly close a file.
with open(filename, 'r') as f:
f.read()
# f.close() invoked automatically here
This prevents trying to close a file that is not opened.The idea of encoding a state machine into the types seems interesting.
let mut my_file = MyFile::open(path)?;
// Note the `?` above. It's a simple operator that asks
// the compiler to check whether the operation succeeded.
// The *only* way to obtain a `MyFile` object is to
// have a successful `MyFile::open`.
// At this line, `my_file` is a `MyFile`, which means
// that we may use it. File::open("foo");
// warning: unused `Result` that must be used
But that's also not necessarily an invalid program, it's just not a very useful one. Failing to open a file doesn't interrupt the program in any way, so File::open here will return either a file handle or an error, neither of which get used. However, if you open a file, don't check for an error, but still try to use the file, then your program won't compile: File::open("foo").read(&mut buffer);
// error: method `read` not found in `Result<File, Error>`
You need to be explicit about what to do in case the file didn't open, and Rust won't let you use your file handle unless you explicitly indicate what to do in that error case!All resources in Rust are implicitly like your Python with example -- when they go out of scope and destructors run, the file will be closed.
RE 2: Rust takes the C++ approach of allowing types to define their own cleanup which gets auto-invoked when the object goes out of scope, instead of forcing the calling code to worry about it. Specifically, Rust types can implement the auto-invoked "Drop" trait, basically equivalent to C++'s destructors. You only need to call close if you want to close a file early (e.g. before the end of the scope), in Rust, which is fairly rare.
I vastly prefer this approach, but it admittedly doesn't play nicely with the... less deterministic lifetimes of objects in a garbage collected system, so I can understand why Python, C#, etc. have more explicit scope syntax for cleanup.
`open` is a static function on the `File` object and an instance of `File` does not have a `close` method at all. So, as in this article, an instance of `File` can only be created if `open` succeeded and the file will be closed when dropped (either implicitly or explicitly). For example:
fn main() -> std::io::Result<()> {
// a file object is only created if File::create is successful
// otherwise main returns an error
let mut file = File::create("foo.txt")?;
file.write_all(b"Hello, world!")?;
Ok(())
} // file closed here, no close method needed.
This contrasts with many other languages where a class instance can be in an invalid state. This is something the typestate pattern helps to avoid.edit: also closing stdin/stdout is common.
https://github.com/rust-lang/rust/issues/32255 https://github.com/rust-lang/rust/issues/59567
In C++ or Rust, non-memory resource management is arguably even easier than Python.
RAII: Resource Acquisition Is Initialization.
(Apologies for the C++; can do the same thing in Rust.)
ManagedFile myfile("path.txt");
myfile.read();
`myfile` gets released whenever the scope finishes.That said, the C++ and Rust stdlibs don't look like that because opening and closing files is not exception-free and unlike Python, C++ and Rust don't have or don't prefer non-explicit error handling.
if let Ok(f) = File::open(filename) {
f.read();
// drop(f) automatic here
}
But Rust's move semantics give you extra flexibility, because you don't have to close it if you don't want to. let keep = None;
if let Ok(f) = File::open(filename) {
f.read();
if random() {
keep = Some(f);
}
}
// the file *may* be usable beyond the first scope,
// and it's still dropped correctly.
Values aren't dropped simply at the end of their initial scope (as in `with` or on-stack RAII), but after their last use, and language semantics allow the compiler to track globally where the last use is.For example, in Kotlin, if a method on state A returns state B, there's no way to invalidate all existing references to state A. Normally this invariant would be enforced on the caller's side, or perhaps by throwing an IllegalStateException if state A is called after producing state B.
1. Type erasure
2. Using sealed classes requires instantiation, while the Rust version is zero overhead.
There is no need to access type information at runtime here.