Serde 1.0.0 for Rust released
github.com
github.com
It is a (de)serialization framework that can be quite easily implemented for various serialization formats like JSON, MessagePack, Yaml, toml, ...
It enables automatic and very performant (de)serialization of you data, into different formats.
Often with a simple:
` #[derive(Serialize, Deserialize)] struct Data { ... } `
(Never heard of it either)
I've never used it, but one benefit of Avro is that it can embed the schema in the serialized output, which allows clients to consume data even though they don't have the IDL, which isn't possible with Protobuf and Thrift. MessagePack does include types and map keys, but doesn't have a schema.
It looks like deserialization is mostly there, but not serialization.
I just found out recently that there's a couple work-arounds but none of them were obvious and for a framework that is zero-allocation based not having great support for arrays was a non-starter since that's 90% of the structures I was working with.
I love Rust but I feel like arrays in general are somewhat of an unfinished part of the language, you run into similar problems with Clone and Debug which can be frustrating.
I guess it's just jarring since the rest of Rust is so well crafted it felt really abrupt to run up against the 32 length fixed size limits. It's also somewhat annoying to have to coerce it into a slice via &array[..] instead of &array.
It's workable and when I finally get the library to a stable state I definitely plan to share what went well and where I saw pain points.
enum Syscall {
Open { pathname: Buffer, flags: u64, count: u64 },
Read { fd: u64, buf: Buffer, count: u64 },
...
}
I was very pleasantly surprised at being able to add a couple of dependencies, add #[derive(Serialize)] right above the struct, change two lines of my main driver program, and get an strace --json with useful output with no further effort. It's the sort of experience I expect from a higher-level language with dynamic types and runtime reflection, but available to me in a systems language.> Zero-copy deserialization
> […] The semantics of Rust guarantee that the input data outlives the period during which the output struct is in scope, meaning it is impossible to have dangling pointer errors as a result of losing the input data while the output struct still refers to it.
To be fair, so does the semantics of every language with garbage collection -- keeping things alive while there are references to them is the bread and butter of GC.
EDIT: I do think it's impressive that Rust can manage this without the overhead of GC. But the sentence from the release notes immediately before 'killercup's quote was: This uniquely Rust-y feature would be impossible or recklessly unsafe in languages other than Rust which struck me as a bit over-hyped.
char buf[1024];
while (read(fd, buf, 1024)) {
messages.push_back(deserialize(buf));
}
for (message: messages) {
print(message);
}
Garbage collection will keep buf alive, but won't guarantee that buf isn't being mutated while it's alive. Rust's ownership system will guarantee that. In Rust, the read() function would require a mutable (i.e., unique) reference to the buffer, and serde's deserialization function also requires a reference to the buffer, preventing read() from being callable while the deserialized objects continue to exist.I think doing reader/writer refcounting at runtime is hard because this is a case where there's nothing reasonable to do at runtime if you have incompatible references. At best you can do copy-on-write, but then you silently lose the zero-copy performance. You really want a compile-time error saying "You structured this code wrong, go redesign it or add some copies."
> Rust's "orphan rule" allows writing a trait impl only if either your crate defines the trait or defines one of the types the impl is for. That means if your code uses a type defined outside of your crate without a Serde impl, you resort to newtype wrappers or other obnoxious workarounds for serializing it.
I hadn't realized this. The justification is reasonable - preventing ambiguity when resolving traits - but it precludes one of the major use cases for traits/type classes: adapting a type from one library to an interface in another library without wrapper types.
import Module (just_one_symbol_i_need)
or to not import a couple of them. I think named instances a la PureScript let you do import Module (instance toJsonText, ...)
which would be great to have in Haskell, but probably breaks something deep inside instance resolution.PureScript as a "Haskell without the mistakes" really is a great idea.
It still takes a bit of boilerplate to write unfortunately.
The good news is this likely won't be the case forever, it looks like specialization (I think it's called) and some other type system features will allow you to write some code to resolve this. Intersection impl's I think it's called. Correct me if I'm wrong anyone.
For example, I've been looking at the CBOR library for Serde [1], and it's not obvious whether the library is full-featured, robust, actively supported, etc. Same goes for many of the other Serde formats. At the moment I'm likely to just choose JSON for new projects since I don't want to build on top of something that isn't known to be solid, but it would be really nice to be able to use binary formats for what I want to do.
Now that Serde is 1.0 it would be nice to do a push on the individual formats so that users coming in can tell what's active and well-supported vs a (possibly inactive) community contribution.
Should it? Or can it? Am curious, not a criticism.
But there's been some innovation done to make that even more efficient. It started with protocol buffers[1]. So you can then basically write .proto files which are based on protocol buffers' own schema[2] which look like this[3]. What's special about these schemas is that they can be strongly typed, and then after a schema is written, which is a .proto file (in case of protocol buffers), code for any language can be generated to receive and parse the binary encoded message properly, with proper error checking. This avoids re-writing code in different languages if a RPC protocol is changed. It also offers other advantages and you can look into the docs for that.
Then, the author of protocol buffers left Google and created something called Capnproto[4], which improved on it in many ways. Now, what I linked to is a rust program supporting capnproto's own schema[5]
[1]: https://developers.google.com/protocol-buffers/
[2]: https://developers.google.com/protocol-buffers/docs/proto3
[3]: https://github.com/WhisperSystems/libsignal-protocol-c/blob/...
But I thought since Avro is somewhat similar to capnproto and it uses Serde (in Rust) then capnproto could/should too.
But it sounds like it is a "should not"
I assume that Cow<'a, str> will only produce an owned string if there are strong escapes that need decoding? If so, that's probably the best approach for decoding appropriately, as you'd get zero-copy as long as no mutations are required, but round-tripping through JSON would still work right.
I wonder how the API would work. Would json::Value be modified to have a lifetime and contain Cow enums? How would it implement ToOwned? It almost seems like there would have to be separate types, a json::Value<'a> and json::OwnedValue.
You could have zero-copy roundtrips if you used escapes in the deserialized strings as well.
(I haven't tried it)