The Serde Rust Framework
serde.rs
serde.rs
It is more convenient to use serde to serialize/deserialize to some standard format like JSON or YAML than it is to write your own half-baked format and serialization/deserialization code --- even for the simplest tasks --- so you just stop doing the latter. This is a fundamental shift, and very good for your software.
"Convenience" here actually covers a lot of ground. The APIs are convenient. Boilerplate code is minimal (#[derive(Serialize)]). serde's attributes give you lots of control over the serialization/deserialization, while still being convenient to use. cargo makes importing serde into your project effortless. All of these are necessary to achieve that shift away from custom formats.
Before Rust I used to write a lot of one-off half-baked formats. Pernosco uses Rust and serde and it uses no half-baked formats. Everything's JSON, or YAML if it needs to be more human-editable, or bincode if it needs to be fast and compact and not human-readable or extensible.
For JS and Python, JSON.stringify/parse and json.loads/dumps are a step up, but you still end up with an untyped mess with no schema validation, which makes them only halfway solutions to me. I'm a static typing guy at heart, sue me.
Let's take JSON or YAML for example:
Rust's closest siblings C and C++ are typed, and you get a pretty much untyped messes when working with serialisation or deserialisation. You have to manually inspect every node or write manual ”NodeType" to struct conversation. They allow unwrapping to primitives at best (eg via template specialisation).
In Haskell it's a bit more awkward, but you can kinda pull something similar off in terms of ergonomics. I've not worked with JSON in Haskell that much that I need to look for it, but I don't remember there being something as ergonomic as serde.
Typescript doesn't support this type of validation either out of the box (there might be tools that add validation).
In Elixir there's the Poison library, or it could be done with Kernel.struct/2 (not 100% sure this will work though).
Of the "mainstream" languages Go is pretty much the only one that has as good ergonomics as Rust in that it supports it pretty much natively (via struct tags)
---
This list excludes codegen tools which can generate (de)serialisers from an external schema (a la capnp, grpc, jsonschema, etc).
var account = JsonConvert.DeserializeObject<Account>(string/stream)
That's it. No inspection or looping through trees. You can often get it to create objects via constructor for validation.
It's been like this for 15 years...
I've worked with the Microsoft tech stack very very little, so I've never bothered to learn C#.
I've completely forgot about Java; I haven't used it in ages. Now that you mentioned Java, I realised Kotlin has it too.
Swift does this too, using the Codable Protocol. The compiler generates the protocol conformance code for you, if you say a type is `Codable`, and the properties the type contains are all Codable types (Int, Float, String, URL, etc. conform to it), you don't have to do anything.
https://developer.apple.com/documentation/foundation/archive...
(1): which is a shame, I remember when it came out that it seemed like a nice language
I think Haskell's most popular JSON library aeson is basically the same thing and probably a source of inspiration for serde.
data MyCustomType = ... deriving (Generic, ToJSON, FromJSON)
or if you want more control you can use deriving-aeson -- I prefer snake case instead
data MyCustomType = ...
deriving Generic
deriving (FromJSON, ToJSON)
via CustomJSON '[FieldLabelModifier CamelToSnake]A generic JSON decoder would give you a dynamic structure that can contain anything, and then you'd have to pick it apart.
Serde just doesn't use reflection to do it.
Performance of a lot of the reflection-based parsing APIs also leaves something to be desired, which means that projects with more stringent performance requirements often have to resort to stuff like Avro. This still happens with serde, but much more rarely--besides having more compile-time information at its disposal and needing to allocate less, it also provides relatively straightforward hooks to achieve things like zero-copy deserialization for strings where the whole buffer is available at once.
How does Serde work with optional extra fields? I know you can use Option<>, but that implies you know they exist - what happens if fields get added in the future silently? I guess the schema changes then, which isn't good, but that might happen?
So you can define a struct with one status field and the right annotations and you will get exactly what you describe without having to write the code and it will still be almost as fast as doing that parsing yourself.
#[derive(Deserialize)]
struct Response { status: u64 }This also informs your second question. If new optional fields in json are added, unless you tell serde to complain about them, it won't, and you can add the new Option at your leisure.
You can define a struct with just the status field.
> How does Serde work with optional extra fields? I know you can use Option<>, but that implies you know they exist - what happens if fields get added in the future silently? I guess the schema changes then, which isn't good, but that might happen?
Unknown fields are skipped by default.
Is there any actual reason something like serde couldn't exist in other languages? None that I know of, but it doesn't, at least not when you consider just how pervasive serde is in the ecosystem (almost every library that stores things you want to serialize will have it). I really hope every language can standardize on something similar; it just makes development so much nicer and faster when you don't have to worry about this stuff.
This has not been remotely my experience. Deserializing has a lot of implications for ownership and you get into complicated deserializer trait specifications pretty quickly.
In particular, there doesn’t seem to be a way to deserialize a type that has a reference because you don’t seem to be able to tell serde who should own that memory. Maybe I’m wrong, but I’ve walked through this multiple times with experienced Rust users and no one could get it working.
Serde is impressive, I’m sure, but far from “don’t even have to think about it”.
We have hit one serde footgun, which is that skip-serializing-if corrupts your data with formats like bincode :-(. https://github.com/serde-rs/serde/issues/1732 really wish that could be fixed.
That said, it's incredibly refreshing to be able to focus on these semantic and migration concerns without having to worry about how to deal with parsing them later. serde doesn't solve enforcing backwards compatibility for you, but it sure makes doing the right thing a lot easier. You never have to worry about some code accidentally forgetting to check the version field, people carelessly mixing and matching messages with different versions, or fields that "should" never be set but are there for backwards compatibility. That's worth a lot to me!
public record Point(int x, int y) {}
// [...]
var objectMapper = new ObjectMapper();
var jsonString = objectMapper.writeValueAsString(new Point(2,1));
It isn't obvious to me what advantage Serde offers that you couldn't get with Jackson or other similar libraries in Java (although I get that Java and Rust are different languages, and that there may not be something this ergonomic in the likes of C++)C# standard library has DataContractSerializer for decades. That thing reads/write text XML by default, but with minor tweaks can do binary XML or JSON.
I like their binary XML format the most. Very fast because doesn’t waste time printing or parsing numbers or Base64 bytes. Also, pre-shared XML dictionary, and session-accumulating dynamic dictionary, makes the serialized representation very compact, sometimes an order of magnitude better than text/xml.
That's still (at least naively) using reflection, but it's a lot safer.
In practice, using reflection for every object being deserialised is far too slow and Jackson at least generates code at runtime rather than repeatedly reflecting. That's magic, but it's nice magic because it's not breaking the language: one can see that it's possible to implement it in pure Java.
So at compile time, just like serde works, you create your type safe serializer.
Manifold [1] an example framework that relies heavily on that, achieving similar things like Serde.
Java developers don't like to use these things too much though, unlike Rust developers who love their `#derive`, so you're mostly right that reflection-based serialization is still more common, but that's by choice.
Dependency Injection is in a similar situation: frameworks like Micronaut [2] can do it without reflection, and things like Google Dagger [3] have existed for several years that do the same thing... but still, most Spring Boot-based projects I know of use reflection-based DI as well... it's fast and the security issues have been largely mitigated nowadays.
[2] https://docs.micronaut.io/latest/guide/
[3] https://rskupnik.github.io/dependency-injection-in-pet-proje...
Reflection is indeed used there, but not to serialize. It's only used once per type, to generate code.
C# has multiple ways to generate code in runtime. One is Reflection.Emit, allows to manually generate bytecode instructions + metadata. For instance, one can build new types in runtime. A typical pattern is implementing manually-written abstract class or interface with generated types: this way the manually-written code can call into runtime generated one.
Another method is System.Linq.Expressions. This one allows to generate code (no new types though, just functions) from expression trees, and provides API to build and transform these expression trees.
Regardless on the method, the generated code is no different from manually written code. JIT compiler can even inline things across, when generated code calls manually written one, or vice versa.
There's a weird (not in a bad way) elegance here that reminds me of lisp.
Yes and no.
No because when you try to do unsupported things like calling a method on an object which doesn’t support one, you gonna get an appropriate runtime exception.
Yes because if you fail lower-level things like local parameter allocation, you gonna get an appropriate runtime exception but that one is (1) too late, I’d prefer such things to be detected when you emit the code, not when trying to use the generated code (2) Lacks the context.
Overall, when I can I’m using that higher-level System.Linq.Expressions for runtime codegen. Things are much nicer at that level. I only using the low-level thing when I need to emit new types, like there: https://github.com/Const-me/ComLightInterop/blob/master/ComL...
I've done some benchmarking on JSON -> in-memory borrow structs and it was pretty darn impressive while being fairly ergonomic to use compared to a streaming parser.
Serde generates heaps and heaps of generic code. This all gets optimized away to be very efficient, but only once it reaches LLVM.
Ever tried working on a crate with hundreds or thousands of de/serializable types? Compile times shoot through the roof really quickly, and serde is often the culprit.
The maintainer of serde also created `miniserde` [1] to tackle this problem, which uses dynamic dispatch instead of generics and can have 4x compile time improvements.
Due to Rusts lack of orphan instances you really depend on a pervasive standard for serialization, though, so the ecosystem is pretty locked in to serde.
For example you can pre-instatiate the most common type parameters and just stuck those instantiations into a binary library.
Also if you have a nice C++ compiler environment like Visual C++, you can even do edit-and-continue for many kinds of changes, and VS 2022 will double down on that capability.
Whatever the solution Rust comes up with I think will look at the practical component so as to actually see adoption. It seems like the Rust team is very good at studying the mistakes other languages make in their approaches to a problem and finding the right fit within the language/targeted domain.
With regards to your specific experience, I don't know that I'd say that Clearcase, HP-UX or aCC are main-stream professional C++ development environments. I think for that you're really looking at clang, gcc, & msvc with icc being a distant fourth that has itself moved to clang recently. The primary target OS would either be server Linux, Android, macOS, iOS with maybe FreeBSD picking up the slack. HP-UX has a niche but it's a small niche comparatively in terms of $ spent on developers in that ecosystem. In terms of source control, everyone is mostly on git or mercurial. Anyone not just refuses to get with the times because they have bought into their existing tool expertise but don't know how to adapt.
The compile time costs for the error handling crates is due to increase in build graph size, not due to the cost of running the error handling proc macros or the cost of compiling the code generated from them. For anything but a small project, that increase in graph size is probably not a big deal (and likely the project is already using much of the same deps). This cost does not increase as a project gets larger.
serde also probably has the build graph size problem, but the bigger problem is the cost generating and compiling the generated code. This cost tends to increase as a project gets larger.
Compilation speed isn't just one number. While 5min isn't bad at all for compiling a whole project from scratch including dependencies, 5min is a really long time for incremental builds.
I'm not sure why. The fact that Rust makes dyn traits significantly more verbose doesn't help, of course, but that can't be the only reason.
The main problem is that trait objects are very limited.
There is no downcasting or upcasting (eg C++ dynamic_cast) and trait objects are limited to a single trait. You can't have `Box<dyn A + B>`.
That leads to lots of headaches in practice.
I hear downcasting might be on the horizon, so that's something to look forward to.
I've dealt with a number of C++ based serialization libraries, and they always have serious downsides to the point I often end up hand serializing structures with one-off half-baked formats just like you, which is error prone and laborious.
It has powered Deno's op-layer since 1.9 (https://deno.com/blog/v1.9#faster-calls-into-rust-with-serde...) and has enabled significant improvements in opcall overhead (close to 100x) whilst also simplifying said op-layer.
I maintain the nodejs bindings for foundationdb. Foundationdb refuses to publish a wire protocol, so the bindings are implemented as native code wrapped by n_api. The code is a rat's nest of calls to methods like `napi_get_value_string_utf8` to parse javascript objects into C. (Eg [1]). As well as being difficult to read and write, I'm sure there's weird bugs lurking somewhere in all that boilerplate code. I've made my error checking a bit easier using macros, but that might have only made things worse.
I'd much prefer all that code to just be in rust. serde-v8 looks way easier to use than all the goopy serialization nonsense I'm doing now. (Though I'd want a serde_napi variant instead of going straight to v8).
[1] https://github.com/josephg/node-foundationdb/blob/c1165539e5...
A big part of its flexibility is how modular it is. Most common Rust libraries support serde serialization (at least as an optional feature), so if you use crates that do, you can plug any backend in and serialize those data structures to it. It doesn't even have to be string-based; I've been using Serde to store arbitrary data structures as objects on Google Cloud Firestore.
My experience with other language ecosystem (I'm thinking Haskell or Scala) is that serialization/deserialization libraries are limited to one format, with incompatible APIs. Even if the libraries in those ecosystem provide the same kind of guarantees, it still doesn't compare to serde. Typically any framework or library (cats, play, scalaz, akka) will be either inflexible with format supports or require glue code (provided by the user of the library or a 3rd party) to even work with your pet codec library, which end up also limited to a specific storage format.
With serde, the assumption is just that things work, and it's fantastic. Example: I have my web server (written in rust) communicating messages to the frontend (in elm) with a websocket. For fun, I decided to test druid's wasm output when it came out (druid is a UI framework), see how feasible it would be to replace the frontend. And I could just pick up the type definitions from my rust server, add a websocket dependency that works on wasm and _it just worked_! I could then switch from json to msgpack and it worked the same.
The strength of rust is the community. There is a strong sense of collaboration and under the rich diversity of frameworks and high level libraries, there is a culture of modular and shared code and contributions that makes every thing that more flexible and reliable.
Pretty cool!
Missing fields default to some kind of default “zero value” - for any type, even full blown structs. You can’t tell the difference of a missing field or the field having the default value.
So if a field has a validation of “must be greater than zero”, you can’t really give a proper error message. If user puts in “0” or omits the field, you always get a value of “0”.
Edit: whoops, misread which language it was you were frustrated with. Nevermind!
Go serialization really has quite a few issues apart from default initialization. Configuration by somewhat weird struct tag strings which are only evaluated/validated at runtime, de/serialization is all done via reflection (unless you want to use code generation), ...
This sounded surprising to me, then I remembered Go doesn’t have generics. I imagine this won’t be as much of an issue when it does?
I'd love to be proven wrong, but even with generics if you start using option types, it's going to be like JS and TypeScript where you'll have some parts nicely typed and others are the wild west
.map() et al only get you so far.
When extracting the value in Go you either have to return a pointer / error pair or abort. Both of which introduce awkward code patterns and potential for misuse.
If type inside of Option<_> has invalid values (like references which can't be null) one of these would be used for None, so Option<&T> is the same size as &T.
[0]: https://play.rust-lang.org/?version=stable&mode=debug&editio...
Can it actually use any general invalid bit pattern? I expect that it only supports an optimization for values that cannot be zero (and then it uses zero to indicate None).
Welcome to Go! This was a real pain for us at my last job, where we used go-swagger to generate REST API code. All the generated structs had pointers for _non-optional_ fields so it could also check that they were populated. Meanwhile, optional fields were not pointers because they could use the default value.
So then we ended up with lots of code doing checks for `is X nil` all over the place. Many of those wouldn't be needed in normal usage, since a nil value would be rejected by the validator earlier, but we also sometimes needed to make these structs by hand and pass them around.
All of this is necessary because of the lack of generics. With generics you can solve this with things like an `Option<T>` type, like Rust does.
In the end what I did is parse the JSON-string twice, once more with a “checkjson” library that tells me about missing fields.
The stylistic difference is that the Rust type is semantic: your optional values are Options. The Go type is silly: why does being a pointer (which is allowed to be nil) imply required?
Go is simple and "default case" which is simple and works for most cases... and then you have edge cases where you need to bend over backwards.
As an exercise, try to use comma (",") in JSON field name.
Say what you want about Go, but gophers get better jobs than you, meanwhile almost all rust jobs are shady crypto startups. Hard truth lol.
The increased complie times are the only real issue with Serde.
Especially if upgrading this crate into std would allow for a way to reduce the compile time “penalty” of very common formats like JSON, etc.
I think there should be a push for rust-lang to provide introspection. If you want to use serde, every crate needs to derive serde traits as well. This is not only bad architecture but it also prevents other libraries to get popular. For example, there are many "mini-serde" variants and I personally think, they are awesome (and surprisingly fast), yet you cannot practically use them because literally no crate supports anything else than serde.
https://github.com/dtolnay/miniserde
https://github.com/not-fl3/nanoserde
https://github.com/makepad/makepad/tree/master/render/micros...
(Or, better still, building generic representation of structs as records into the language/compiler, where the whole ecosystem can depend on it and we don't have to bloat compile times generating it with a macro).
KotlinX serialization works basically the same way too; it definitely seems like a case of convergent evolution
Anyway, it has nothing to do with Serde, which doesn't have a network component. For Serde you'd typically buffer the input first, or use external framing (like line-oriented JSON).
Tokio or async-std timeouts only cancel futures. The dedicated spawn_blocking threadpools also do not support cancellation.
To prevent such cases you would have to spawn the work into a dedicated thread pool and then kill the thread on timeout.
Terminating threads is a very complicated topic though and not generally possible for compiled languages, especially without introducing memory unsafety.
The only real workaround is to have the synchronous code regularly check for cancellation via an atomic bool or similar, and terminate if required.
Native busy-looping or deadlocked threads are tricky to kill safely, but that’s a general problem with threads in non-interpreted languages. Not specific to Serde, not related to Slow loris.
https://github.com/wulf/create-rust-app
Although there's a lot of work left to be done, I'd love to hear feedback, painpoints and ideas you have :)
Hoping to have documentation up at create-rust-app.dev soon~
So I think the work you are doing is going to be valuable for the community. I guess giving an option between the three popular web frameworks (actix, warp, rocket) won't be so easy as you might have very different way of integrating the authentication etc with each. Probably one way to do this is to keep things more modular so that the authentication can be called independently (among other features).
It is doable (e.g. see SafePSBT deserializer here https://github.com/sapio-lang/sapio/blob/master/ctv_emulator...), just not particularly ergonomic.
There are varying definitions of framework but what tend to be common is that they have a strong effect of "locking you into their ecosystem", which isn't really the case for serde.