Is there any actual reason something like serde couldn't exist in other languages? None that I know of, but it doesn't, at least not when you consider just how pervasive serde is in the ecosystem (almost every library that stores things you want to serialize will have it). I really hope every language can standardize on something similar; it just makes development so much nicer and faster when you don't have to worry about this stuff.
This has not been remotely my experience. Deserializing has a lot of implications for ownership and you get into complicated deserializer trait specifications pretty quickly.
In particular, there doesn’t seem to be a way to deserialize a type that has a reference because you don’t seem to be able to tell serde who should own that memory. Maybe I’m wrong, but I’ve walked through this multiple times with experienced Rust users and no one could get it working.
Serde is impressive, I’m sure, but far from “don’t even have to think about it”.
We have hit one serde footgun, which is that skip-serializing-if corrupts your data with formats like bincode :-(. https://github.com/serde-rs/serde/issues/1732 really wish that could be fixed.
That said, it's incredibly refreshing to be able to focus on these semantic and migration concerns without having to worry about how to deal with parsing them later. serde doesn't solve enforcing backwards compatibility for you, but it sure makes doing the right thing a lot easier. You never have to worry about some code accidentally forgetting to check the version field, people carelessly mixing and matching messages with different versions, or fields that "should" never be set but are there for backwards compatibility. That's worth a lot to me!
public record Point(int x, int y) {}
// [...]
var objectMapper = new ObjectMapper();
var jsonString = objectMapper.writeValueAsString(new Point(2,1));
It isn't obvious to me what advantage Serde offers that you couldn't get with Jackson or other similar libraries in Java (although I get that Java and Rust are different languages, and that there may not be something this ergonomic in the likes of C++)C# standard library has DataContractSerializer for decades. That thing reads/write text XML by default, but with minor tweaks can do binary XML or JSON.
I like their binary XML format the most. Very fast because doesn’t waste time printing or parsing numbers or Base64 bytes. Also, pre-shared XML dictionary, and session-accumulating dynamic dictionary, makes the serialized representation very compact, sometimes an order of magnitude better than text/xml.
That's still (at least naively) using reflection, but it's a lot safer.
In practice, using reflection for every object being deserialised is far too slow and Jackson at least generates code at runtime rather than repeatedly reflecting. That's magic, but it's nice magic because it's not breaking the language: one can see that it's possible to implement it in pure Java.
So at compile time, just like serde works, you create your type safe serializer.
Manifold [1] an example framework that relies heavily on that, achieving similar things like Serde.
Java developers don't like to use these things too much though, unlike Rust developers who love their `#derive`, so you're mostly right that reflection-based serialization is still more common, but that's by choice.
Dependency Injection is in a similar situation: frameworks like Micronaut [2] can do it without reflection, and things like Google Dagger [3] have existed for several years that do the same thing... but still, most Spring Boot-based projects I know of use reflection-based DI as well... it's fast and the security issues have been largely mitigated nowadays.
[2] https://docs.micronaut.io/latest/guide/
[3] https://rskupnik.github.io/dependency-injection-in-pet-proje...
Reflection is indeed used there, but not to serialize. It's only used once per type, to generate code.
C# has multiple ways to generate code in runtime. One is Reflection.Emit, allows to manually generate bytecode instructions + metadata. For instance, one can build new types in runtime. A typical pattern is implementing manually-written abstract class or interface with generated types: this way the manually-written code can call into runtime generated one.
Another method is System.Linq.Expressions. This one allows to generate code (no new types though, just functions) from expression trees, and provides API to build and transform these expression trees.
Regardless on the method, the generated code is no different from manually written code. JIT compiler can even inline things across, when generated code calls manually written one, or vice versa.
There's a weird (not in a bad way) elegance here that reminds me of lisp.
Yes and no.
No because when you try to do unsupported things like calling a method on an object which doesn’t support one, you gonna get an appropriate runtime exception.
Yes because if you fail lower-level things like local parameter allocation, you gonna get an appropriate runtime exception but that one is (1) too late, I’d prefer such things to be detected when you emit the code, not when trying to use the generated code (2) Lacks the context.
Overall, when I can I’m using that higher-level System.Linq.Expressions for runtime codegen. Things are much nicer at that level. I only using the low-level thing when I need to emit new types, like there: https://github.com/Const-me/ComLightInterop/blob/master/ComL...
A generic JSON decoder would give you a dynamic structure that can contain anything, and then you'd have to pick it apart.
Serde just doesn't use reflection to do it.
Performance of a lot of the reflection-based parsing APIs also leaves something to be desired, which means that projects with more stringent performance requirements often have to resort to stuff like Avro. This still happens with serde, but much more rarely--besides having more compile-time information at its disposal and needing to allocate less, it also provides relatively straightforward hooks to achieve things like zero-copy deserialization for strings where the whole buffer is available at once.
How does Serde work with optional extra fields? I know you can use Option<>, but that implies you know they exist - what happens if fields get added in the future silently? I guess the schema changes then, which isn't good, but that might happen?
So you can define a struct with one status field and the right annotations and you will get exactly what you describe without having to write the code and it will still be almost as fast as doing that parsing yourself.
#[derive(Deserialize)]
struct Response { status: u64 }This also informs your second question. If new optional fields in json are added, unless you tell serde to complain about them, it won't, and you can add the new Option at your leisure.
You can define a struct with just the status field.
> How does Serde work with optional extra fields? I know you can use Option<>, but that implies you know they exist - what happens if fields get added in the future silently? I guess the schema changes then, which isn't good, but that might happen?
Unknown fields are skipped by default.
For JS and Python, JSON.stringify/parse and json.loads/dumps are a step up, but you still end up with an untyped mess with no schema validation, which makes them only halfway solutions to me. I'm a static typing guy at heart, sue me.
Let's take JSON or YAML for example:
Rust's closest siblings C and C++ are typed, and you get a pretty much untyped messes when working with serialisation or deserialisation. You have to manually inspect every node or write manual ”NodeType" to struct conversation. They allow unwrapping to primitives at best (eg via template specialisation).
In Haskell it's a bit more awkward, but you can kinda pull something similar off in terms of ergonomics. I've not worked with JSON in Haskell that much that I need to look for it, but I don't remember there being something as ergonomic as serde.
Typescript doesn't support this type of validation either out of the box (there might be tools that add validation).
In Elixir there's the Poison library, or it could be done with Kernel.struct/2 (not 100% sure this will work though).
Of the "mainstream" languages Go is pretty much the only one that has as good ergonomics as Rust in that it supports it pretty much natively (via struct tags)
---
This list excludes codegen tools which can generate (de)serialisers from an external schema (a la capnp, grpc, jsonschema, etc).
var account = JsonConvert.DeserializeObject<Account>(string/stream)
That's it. No inspection or looping through trees. You can often get it to create objects via constructor for validation.
It's been like this for 15 years...
I've worked with the Microsoft tech stack very very little, so I've never bothered to learn C#.
I've completely forgot about Java; I haven't used it in ages. Now that you mentioned Java, I realised Kotlin has it too.
Swift does this too, using the Codable Protocol. The compiler generates the protocol conformance code for you, if you say a type is `Codable`, and the properties the type contains are all Codable types (Int, Float, String, URL, etc. conform to it), you don't have to do anything.
https://developer.apple.com/documentation/foundation/archive...
(1): which is a shame, I remember when it came out that it seemed like a nice language
I think Haskell's most popular JSON library aeson is basically the same thing and probably a source of inspiration for serde.
data MyCustomType = ... deriving (Generic, ToJSON, FromJSON)
or if you want more control you can use deriving-aeson -- I prefer snake case instead
data MyCustomType = ...
deriving Generic
deriving (FromJSON, ToJSON)
via CustomJSON '[FieldLabelModifier CamelToSnake]