Fury – Fast multi-language serialization framework powered by JIT and Zero-copy
github.com
github.com
It's a hard choice to select one for the title, so I use JDK for it, which may be not a good choice.
BTW, fury support jit serialization for jdk17 record, which is super fast compared to other serialization frameworks such as kryo
Correct serialization can only be done for classes that are designed for it by having a well-known construction protocol, which are currently basic collections, enums, Strings, records, and classes that register specific serialization code.
Fury put much work on this to avoid the open dynamic deserialization risks.
But although `basic collections, enums, Strings, records, and classes that register specific serialization code` are the only objects should be allowed for serialization, and they are serialized by the construction protocol in fury already. There are so many applications has used the existing serialization protocol assumption, we have to keep compatible, otherwise most of application can't use fury.
Zero copy state transfer is a viable and high performance alternative.
Security and integrity can (should?) be implemented at a different layer.
The
Completely agreed, but also, adding my own emphasis there.
The technique has has enough gotchas, edge cases, and additional security considerations that it should really be an alternative that people can opt for when they need it, and never be the default approach.
What I am saying is that it is NOT necessary to implement secure serialisation and deserialisation because you CAN (and should?) implement it at a different layer and transparently to the application code.
You mean XML/JSON/binary etc? This is impossible :) - the only thing that comes close is ASN.1.
But serialisation can be easy: just don't do anything and let runtime handle that. Any compacting GC can be seen as serialiser/deserialiser (it does move object graph from one place to another) - and it does not need to run any application code to perform its task. Criu and any other live program ("object") migration methods are similar.
> and that accounts for cases where not both sides of the communication are trusted
I don't believe Java serialisation will ever be capable of handling deserialisation of untrusted input.
Item 85: Prefer alternatives to Java serialization
Item 86: Implement Serializable with great caution
Item 87: Consider using a custom serialized form
Item 88: Write readObject methods defensively
Item 89: For instance control, prefer enum types to readResolve
Item 90: Consider serialization proxies instead of serialized instances`writeReplace/readResolve` are an useful pattern, and can be used when needed.
And having an IDL is an advantage imho. Clear specs and interface design.
But on the other hand, the success of fury does not depend on whether I put in enough effort but on whether I can build a thriving community to involve more people to join us. I must admit that I am still learning in this area and have a long way to go.
The best way to avoid it I know of is to delegate responsibility and find other people to maintain the library with you (and/or charge for development/encourage corporate sponsorships!)
That said, maybe consider doing the best you can to encourage people to work on the issues that come up -- if an issue comes up, prep as much context as you can and hand it to the person if they're capable of committing code. And then, help people with their PRs once they have something up.
Also do things like post about your project in places Java places might look, or going into "Java weekly" style newsletters (they often have a "call for participation" at the bottom).
Oh and don't forget things like Hacktober (https://hacktoberfest.com/)
Good luck out there! Make sure to take care of yourself -- delivering value for free is not something many people do (thanks for even trying!), but doing it sustainably is harder than it looks, much better for the world long term to have people like you not burn out, even if a feature ships 6 months later (or never at all).
The suggestions about sharing the project in Java-related communities are also excellent. Things like Hacktober are also fantastic opportunity to involve more contributors. Thanks for those suggestions.
Your words of encouragement and reminder to take care of myself are truly heartwarming. It means a lot to me that you acknowledge the effort I'm putting into fury for free. I will strive to find a balance that allows me to continue contributing in the long term.
Thank you once again for your kind words and support. It's people like you who make the open source community a wonderful place to be.
A while back I actually implemented some utilities for generating classes at runtime and "linking" constants into them using constant dynamic. One day I might clean it up and release it as a library.
I'm not sure whether `MethodHandle` can generate the most complicated code, since the serialization logic here are more complicated even than the manual written code.
Janino can generated the bytecode for fury generated java code.
I must agree that generating bytecode directly has it's advantages, the abstraction is more low-level, thus more flexible, except more complicated for developing.
2. Data transfer for bigdata distributed systems: bigdata systems will handle much data, the data needs be transfered between workers. The serialization can be the bottleneck too. Spark RDD,Flink DataStream all have this bottleneck. The use bianry format such as arrow/tungsten format to reduce serialization overhead. But it's limited, sql oritiened, and can't express complex logic such as graph/event/domain.
3. Task scheduling: Image you have a mpp distributed systems, you need to schedule thounds of tasks in a process every sub-seconds. Serialization will be the bottleneck too.
More popular/newer examples are https://github.com/Cysharp/MemoryPack (which is similar to Fury with its own spec, C#-code first schema), https://github.com/MessagePack-CSharp/MessagePack-CSharp or even gRPC / Protobuf tooling https://github.com/grpc/grpc-dotnet
see https://github.com/chaokunyang/fury-benchmarks for detailed benchmark code.
- Does Fury support versioning?
- What is source of truth for the schema when dealing with multiple languages and the same data?
Here is the code for reproduction: https://github.com/chaokunyang/fury-benchmarks#fury-vs-micro...
Right? Flatbuffers supposedly being slower than protobufs is circumspect since most benchmarks I’ve seen for it in C++/Rust show it outperforming, especially since it too does zero-copy deser.
For http rest api, json is more suitable. For RPC/bigdata/game/scheduling, fury may be better