Oracle plans to dump risky Java serialization
infoworld.com
infoworld.com
Apart from the horrible security, what annoyed me the most with serialization is the lack of control you have over the process. There doesn't seem to be a way to access serialized data as a simple parse tree or record sequence - you have to construct objects of the actual classes. If only one class is not available or has breaking changes, there is no (built-in) way to access anything inside the blob.
This is particularly fun if you want to refactor things. Suddenly package names, class names and names of private fields (!) are part of your public interface.
So if we could drop reflection/unsafe-based serialization and instead just got a simple parser/writer for java's binary object graph format, I'd be very happy.
If you want data then define a schema and write/read your data. Serialization should really never be used and certainly not for long-lived data storage.
There is one place where serialization becomes useful and that is for storing objects off-heap, IPC, and for short-term storage (ie snapshots like Android's parcelable). For these cases I'd like to see the JVM embrace not just immutable value objects but full-fledged structs that have a well-defined memory layout. You can do this today using off-heap Buffers and interfaces but language support is always good so there's a universal standard that everybody can build upon. Once that's in place there'd be no need to ever use serialization.
That said I can't imagine Oracle will simply remove support object serialization. It may be kicked out of the "core" JDK and become an optional module. The classes may be deprecated. But the functionality likely isn't going anywhere in the next ten years.
Even if they did remove it nobody should be using the standard object serialization anyways. If you're going to use serialization (and you shouldn't) then you should absolutely be using FST [0].
Lisp would, of course, disagree. :-)
(define data '(display code))
(define code (let ((pe primitive-eval)) (match data ((fun arg ...) (lambda () (apply (pe fun) arg))))))
(code)
Note that even when it's not an explicit feature serialization libraries have a tendency to become unexpectedly Turing-complete [1]. There are many, many systems out there that are vulnerable to "surprise Turing" attacks.
The only real defense against this is to have well-documented schemas in place and to thoroughly validate all incoming data. Ironically all the guys churning out schema-less JSON microservices think they're okay. They're not. As they say the truth is even worse than it appears...
[0] http://blog.diniscruz.com/2013/08/using-xmldecoder-to-execut...
[1] https://medium.com/@cowtowncoder/on-jackson-cves-dont-panic-...
In the common vernacular,
data = non-Turning complete commands
code = Turning complete commands
As I recall, the default behavior of Java (de)serializer, would only write the data fields one by one in binary format. The code itself, i.e. the bytecode of the methods and the class definition itself have to be in the classpath at application startup, so how could this malign code be injected as part of the attack?
It is similar to “return oriented programming”, which is one way to escalate c stack overflows to arbitrary code execution.
This sentence neatly describes one of the major ways things have changed since the 90s. Computing is now a war zone, rendering many of the elegant and beautiful distributed systems concepts discussed back then off the table. Instead we have walled gardens, closed platforms, closed systems, and closed pretty much everything. Anything open is immediately spammed and exploited to death.
this is not true. the jvm may choose to make serialization more complicated than it needs to be, but an object is nothing more than a collection of locations in memory
what's fundamentally wrong is sun/oracle's insistence on running constructors, not serialization itself
Many other language-specific serialization mechanisms just write arbitrary binary data into stream which can then only be parsed by ad-hoc recursive descent parser implemented by equivalent of java's readObject().
[edit] so why the downvote? We have YAML, JSON, protobuf, Thrift, Avro. Yup, these serialize "contents" rather than "structure + contents" but one gets interop with other technologies for free. Every tech mentioned above is so simple to use that removing Java serialization is a no-brainer.
It's not necessarily the serialization format at fault, but rather developer assumptions.
YAML is of course a format that is designed to instantiate objects as described in the source, rather than as checked by the destination, so I wouldn't want the world to adopt YAML instead of Java serialization.)
Its good that its been added, although it would have been better if it'd been there from the beginning and the safe version was default and you had to invoke 'unsafe_load' if you wanted complex object instantiation, to hopefully encourage even novices to think twice before doing it with tainted input.
Huh. Where can I read more about this?
* Python's yaml.load() hapilly executes arbitrary code.
* Some Perl YAML libraries deserialize objects by default:
http://blogs.perl.org/users/tinita/2018/02/safely-load-untru...
Do they?
https://github.com/yaml/pyyaml doesn't even mention safe_load().
And this is the docstring for yaml.load():
Parse the first YAML document in a stream
and produce the corresponding Python object.
This doesn't sound very scary, and certainly doesn't imply possible code execution.For java backwards compatibility used to be very important, so this is big news for the platform.
I think this remoting the bytecode with serialisation madness was once upon a time very important part of RMI/serialisation - back in the thin client java days this was supposed to be the way to distribute code across a network link, security was not the very first priority in the nineties (beats me why they made JRMP non routable)
java.rmi.server.hostname=myhostname.com
java.rmi.server.useLocalHostname=true
You also can tunnel JRMP through HTTP - there was a CGI script called java-rmi dating back to the late 1990's that I think was still distributed through Java 8 (!), and also an RMI Servlet Handler which was a bit more robust/performant. Spring also still has the RmiServiceExporter and HttpInvokerServiceExporter.I remember building Java applets and servers that did fixed income quotes & bond trading systems via streamed encrypted serialized Java objects circa 1999-2000. What a security nightmare, but no one knew better.
I feel old.
Java serialization has not many ways around it, you have to trust the sources.
XML, JSON, ... used naively with reflection (e.g. new XStream().fromXML(...)) exposes the exact same issues. Custom homemade parsers are also very likely somehow vulnerable.
The deserialization attack is a way to make the deserializer exploits vulnerable classes that exist in your classpath. It's not generating or executing malicious code by itself.
The standard classes should be ok, or at least fixed quickly.
Things like commons-collection will likely never be fixed. It might be considered as a feature from some point of view.
Check this out: https://github.com/frohoff/ysoserial
My 2 cents: - Secure your sources - Know you format, do not rely on reflection to parse text or binary data - Watch out with your classpath, but you can never know what new vulnerability will pop next weeks
JVM based distributed grid computing services use Java object serialization to distribute queries throughout the grid. This allows users to write their queries as arbitrary Java code which will be executed on every node without having to deploy the code on every node.
> How do closures get shipped around?
> Every closure is an object of a particular class. When the closure is being sent it gets serialised to a binary form, send over the wire to a remote node and deserialised there. The remote node should have the closure's class in its classpath or enable peerClassLoading in order to load the class from the sender side.
I wish they would just rename it something sufficiently ominous sounding that people wouldn't think about using it on untrusted data sources.
Maybe AribitraryCodeAndDataSerialization
As you can imagine, any sufficiently large codebase is likely to have a million different ways this can be leveraged to get arbitrary code execution.
Be curious to do a Github wide grep for ObjectOutputStream or something similar and see what it's like in open source land.
Wait, what? Java serialization does not serialize and deserialize code. The only thing encoding behavior when serializing is class names. The receiving system needs to know those classes/be able to load them.
Apparently there are some vulnerabilities in the implementation, but they are not an inherent part of it as some people seem to think.
...but then, if the developer is writing custom deserialization code that does black magic (like, executing the equivalent of eval() on a content of a String field), any serialized format is affected, be it binary, JSON, YAML, etc
A lot of things are currently built on top of serialization, from JMX to almost everything in Java EE including servlet sessions.
https://www.alphabot.com/security/blog/2017/net/How-to-confi...
In the JSON case linked, I presume there is a root type given when you are attempting to deserialize a document. However, if one of the properties of that type is ambiguous (say System.Object), and the deserialization algorithm looks for a 'type' property in the JSON with a class name to determine what is instantiated, then there can be all sorts of unintentional types that might be built by the processing of that malicious JSON.
Json.NET seems to allow the same behavior, but has it disabled by default.
> In fact the only kind that is not vulnerable is the default: TypeNameHandling.None
> Binary serialization can be dangerous. Never deserialize data from an untrusted source and never round-trip serialized data to systems not under your control.
The killed it in .NET Core 1.x; but brought it back in .NET Core 2.x due to back compat and interop complaints
I've been using .NET binary serialization for dirt cheap local snapshots for the purpose of undo/redo. It's perfect for this.
Using the Newtonsoft library you can configure it using `JsonSerializerSettings.ReferenceLoopHandling`, and with with protobuf-net you set `AsReference` on your class's `ProtoContractAttribute`. I don't think MessagePack-CSharp or msgpack-cli support cyclic references tho.
[1] https://googleprojectzero.blogspot.co.uk/2017/04/exploiting-...
Full disclosure, I’m the author of that blog post.
(Everything that tends to come up on HB, anyway)
Do you mean that you don't trust the protobuf implementation itself to write the correct bytes, or are you worried you may have written the proto file wrong?
If it's the first case, protobuf is a widely used format that has been rigorously tested in the field. Provided you use it for a major language, you should be fine.
If it's the second case - could you not just serialize and then deserialize, and check that the objects that pop out again are the same as the ones that you sent in?
I've even seen Google projects use text protos for config files.
Or to decode binary files, use the protoc tool.
Examples: Java Serialization, Python marshal and pickle, Ruby marshal, Perl Data::Dumper.
But on a deeper level, if a programmer doesn't know anything about security, then this sort of hole will continue to happen even if java serialization is disabled (just a bit harder to screw up). I m not that big of a fan of making security decisions without programmer's input, since you'd assume a professional programmer should know better anyway.
Since there won’t be any code that gets executed, that particular vector is closed.
Of course if the code reading that object does something stupid…
Recent releases added an opt-in feature to filter which classes are allowed to be deserialized, but there's still a horrible amount of open, unauthenticated network ports that take in serialized java.
The programmer trying to get the same failures in a post-serialization world would presumably have to find or build a new system with the same design issues.
This is the big one: as soon as you deserialize incoming data, any library on the classpath becomes a potential source of remote-callable snippets. And when a vulnerability resides in a library, exploits will tend to be compatible across applications, which makes them far more likely to actually hit you than any custom weaknesses.
There has to be a few specific classes that are common, and that have methods that the attacker can expect will be run (such as toString conversion, comparison etc), where the attacker can control the data used.
E.g. if he knows that serializing a certain collection type containing Font objects will use some platform native code that reads the font data, he can then pass a corrupt font in and fool the deserializer into an out of bounds read. Or something like that. I vaguely remember hearing about one of these attacks and I can't find it. It would be interesting to hear about some real world attacks.
This class of problems was very common when that serialization appeared in Java. It was possible to inject code into pretty much anything, e-mails (1), office documents (2), web pages. Since that time, most developers / managers have learned their lesson and started to pay attention / allocate resources so modern software/protocols/formats tend to be more secure, at least on average.
Both that VBS, and DOC/XLS, contain some code inside. But users though they are just data files, so they opened these things, running the code.
Also, old MS office files were not limited to embedded VBA. They are OLE compound files, so a specially crafted file can create and run ActiveX objects (often implemented as Win32 DLLs registered under HKEY_CLASSES_ROOT key of the registry) installed on user’s system. This is very similar to the way Java serializer creates and runs objects when desterilizing data received from untrusted source.
I’m just not seeing how this is a problem with the language. Leave anything on an unprotected port taking unsanitised input and it’s vulnerable no matter what it is written in.
Sanitizing the input yourself means not using it. You can't sanitize it.
I don’t think you’re comparing apples with apples.
Even in a default configuration, the majority of services that show up on a network don’t give you instant code exec. Maybe tomcat servers with default passwords, and a few other things.
But (de)serialization isn’t securable at all. You can’t add auth, you can’t WAF it, you can’t fix the underlying vulnerabilities.
There is absolutely nothing which has that level of vulnerability and lack of security on a modern network.
OK, let's say I overwrite part of your system with something evil, knowing that your app will Class.forName() it. Is that a problem with the ClassLoader or is the problem that your perimeter is already compromised anyway?
Java has serialization, and everyone uses it, and it’s a security nightmare.
As uncomfortable as it may feel to me, I kinda agree with Oracles approach here: kill serialization, make things more secure. If they have to change core language/libraries, so be it.
You were trying to say that anything left on a network socket will be compromised, and that’s simply not true. Most software is pretty solid. Anything in C will have more than it’s fair share of upcoming patches, but unlike Java serialization, you can’t fix a single lib to fix 99% of bugs.. or we would’ve surely done that.
The default Java serialization is one of the easiest way to serialize instance of objects but there are many other ways, and many other risky ways among them.
It seems to be always the same problem: ClassLoader access. Couldn't there be a way to let the deserializers use a specific ClassLoader?
I mean some sort of (Sandboxed)ObjectInputStream that uses a specific ClassLoader defined in the JRE config. The sandboxed contexts could be defined in something like java.security, .policy, to define what it is supposed to know and when/where it is supposed to be used.
Edit, they switched to Protocol Buffers in 0.23
In these cases, the changes would break existing code. However, who knows when will Oracle decide to remove the Java serialization API. I expect it will take a few years and the situation will be different on the Spark side then.
This is going to break a lot of code.
This might be your approach if say your session cookie is based on serialized Java.
(However, most people give up on this approach - java serialization is also very inefficient space-wise, and the cookie will get too big for the browser to honor)
> Warning
> The pickle module is not secure against erroneous or maliciously constructed data. Never unpickle data received from an untrusted or unauthenticated source.
As with many of the dynamic language deserializers, such as the Ruby YAML one, for pickle it's not even an exploit or something... it's a feature of the code that it can call methods, and getting to arbitrary methods isn't that hard.
RMI required it. Mr. Reinhold has apparently forgotten the scene in late 90s. CORBA anyone?
Or maybe I'll just move on to something else anyways as I'm really rather sick of writing these syntactically crippled lambdas. The streaming stuff is almost good.
And yeah, it would be nice if eventually a spring cleaning could remove some of the mistakes - Enumeration, Date/Calendar, Hashtable, ClassLoader, etc.
I put mistake in quotes because there are situations where RMI (and Java serialization) work fine: trusted, reliable networks like cluster or grid computing.