Log4Shell Still Has Sting in the Tail
spectrum.ieee.org
spectrum.ieee.org
You may be surprised when I tell you what the Apache Software foundation's yearly budget is. You'd think for software that is used by practically every Fortune 500 company and most governments, it would be something reasonable. Maybe a few hundred million dollars a year to pay for a reasonable full-time staff, right?
It turns out... it's about $2 million a year. (Wikipedia[0])
This helps explain to me why the devs of Log4j directly uploaded the file "JNDIExploit.java" (the POC) to GitHub while they were patching. (Here is a full analysis about what happened[1].)
They're not security people. They're volunteers working on this in addition to their full-time job.
What kind of brave soul wants to trudge through and maintain log4j in their spare time for zero compensation? I appreciate the people that are capable of doing that, but I think they are rare!
This whole entire vulnerability was eye opening for everybody and I have actually spent the last year building tooling on GitHub to help fix the problems that Log4Shell exposed.
If you have 2 seconds to try that out or just Star the repo[2], it would be very helpful!
0: Log4j revenue https://en.wikipedia.org/wiki/The_Apache_Software_Foundation
1: "How to Discuss and Fix Vulnerabilities in Open Source" https://www.lunasec.io/docs/blog/how-to-mitigate-open-source...
2: GitHub project building better dependency patching tools https://github.com/lunasec-io/lunasec
Can't this sentence be changed for almost any open source project?
Compare that to something like Bun.js[0] which is "sexy" and written in a "cool" programming language (Zig). Or Wasp[1] which is built with Haskell and is trying to define a new programming language designed to make common dev patterns less painful.
Those projects are naturally going to soak up smart people that have extra energy to share because they hate their day job but need to pay their bills. (imo)
Who is left that wants to bang their head against a legacy codebase like Log4j? Maybe somebody that feels there is "clout" to be had from it? (Spitballing here, I honestly don't know!)
There are a LOT of Java SWEs floating around...
There are people that like to clean messes and make people life better from the shadows. These tend to be called engineers.
You may be interested in https://www.fordfoundation.org/work/learning/research-report...
It's not clear to me as an outsider what exactly the Apache foundation is doing for these projects. It feels like Apache is willing to accept code donations from anyone and is willing to attach the foundation's name to code that isn't widely used, actively maintained, or may just be abandonware.
I have soooo much more confidence in CNCF projects. The conditions for graduating as a CNCF project include criteria like that your project must be in use by multiple real companies, have maintainers who are (paid) employees of multiple different companies, and get a professional security audit.
That's why I'm allergic to Apache software. A lot of it is overengineered, insecure, legacy abandonware.
That’s incorrect. Projects need to report quarterly and need a Project Management Committee of at least three people, or they are retired. Retired projects may not make releases.
(Source: past ASF board member, who used to review those reports each month.)
There are a fair number of retired projects, and others that may become retired within the near-to-medium term. The ASF has been around for a while, and every software project has a life cycle. Those are still associated with the ASF brand because Google, whatcha gonna do? An explicit retirement policy overseen by a board is still superior to how the vast majority of open source projects approach end-of-life.
The project has few, if any, volunteers, and there are security problems known to be actively exploited, yet the ASF is not willing to work to find a viable solution.
Something I have a harder time understanding is how it came to be that Apache Thrift and Facebook Thrift both exist as competing implementations of the same software originated by the same company.
Only a skeleton crew of paid developers stayed with Open Office, enough to cut releases regularly but not even to fix the security issues actively exploited. All distributions moved with the developers, but there is a discoverability problem which has led to mostly Windows continuing to install the unmaintained version.
The ASF could have fixed this quickly, either by helping out with the trademark issues, moving with the developers, or at least moving the unmaintained version to the attic and steering new users towards the actively developed version.
But they collectively decided to sit on their hands as users continued to install unmaintained software rather than take the slightest risk of offending one of their members. From an outside perspective, all of this was completely unnecessary.
The Thrift situation is another example where some active stewardship could have made a difference.
Hey Oracle.
— Larry Ellison, probably
The current Apache OpenOffice "Vice President" ie the leader of its developers such as they are, is Jim Jagielski, who was also an Apache Director back when a previous "Vice President" suggested the project should be wound up altogether. ASF should have accepted, they did not.
There should be billions in funding of OSS from governments each year. Think that's ridiculous? Consider how much Microsoft makes in support payments from governments around the world. The US government has a 92 billion dollar software budget.
I mean, look at this: https://www.reuters.com/article/us-usa-cyber-microsoft-exclu...
Society figured out a way of forcing people to pay for common goods a long time ago, and it's taxation.
If Apache tried to use it without doing it in the context of a formal deal or sponsorship, it wouldn’t last long.
No such thing as a free lunch, or whatever.
The core infrastructure team is paid, and that’s been the case since fairly early on — there’s too much work not suitable for a purely volunteer crew. Infra is half the budget.
In all practical sense, no one "at Apache" has released any remediations for log4j. There's just a bunch of random people, probably like 3 around the world, who released these patches.
"how reliant many companies are on third-party and open source code over which they have little control or visibility" - that's crap. They have complete visibility, and a good deal of control. Companies just choose not to spend money doing so. Which is true of their own code, as well.
What the hell is "log4jshell" that is a completely stupid name. It confers some meaning that isn't there.
It's a pun on Log4J and describes the vulnerability: all you need to get reverse shell is for the application to log input.
What meaning do you feel it confers that isn't there?
(The trend of subsequent vulnerabilities being dubbed "...4Shell" is infuriating, though.)
It is an interesting exercise in the power of language though. The name "Log4Shell" was literally picked because there wasn't a CVE and people running searches for "log4j RCE" were seeing a vuln in Log4j v1 from 2016. There needed to be a unique identifier that people could search for and tag on Twitter.
The name just... sort of got picked arbitrarily while I was very tired after doing many hours of security research and reading through Enterprise Java code lol. (You're right that "shell" was used to express RCE, ala Shellshock or other major vulns.)
I think the subsequent vulns (Spring4Shell, Text4Shell) just use the format of the name to be more catchy and ride the coat tails of Log4Shell's notoriety. (Like a meme!)
> I think the subsequent vulns (Spring4Shell, Text4Shell) just use the format of the name to be more catchy and ride the coat tails of Log4Shell's notoriety. (Like a meme!)
100%. The problem is that it's not only lazy (like how every controversy is X-gate), it also makes executives and senior leaders shit their pants. Log4Shell was a serious wake-up call for many companies and forced them to take security and software seriously. Subsequent clout-chasing or memed vulnerabilities that evoke its name risk creating "alarm fatigue" — "The Boy Who Cried '4Shell'", so to speak.
When i read here about Log4J, is was a real "Holy shit moment". It was so easy to use, it could've been a nebula exercise.
I contacted my PO asap, we found out that we had no inventory of what our VM were and what java versions were installed (worked at a bank btw). We then contacted the secops team, who just found out, and triggered all alarms. It was all hands on deck time. Quite a fun two weeks tbh.
Not familiar with the term PO in the way you are actually using it. What management level does that represent?
It's basically the same managment level as the tech lead/lead dev.
I understand this sound like scrum/agile gibberish, but honestly this was like having a technical manager/salesman who could also help you debug stuff, probably the most usefull post of "agile at scale"
All the log4j code is open. All the discussion are open. All the commit log is open. This is the case for the overwhelming majority of "open source code" projects that many companies. It's hard to see how more "control or visibility" could be reasonably enacted for a project like log4j.
All you have to do is actually get involved instead of being an apathetic consumerist and then start crying when it's not all perfect.
I've blogged about this to mixed reception:
https://dave.autonoma.ca/blog/2022/01/08/logging-code-smell/
> The Apache Software Foundation, which maintains the open-source tool, quickly released a patch...
Apache horribly mismanaged this and did not release a patch until it was already widely known and being exploited in the wild. They also messed up and had to release several subsequent patches to actually fix the vulnerability. This greatly contributed to "[the] mad scramble to plug the gaps".
You can watch the chaos unfold by reading through the GitHub PR. Remember, this was disclosed to Apache in November. https://github.com/apache/logging-log4j2/pull/608#issuecomme...
It may be technically incorrect, but most people wouldn't understand if you got into the minutiae.
It's a distinction without a difference for most people, and trying to clarify the structure of the foundation to executives and senior leaders (technical and non-technical) only makes their perception of any Apache-branded or open source projects worse.
That isn't comparable.
- the name of the project is "Apache Log4J"
- the tag-line is "*Apache Log4j* 2 is an upgrade to Log4j that provides significant improvements over its predecessor..."
- it's hosted under the Apache organization
- the documentation, release notes, and bug tracker is under the apache.org domain.
- the artifacts are published under the org.apache namespace
- etc.
No competent person in the industry is going to think that GitHub is associated with the projects it hosts. There are also hundreds of projects stewarded by Eclipse, Red Hat, and other similar organizations that avoid this confusion by using sensible branding.
What you're saying is more equivalent to "no reasonable person would think that Vitamin Water contains vitamins and is healthy".
https://www.businessinsider.com/coca-cola-glacau-vitaminwate...
I am not sure how to make my point any clearer.
If the concern is that "most people" will have a wrong perception, I don't see how talking losely about Apache will help. Apache is irrelevant in this context. Is your goal to make "most people" steer away from anything Apache?
They put in an "attempt to evaluate strings" feature. Literally as bad as someone doing a `console.log(eval(input))`.
Not only is it that bad, but also it's something with just a tiny bit of inconvenience you can get with either String.format or the likes of slf4j's formatting ( `logger.info("{} value", foo.get());` )
They took the worst case for an injection attack and applied it everywhere.
So this is one of the cases where I'd personally would not be surprised if in the far future someone admits on their death bed that malice was involved.
It's really easy for someone to see a feature like "interpolate strings better than ever" in their string library, and say to themselves, "Hey, that sounds good, let's turn it on!" without realizing that "better than ever" means "with arbitrary lookup from a string into a class of some sort", and then chain a few things together and hey presto you've got an arbitrary code execution vulnerability.
I personally think allowing arbitrary string -> class lookup for all classes in the runtime is automatically a security code smell at best if not outright antipattern, but I don't think this is at all well understood yet. (While the general pattern is hard to avoid, you should have to register all classes/structs/values/whatever you want to automatically load from somehow. This prevents things like "I can use the OS object to execute arbitrary shell code" from automatically and unexpectedly creeping in.) Dynamic languages are all but based on such a capability, with their "eval" support. It's probably still going to be decades before this is generally understood "programmer wisdom", partially because the particular confluence of features that enables this bug is not common. It's just that when it does happen, it's catastrophic. But it's not actually that common.
Spot on. Case in point: the maintainer of snakeyaml furiously arguing that a similar vulnerability in his code is the fault of absolutely everything (client code, any code that can be used as a deserialization gadget, "low quality tooling") except his code: https://bitbucket.org/snakeyaml/snakeyaml/issues/561/cve-202...
...
> 100% of the application which use SnakeYAML do not parse data from untrusted sources.
Wow, what a trainwreck of a thread. I don't know Jonathan but I would love to buy him a beer.
Thank you for sharing. ;)
Then later when someone says they are, he doubles down saying that parsing arbitrary YAML provided from people on the internet isn’t parsing “untrusted” YAML because they’re authenticated. Apparently it’s reasonable that anyone that signs up for a SaaS service gets code execution privileges on the prod infrastructure?
I’m pretty sure that thread had me shaking my head hard enough to give me a concussion.
if library have interface being
.log(format,value...)
.log(format)
It is kinda easy to assume second invocation would not apply formatting and just act as plain value, because it makes zero sense to have format string with no inputs, yet that's what Log4J does.Separating it for "log" and "log with formatting" calls helps a bunch, at the very least makes easier to search for potentially vulnerable code.
You have to evaluate to strings eventually though. A shell console/command prompt has no concept of SQL-style bind-variables. If you are logging "username: {}" that eventually has to be evaluated to a string sometime.
> It was a complete misfeature from start to finish.
No, there was a legitimate demand for this at the start. If you want to log instance-specific information (like JVM container name) the facility provided for that in Java is JNDI. And obviously it would be nice to have "what's the instance that produced this particular message" in the error logs, that's self-evidently helpful and desirable.
The problem is that JNDI is a foot-howitzer, it can also be used to load and execute arbitrary code from remote locations among many other things. Similar to how binary deserialization from remote sources basically cannot be used safely in Java (because a deserialized object can be in an "illegal" logical state and bypasses constructors and other assertions, an object type cannot be known before it is deserialized (so it can be of any class-type), and deserializers can run arbitrary attacker-controlled code before it hits any validation layer) this is not immediately apparent just from the fact that you are using JNDI, but, it's safe because we're not going to call attacker-controlled remote code through JNDI, right? Why would we attack ourselves?
And then the final piece is that (as you note) they fucked up the order of evaluation. JNDI is evaluated at the very end of the stringification process, meaning an attacker can include their own JNDI calls in the string and then it will be evaluated like developer-controlled JNDI. That is the fundamental problem here - this is bog-standard SQL-injection-attack style stuff, you should never be able to escape an injected/bound variable, but JNDI evaluation takes place after stringification.
All that really needed to be fixed was the last part. But honestly the dogma of the Java world has moved away from this whole approach in the first place - people were doing multi-container-per-JVM back in 2000s when memory was scarce, so you wanted to know that instance information so you could tell what instance was generating a message. But nowadays people do 1-container-per-JVM (and many JVMs scaled in a cluster) and you can handle that with runtime variables/etc that let the instance know where it is from external oracles. So there is no need to touch that in the first place.
It's all very much functionality that was demanded by real-world users... in the early 2000s. But Java refuses to clear some of the cruft from their APIs because it would produce (gasp) breaking changes in legacy applications... sort of the exact opposite of the Python 3 mindset. And JNDI is one of those interfaces that was just way too powerful when it was created. As mentioned, binary deserialization is another... you can't really use that in almost any scenario because the attack surface is just too large.
https://cheatsheetseries.owasp.org/cheatsheets/Deserializati...
(JNDI plus deserialization is a good one too, up until a couple years ago you could specify the class/deserializer as living on a attacker-controlled remote source, because JNDI can do anything!)
https://www.veracode.com/blog/research/exploiting-jndi-injec...
https://xeraa.net/blog/2021_mitigate-log4j2-log4shell-elasti...
Log4j 1.x didn't have the log4shell vulnerability. Log4j 2.x did. Those who stayed behind on 1.x never had the log4shell issue, even though the log4shell issue raised some other new vulnerability issues on the old code.
Worse yet, when that happened, the Apache project wouldn't reopen the 1.x branch to allow for fixes, necessitating a fork called reload4j. What could have been an easy version update overriding old versions became more troublesome. Now one must add exclusions for the old log4j to any dependency that has log4j 1.x transitively, and then declare the reload4j dependency to replace them. It is a very suboptimal solution which requires constant vigilance instead of a one and done fix.... all in an apparent effort to force dependency consumers to update to their new version which added major vulnerabilities.
Making something work in Java is easy. Proving it is secure is not.
It is almost certain that we over-reacted and shot ourselves in the foot with this, but our customers were happy with the way in which we responded. Better safe than sorry was the theme throughout.
Ok, so which language it is easy to prove secure?
Any use of Java will trigger additional diligence work to prove it does not use vulnerable log4j. So in this sense almost anything else is better.
But, we do have sufficient resources to justify and defend decisions made with our primary development language.
In this context, "prove" is more like "convince".
I've never experienced the problem of "nobody in my company knows Java well enough to trace the dependencies of a JAR". Do you all works in Rust and Javascript or something?