Log4j 2.16: Certain strings can cause infinite recursion
issues.apache.org
issues.apache.org
I read a lot of negative things about log4j these days, but people forget how great log4j was compared to whatever the JDK provided us with. It served a useful purpose for most developers, and inspired many libraries in other languages.
Take care!
From a quick glance at the comments this looks like a minor issue due to the attack vector being very, very small — i.e., the attacker must have access to where the logging pattern is defined, and if that is the case, this attack is probably not the most worrisome they could pull off.
I hope I'm not wrong, otherwise we'll be patching everything again.
Plenty of existing code does things like:
log.info(“foo: “ + request.getFoo());
Rather than using fixed format strings and {} place holders. You’re not supposed to, but it’s far from uncommon.Why would you still do be doing that work? If a value goes into a no-op, then the value isn't computed. (In theory - I'm sure it doesn't always work out 100% of the time in practice.)
Are you saying that Java does or should "optimize" in such a way that it branches to the logic within each method called during runtime, inspects which code would be executed per possibly mutable runtime values (like log level), considers possible side effects (in-method and in parameters) then decides whether to even invoke the method at all?
One of the bigger wins with the JVM is assuming that a virtual method can be called statically if there are no derived classes. That's a huge win for every leaf class in the class tree. The assumption can change every time a class is loaded of course. The optimization is called devirtualization and it can be combined with inlining to get even bigger wins.
Should the JIT call incrementAndReturnFooCount if debug logs are disabled? This is a long-recognized pitfall of C preprocessor macros that look like function calls but may simply be defined away, causing unexpected behavior.
Did I fall into an alternate dimension where 90% of Java seems to only use preexisting objects now, and people hardly ever write their own code?
(No disrespect to the asker; it's just such a surprising question to stumble across, that mentally speaking, I had to deoptomize my mental model of the world to accommodate for people who may not have worked much with Java).
I get functional is all the rage, but wow. Default toString() impl is inherited from Object, and basically gives you class and instance number.
The whole point of a JIT is that it can make assumptions. 'Assumption' is literally the term used in the JIT to track things like this.
If logging was part of the language, one could simply rule by fiat that arguments shall not be evaluated unless logging is enabled, but log4j is just another collection of user-defined classes, and gets no special treatment.
I'm not saying that when it comes to the JVM's optimization behaviors that I'd jump without checking that there's water on Chris's assurance, but I'm not saying I wouldn't either.
To be clear, I am not suggesting that these sort of expressions are inherently insecure - that would depend, I think, on whether they involve user input in a way that allows an attacker to take control of the call. A hard-coded JNDI lookup from the arguments list of a log4j call might be inefficient, but no less secure than if made anywhere else.
Assuming a global mutable variable will never change is generally not a safe assumption.
I suppose, at least in principle, that something like the GraalVM AOT compiler has the option to scan all of the code to verify that it never is mutated. But HotSpot cannot, because it only JITs code as it is loaded.
Put your logging level behind a switch point and C2 will treat it as a constant but still let it be changed - that works today.
For example, an interface implemented by only 1 class might get the class inlined. If a second implementer pops up (which can happen at runtime for e.g. some dynamically generated class), all of this will get undone.
There was a series of small articles with all these things, but I can't seem to find them right now.
Java HotSpot can make and verify that assumption. And switch code when the level changes.
If (a) DoExpensiveStuff());
Becomes
When modifyA() RecompileIfWithNewAValue();
As long as you call the initial if more often than you modify (a), you’re fine (and the JVM was able to see that you called your if 10k times without modifying the value of (a) even once)
log.debug(xmldoc)
The debug method took a string, and Java was converting the xmldoc to a string. This was back in the 1.4/1.5 days.
log.Print(fixed_string)
log.Printf(format_string, args...)
Simple, unambiguous.If it were a special type then the compiler could do it.
So if this is true:
> In 2.16.0: if the suspect string is put in the PatternLayout, then that specific patternLayout will crash when loaded and replace itself with a PatternLayout that just logs what is handed to it with no formatting
> logging the suspect string seems to have no affect, and it is passed untransformed to both System.Out or the file I specified as expected.
Then not a big deal. I'm sure there's a number of ways I could mess up the config file and end up with no formatting. Should I be submitting CVE's for those too?
12/18 - https://nvd.nist.gov/vuln/detail/CVE-2021-45105 Score: -
12/14 - https://nvd.nist.gov/vuln/detail/CVE-2021-45046 Score: 3.7
12/14 - https://nvd.nist.gov/vuln/detail/CVE-2021-4104 Score: 8.1
12/10 - https://nvd.nist.gov/vuln/detail/CVE-2021-44228 Score: 10.0
And the new CVE-2021-45105 is self-assessed to have a CVSS of 7.5/10 (see the same page above).
It's goods new so far that with more people/time and attention paid to the log4j exloits, the vulnerabilities are just getting narrower in scope and lesser in impact.
If you shine a thousand spotlights at a problem, you'll find more problems.
Glad this library is finally getting the code review it deserves. Hope the whole SDK gets the fine-tooth comb treatment it is overdue.
Like it or not... 50% of the Internet runs on Java (and my statistics are 50% accurate. I swear, 50% of the time.)
Polished, masterful, crafted product is hard to KPI for.
Highly unlikely. Most of the internet runs on PHP, WordPress more often than not (there are stats on that, but I'm on mobile and can't check right now).
Wow, a LOGGER engine that execute arbitrary code. wow
Ironically, tho, my tool to scan all our jar files for log4j has revealed a panic in archive/zip, something to do with a zero length file name.
The last time I benchmarked java.util.logging though, it lost out to log4j by a wide enough margin. Has anyone done any benchmarking lately?
Who needs decades of battle tested, proven methods and tools? Who needs thousands of human hours behind thousands of GitHub issues and pull requests?
Let’s do everything from scratch!
Hm, why is there JavaScript fatigue and 10 new frameworks every month?
I'm not a Java developer, so I don't know, but is there something "wrong" with java.util.logging?
Our field is royally screwed, whole thing needs to be rethought.
Any such field is in varying degrees of messiness (medicine, agriculture, architecture, you name it). Mistakes are made constantly, have real consequences (people die), and hopefully are discovered and admitted sooner rather than later.
All we can is (cliché warning) do our best, keep growing, act responsibly and be humble.
It's quite depressing.
Sometimes I wish I had more time to rewrite all Java stack, from logging to http server with dumb simple inefficient but understandable code.
May be I should just move to Go. It seems to better reflect that approach. But I like Java language...
At one point, it seems that "everyone use this so it must be secure enough" replaced "we're a large company, did we spend enough time reviewing code of the open source stuff we use?".
It seems the Linus's quote "given enough eyeballs, all bugs are shallow", is not really true.
Imagine if a well funded agency like the NSA employed at least hundreds of full time developers whose job would be to sniff for those vulns. I'm pretty sure you could automate searching for those vulns, and that only the NSA has such tool.
The problem is, there are a vanishingly small amount of engineers talented enough to find these kinds of things.
And if NSA wanted to hire a few hundred they’d have to go the defense contractor route which will inevitably lead to them getting 1 great engineer and 499 extremely mediocre ones.
It is about the willingness to invest, not brainpower. If I was tasked to "improve security" at my current project, I could instantly tell you things that could be looked into, even though reasonably precautions are already taken and there are already processes around it in place. I bet it is the same for most programmers.
Given that you're mentioning fuzzers and a decent source of inputs you could use, I think you're either underestimating yourself or overestimating the actual "average" of software developers.
Well that's the issue isn't it. Very few have read the log4j code, even though thousands of developers have incorporated it into they projects. Right now all eyes are on log4j and the bugs a showing up quicker than the log4j users can patch.
If you want to find issues in a code library, embed it somewhere in the hot code paths of a game or DRM system to maximize the number of eyeballs looking at every single instruction.
I think it is called Linus' Law but the quote is by ESR
I am 1000% in favor of apps managing their own logs. Generic system agents can collect and ship arbitrary logs remotely for storage/processing.
Putting /var on its own partition has always been recommended.
When default installer/partitioners fail to do so, I assume it's a habit picked up from desktop-oriented Linux distros.
Logging to /opt would clearly be a bug of the brown M&M variety -- it would make me question other decisions made by the developers.
I'd describe single-partition as "slightly simpler, but requiring additional support mitigations, and less effective at the primary job of keeping systems running".
I prefer:
One partition for / which is required to have an operable system for diagnosis of problems
One partition for /var which is expected to grow with limited predictability
One partition for the rest, which should be largely static depending on function, but sometimes surprises you regardless.
...
With elastic scaling compute nodes and centralized logging, most of these issues are subsumed into the enormous support infrastructure and can be ignored. But some applications don't map well to that environment.
And looking at the documentation, it has a very Java idea of rotation, where it supports every possible use case as a main case, and the examples expect that you will set it with the most insane defaults from day zero.
I decided to just ditch it and write from scratch something api-compatible, but extremely cut-down on "features".
Maybe someone would release something along this line. I can't.
LibreSSL is a complete rewrite of the openssl functionality with drastically fewer features. Same goes for CARP.
Maybe running software with minimal defaults is a good thing, as it forces the users of the system/library/whatever to think about its behaviour and usecase.
That's actually a great thing for the long term health of the project - it's getting a whole lot of free auditing right now. Many millions of dollars worth.
There's an old joke, goes something like: “Recently, I was asked if I was going to fire an employee who made a mistake that cost the company $600,000. No, I replied, I just spent $600,000 training him. Why would I want somebody to hire his experience?”. That's what's happening for log4j right now. A lot of eyes, a lot of attention, stuff's going to come out - and long-term it'll be better off for it. Shutting it down and moving to something else that hasn't been tested in battle isn't the pay-off it sounds like.