Third High Severity CVE in Log4j Is Published
logging.apache.org
logging.apache.org
We won't hold a grudge against you; open source means collaboration, and you don't blame hard-working people that give away their software for free for honest mistakes. Your software still did far more good than bad to the software world.
Thanks!
This wasn't just a predictable scenario, it was predicted. Or more accurately, it has occurred already repeatedly in the NPM ecosystem, but for some mysterious reason those incidents were simply ignored by security teams world wide.
Instead of chilling them to the bone, they simply shrugged their shoulders and said "Well, we don't use NPM... I think. Probably?" and went on with their paper-pushing or whatever it is CISOs do these days.
If you think Log4j is bad, wait until Rust gets popular and has something similar happens with a commonly used crate. How would you scan for a vulnerability in compiled code that doesn't even have separate module files?
I had to help an organisation with log4j that had on the order of 5K distinct executables/applications across 3K servers on two clouds and three on-prem networks. At least with log4j it's a simple matter of finding JAR files and scanning their contents. They're just zip files! With languages that output a single binary by default such as Rust and Go, we would have been screwed. No way to scan, no way to self-help update.
Introspection and post-shipping updates for security are mandatory. Rust and Go like to pretend they aren't, because they originate from organisations that are in total, end-to-end control of the software on their own network. Google famously uses a monorepo and can build everything they run from scratch in short order. The shortcuts they can take with their post-release management will never work for ordinary organisations. Never.
So let this be a lesson: Log4j was made for a language that at least allowed us all to find the issue and fix it ourselves. We have Sun's forward-thinking and the enterprise-friendliness of the Java ecosystem to thank for that.
> I lay any blame squarely at the feet of IT security of large organisations that were entirely unprepared to update a widely used dependency that wasn't an operating system or a runtime.
If there is the possibility to plan ahead from a sufficiently strong bargaining position, you will not end up in the situation you described. The seller should either continue to support it with a good SLA, or you get the source code with a turn-key build system if the seller discontinues its support. Or you switch out the software if the SLA ends. Should you not be able to agree on such terms, I don't think you would have enough influence to tell the seller to write their software in such a way that dependencies are easy to hotfix/upgrade.
Ok, after putting the unicorns herd to bed, an event happens and congratulations, you now own a license to PeopleSoft v.whatever source code. Oh yeah, you also got the management platform for your network provider too.
Now what?
There is typically no link back from a deployed system to its source, even if you compiled the binary within your organisation.
If it came from an external organisation, things are exponentially harder.
Modern deployment systems are largely "one-way", with no way to trigger a full recompile from the deployment end of things. If you have a VM with "SomeRandomBinary.exe" running on it that has a vulnerability in it, you don't have many options.
IMHO, typical IT as done in the wild suffers from this write-once attitude, where the chain of provenance is too easily broken.
Companies like Google have a very non-typical setup that can be copied only by very small orgs like startups. Typical enterprises look nothing like a startup or a FAANG.
IDK where this idea that no one can know what's in a dependency tree is coming from but is not ops best practice for more than a decade.
I’d guess people running commercial non source available binaries in production. And outside startups and unicorns, that’s almost everybody.
How do you find the dependancy tree for your on prem Oracle db, or your self hosted Atlassion stuff, or your non cloud ServiceNow or PeopleSoft stuff, or your Huawei network management stuff, or or or…
And what can you do for any of your cloud hosted saas stuff, beyond telling legal “I dunno, their website hasn’t said whether they’re vulnerable to exploit de jour yet”?
Depending on the product, maybe forever.
I have some bad news for you: viruses and hackers don't care about your support contracts and the delays they cause.
Actually, I tell a lie: the hackers love them.
a) It's very rare to patch things manually
b) It's very rare to develop patches yourself (for 3rd party software)
I'm sure some companies out there with many decades old deployments have to do stupid shit like that, but it's hardly the standard.
If it’s some important piece of software you also can’t stop or the business stops - well, you and everyone else is just going to take it in the pants.
Many enterprises are in exactly this bind. And support contracts only get you so far.
Many people don’t even realize how long it can take to get a real issue fixed even with a super sweet support contract. For instance, I once was a wet behind the ears Oracle DBA, and when creating a schema i dumbly used a feature in the documentation without checking to see how long it had been released. Unfortunately, it had been released about 6 months before, as I found out later.
It took 6 months for them to figure out that it was the cause of our sporadic ORA600 errors (aka core dumps), by which point we’d already migrated to MySQL for many of the workloads because the database was having at best 95% reliability because of it.
Once we migrated the table off with the feature, somehow our problems ceased, but that was still well before they told us what was going on.
Moreover he is arguing that java, as opposed to fat binaries with potentially stripped symbols, made things solvable. He fears that many companies would be much worse off with a big vulnerability in e.g. a rust crate.
I don't think we should optimize for garbage companies with garbage practices.
I think we should optimize security for say the 80th percentile company as far as not following best practices (i.e. 80% of companies have better practices than what I suggest targeting).
The vast majority of cases, and this has been the case for decades, will be that you take responsibility for your own dependencies, and you let your vendors take responsibility for theirs.
> The go command now embeds version control information in binaries including the currently checked-out revision and a flag indicating whether edited or untracked files are present.... Additionally, the go command embeds information about the build including build and tool tags (set with -tags), compiler, assembler, and linker flags (like -gcflags), whether cgo was enabled, and if it was, the values of the cgo environment variables (like CGO_CFLAGS).
https://go.googlesource.com/proposal/+/master/design/draft-v...
At the time that Microsoft started doing it, paranoiacs thought it was some kind of anti-piracy software fingerprinting mechanism, embedding the compiling computer's MAC address or something. But really, it serves exactly the same purpose that Golang's effort does here: to let you map compiled-artifact back to inputs.
Yes...? What's the problem?
> why they don't instead hash all the source files and then embed that hash.
This is almost exactly what the git sha is. If you're arguing that README.md files shouldn't be included in the checksum, that's subjective. Many people would argue otherwise. And it doesn't hurt to be more accurate than less.
Do you not have this? Our docker images are tagged with the git hash they were built from, so at any point, for any of our envs, I can pull up the lock file of that build.
Our deployment config also describes everything that is running the relevant env, and nobody has machine access- going through the pipeline is the only way to get access, so there’s nothing deployed that’s a mystery.
Imagine you are Mr SecOps guy, and you've just ran some sort of Log4j tool across literally three thousand servers. Of those, several hundred came back positive.
Those included about a dozen flavours of Linux, a smattering of manually built(!) containers, and every version of Windows from 2008 R1 to 2022. Most of the code was built by third parties, some under support contract, some not. Most was built and installed manually, with developers using RDP or SSH to edit config files an whatnot directly on servers.[1]
So what you have now is literally just a string to a path, something like "D:\apps\foo\bar\baz\libs\stuff\thingie\log4j-core.jar" or the Linux equivalent.
Now what?
No, seriously, now what do you do? You're in SEC OPS. Not dev ops. You're certainly not in the dev team with access to the Git repo of some random vendor product like Tableau, or JIRA, or whatever[2]. You didn't deploy it. It got installed by a contractor during a short-term project three years ago.
A random hash string is totally useless to you. Even a repo URL and the commit hash will more than likely just end with an "Access Denied" URL, assuming you even have a network route to the Super Secure SCM Server.
There is no way you can figure out who needs to do what to make this go away. Not at this kind of scale at any rate. Not without first-class automation for literally everything. Which you can't have, because third-party software just doesn't play nice with any one tooling you'd like to use.
Containerisation? Bahaha... haha... snort. You're dealing with vendors that literally advertise "now with 64-bit support" and are unable to comprehend the concept of unattended command line installers. Vendors that insist on USB dongles for licensing. License that expire. Annually. And are tied to CPUID values. And on, and on.
[1] Oh, you think you can dictate release methodologies to these people? They're bureaucrats and they play politics better than you. Any word of changing their workflow in any way will immediately bring their boss, their bosses' boss, and maybe a few more levels up down upon your lowly head. You will have people literally screaming at you that your fanciful notions of build pipelines is "too much" and would "impact the work". That's the end of the conversion. I said THE END, and good day sir.
[2] Just get them to update it under their support contract? Ha-ha. Ha. Haaaa... We had vendors straight up lie about the vulnerability of their software. Then another vendor said that updating the JVM already mitigates the issues and hence they're not going to release an update. (Narrator: JVM updates aren't sufficient.)
Goodbye and thanks for all the fish. Run!
Now what?
The remediation (until you can get an update from the vendor) is to remove the JndiLookup.class file from that jar. It's been fairly well publicised, as has the way to do it.
I can't see how this is going to work as a strategy. Ad hoc cleverness? Sure, and that's always good to have when you're smart and lucky enough to pull it off. But it'd be nice to have a strategy that works without relying on being smart, willing and able to engage with tricky ad hoc solutions, and somewhat lucky.
Your company has an inventory of hosts to teams which you can look up and contact. In addition the package manager tracks which package the file belongs to. Therefore you raise a ticket with the team from the inventory tracker that package XYZ on host ABC contains vulnerable file /x/y/z/log4j-xyz.jar (The reverse is also true, for each package we can find the corresponding repo and set of hosts its deployed on). Fixing it is now that team's problem, and secs ops guy's problem is now just to verify when they claim to have fixed it.
If the software is JIRA? Doesn't matter, still needs a package built before it gets deployed on our machine. "No really, the vendor's only supported installation method is sudo curl | bash" - into a container it goes, the owning team of the container can be identified much the same way as the owning team of the host.
Now the problem you find is sometimes the result is that the host belongs to team XYZ and you find the team was reorged by some exec's great idea in 2019 and only one guy who used to be on the team is still in the company, but he left the team in 2018 and has no idea what that team did since he left, but the application has been still runnning and power $millions of real revenue. That one is much harder to fix by automation. The business reality is the security team can't enforce to the execs to not lay anyone off or disband any teams without a concrete transition plan for their systems.
ELF binaries are just collections of another sort. You can scan for symbol names or assembled instructions common to a dependency, and you might be able to update with an LD_PRELOAD library that patches the symbol table.
Just look at what game modders have accomplished without access to source code.
> Just look at what game modders have accomplished without access to source code.
Game modding is usually done on Windows given the target market (up until recently of course) and thus usually means Windows PE's, which do not have debug symbols attached - ever. Windows debug symbol databases are emitted as PDB files and are almost always omitted from game releases.
This means modding has to do a sigscan[1] in most cases in order to find something interesting - especially in the case of ASLR[2]. Then whatever modding framework (or hack) can set itself up and hook into the game.
These techniques are NOT what we should be encouraging, and are certainly not common, even for vulnerability mitigation. Re-compilation and re-deployment should be the defacto mode of operation in production, especially for systems that have sensitive information. Relying on bin-patching is not something many security specialists would regard as a "good" mitigation.
[0] https://en.wikipedia.org/wiki/Strip_%28Unix%29
[1] https://wiki.alliedmods.net/Signature_Scanning
[2] https://en.wikipedia.org/wiki/Address_space_layout_randomiza...
Imagine the vulnerability would have been in the vendor code instead of in the log4j jar. You’d. still need a way to fix and rebuild it.
Note that this doesn’t have to mean a public MIT repo on GitHub. You can have more restrictive gray box licenses between your b2b partners only.
What are you doing about it?
With log4j you have a brutal combination of:
1. RCE
2. Exposure
RCE happens frequently but often attackers don't have an easy time getting to the exploitable code. With log4j it's trivial - every app can be owned.
But here are some questions:
1. Why do those apps have the ability to make network requests?
2. For the apps that need to make those requests internally, why aren't those over mTLS?
3. For the apps that need to make them externally, why isn't that going through egress proxying?
4. Why are there credentials spewed all over your environment variables across your services? Are they short lived?
It's actually not super hard to have a worst-case scenario vulnerability like log4j be not that bad for your organization. A bit of hardening and even if something like this happens you're in a good position to wait, monitor, and patch.
not just for security either. i can't count the number of times i've traced a performance problem or bug to a piece of code that was doing some io when it didn't need to
That's a bummer, but I suspect many on HN can in fact build software that isn't total garbage.
> What these applications are supposed to connect to are likely unknown. This is a major problem at enterprises
Only if they fail to consolidate vendor integrations and/or have already insufficient staff. (Which is often indeed a problem.)
Simple, if brutal solution; Infosec audits logs and works with application teams to confirm existing traffic and pre-document new traffic out. With properly staffed teams you should run out of surprises within a year.
> Simple, if brutal solution ... properly staffed teams
The recruitment part is not simple, maybe bordering to impossible? for some larger companies. I say, based on my past experiences in how clueless big companies can be concerning what technical people they hire.
Depends on how you game OKRs. Implement auditing Q1. Contact all teams Q2. Have 75% of teams Q3. 90% of teams Q4. Just an example but its all in how you sell it. Big IT understands icebergs, you just have to give a good/accurate roadmap and execute if they agree on it.
You might be correct that a lack of insight on this sort of thing is an organizational problem. My experience indicates that if you can't get a list of dependencies based on a series of e-mails down an org chart, something is wrong. Either someone is managing too many services for the level of tooling you have, or there's a complete breakdown in org comms (which happens more in larger enterprises.)
I find it quite ironic to praise the forward-thinking of the company that has been instrumental to bring us into this mess via their vision of loading dependencies at runtime from an online repository. Sun imagined that the code and its configuration does not have to worry about how to fulfill its dependencies, but instead JNDI [1] could magically fetch the appropriate objects from wherever, don't worry.
I'd say I would even find it ironic to praise the Java ecosystem, which, due to its excessive over-engineering, has produced best practice frameworks where fully documented behavior has the latent potential to just be catastrophically exploitable by accident, as nobody is able to reasonably understand how all of it plays together. I have a hard time to imagine how you could enable loading of external code in Rust or Go by accident using run-of-the-mill logging frameworks.
[1] https://en.wikipedia.org/wiki/Java_Naming_and_Directory_Inte...
Designs like JNDI were widespread in the IT software industry at the time.
The academic prototype for all this was CORBA, the Common Object Request Broker Architecture. An object-oriented was to transmit data between systems running on separate processes, CPUs, remote servers.
CORBA was a next-gen RPC (Remote Procedure Call) framework, the data format on the wire closely follows the format for C structs using Sun-RPC, but was way better because these CORBA things were discoverable and described at run-time, rather than compile-time libraries.
That's right: RPC was initially a Sun UNIX thing.
Microsoft needed a way to do this, so they had COM, the Component Object Model, for run-time discovery and linking (late binding).. Microsoft got COM to reach across network connections, "DCOM", for Distributed systems.
Meanwhile, corporate IT was still pulling text out of databases and shoving it into C structs. It sucked.
Microsoft's object model was better, and Java was better than better, because you didn't have to license it to use it. Well, you did; it was like an MIT License.
Reason why IBM writes huge Java systems. And why Oracle owns Java. For some value of ownership.
Yay objects!
In these Java attacks the code is downloaded and run locally, with all secrets and data access of the attacked program.
It’s funny because when I worked in .gov IT, we were all characterized as stupid donkeys by because we weren’t able to just use NPM, etc to do cool kid stuff.
Good luck getting security people with the power and brass balls to do that in any company.
This works because the binary is compiled against the dynamic library, using the system-provided library. It isn't statically compiled, because even though that frees the developers from worrying about library versions, it makes the sysadmin's job tracking library versions be much harder. It isn't vendored, for the same reason.
Sometimes you can get away with just depending, and that’s nice, but other times there are strong reasons to embed. At that point, you own the code, and can actually do the things you need to do. (Of course, it follows that you own the code - and need to then actually do the things you need to do.)
Both approaches have their advantages, and both have potential for technical debt. Which approach offers the most advantage for the least debt will depend on the individual circumstance; I don’t feel that you can really advocate for one over the other in a general sense.
Our software scans take 3 forms:
1. Searching for file names
2. Running test exploits that have a payload that just logs the vulnerability with a central monitoring system
3. Dependency resolution and checks against a DB (our solution _already_ supports Cargo.lock files.
Rust (and Go, C++, etc.) only prevent option 1. Thanks to bundlers it also can miss it in JS, or shaded fatjars it can also fail kn Java too. So it's not a new challenge
I think this attitude is one of the main issues for this kind of incidents in larger organisations. I'm not sure why you would be willing to lay blame at your colleagues for this. Most of the time, development teams don't tend to like some governance over the code and the dependencies they are pulling in. Not just because they know of course what they are doing, but also because they are under pressure to deliver features. That is what matters for business. Also, convincing management that budget is needed for correct tooling to track all stuff deployed; is also not as straight forward as you seem to suggest.
In this case, the problem goes even beyond just the code of your own dev teams. This is embedded in countless software packages deployed all over your organisation. Same here, people want to buy and use whatever they want. And all processes to keep some form of control over it, are mostly seen as overhead.
And I'm sure they are lots of "security" people who are just producing documents and policies which are complete detached from reality. But developers who consider security completely as somebody else his responsibility, are a problem as well.
The only good think I see coming from this mess, is that security teams probably will get the means to try to get more control and insight over this. For the coming weeks/months at least. After that, everyone in management of dev will be forgotten about it. But something tells me the security people who are working on this right now, won't.
no, best practices include checking in your Cargo lock file
Githubs of the world could just gate downloads, pull requests etc. behind a payment to see what is the real valuation of open source software; I imagine it'd mostly settle around $0 excluding couple of big projects.
But just be be sure, we should ask legal!
If we can't find a way to pay Google, there is no hope for some random developer in Nebraska.
Security teams are stuck with securing the tire fire; they didn't choose the library or platform. If anyone should be advocating for supporting open source, it's the developers who benefit by using open source libraries.
Log4j2 partially exists because of pushback on adding features to v1
Honestly the biggest problem of package managers is finding a reputable package that one can trust.
The best one can do to aid this situation is reviewing, vetting and warning of which packages that can be trusted and not. Automatic scanners to find code smells and vulnerabilities.
“Security rating: 2/10, This package seems to use JNDI loading, are you sure you want to continue (y/n).”
Logging as it's understood today should ship with most STD libraries. Things like logback an log4j really should just eventually be rolled into the std.
---- This wasn’t a process failure, the vendor did everything right. Mozilla has a mature, world-class security team. They pioneered bug bounties, invest in memory safety, fuzzing and test coverage.
NSS was one of the very first projects included with oss-fuzz, it was officially supported since at least October 2014. Mozilla also fuzz NSS themselves with libFuzzer, and have contributed their own mutator collection and distilled coverage corpus. There is an extensive testsuite, and nightly ASAN builds.
I'm generally skeptical of static analysis, but this seems like a simple missing bounds check that should be easy to find. Coverity has been monitoring NSS since at least December 2008, and also appears to have failed to discover this.
Until 2015, Google Chrome used NSS, and maintained their own testsuite and fuzzing infrastructure independent of Mozilla. Today, Chrome platforms use BoringSSL, but the NSS port is still maintained.
Did Mozilla have good test coverage for the vulnerable areas? YES.
Did Mozilla/chrome/oss-fuzz have relevant inputs in their fuzz corpus? YES.
Is there a mutator capable of extending ASN1_ITEMs? YES.
Is this an intra-object overflow, or other form of corruption that ASAN would have difficulty detecting? NO, it's a textbook buffer overflow that ASAN can easily detect. ----
[1]: https://googleprojectzero.blogspot.com/2021/12/this-shouldnt...
I do not know about your security teams, but my security teams are now sitting 9 days straight trying to stop that shitshow. With teams complaining that they have to patch instead of doing Great Things.
After that these security teams will be back to invisible work and forgotten again. Until the next issue.
If your company has a truly different approach to security, tell it here - you may get really talented people knocking in because they are fed up with the security theater at their current job.
Bugs come in bunches. If you have fixed just one or two, more lurk right there. And, anyplace else that coder worked. The more you have found, the more remain to be found. Look at other places that coder worked that week. Or year.
But looking in your own code is probably next best. Ask yourself: if there is a bug somewhere in my code, where would it be? You may be surprised at how immediately the answer comes to mind. Look there.
A big subset of mistakes people make when writing software are the result of people not having the same understanding of a concept, policy or practice.
Even a good writer will make mistakes, otherwise we wouldn’t need proofreaders.
And that's a point in time audit. To maintain that value we'd have to redo the audit periodically.
It's just not gonna happen.
It was especially difficult for us because we’d shipped so much code that used the library, and replacing the library was unthinkable.
It's just not as simple as "security audit finds all the vulnerabilities, then you fix them." You invest X in the review, you get the results that X/(hourly rate) finds. This is a lot of software with a ton of configurability-- that's a lot of variations to review and test.
Now that someone found the first lump of gold and gave it away, there are thousands of eyes searching for the next one. These recent findings are all abuses of this same chain of functionality, just along different sets of settings. In another month we might have half-a dozen more of varying severity and scope. That _still_ won't prove that the overall library is then safe, but we will probably have a a little more confidence in this particular bit of crazy template formatting flexibility. Maybe not as much as we had had three weeks ago, but more than we do now.
https://issues.apache.org/jira/browse/LOG4J2-3230
It's pretty interesting technically, and it's not really affecting 2.16.0 much except in weird edge-cases.
Switching over merely involved changing the jar files we used, and redoing the config.
Hugs for all of you. This can't be fun.
Why are people still using JDNI? It is frail enough to have logs depend on a dns lookup, I can't imagine depending on a LDAP lookup. This is not such a fundamental issue as the above but it is pretty dangerous.
Oh yes and of course one of the worse and least pythonic package in the Python standard library was "inspired" by the log4j monstrosity
Because "java:comp/env" is the way for a J2EE web application to get values or objects from the configuration outside its container. Yes, nobody actually uses the remote parts of JNDI anymore, but the local variant (within the same process) is still in use.
And that's what led to this vulnerability: someone wanted to be able to get these local values within the logging configuration, which is a valid use case. Not realizing that this opened the door to remote values and objects.
But to be clear: nested/recursive template expansion and expansion of user provided strings were never by design, correct, and the removal shouldn’t just be by configuration or considered breaking - it should simply be corrected (just like the jndi should be removed and not even be optionally possible)?
The latter option represents a liability, which in this case (as with others) has shown can be a tremendous risk. Is the time saving really worth it, at the cost of risking disasters like this?
I get that the tradeoff it is worth it for complicated things (e.g. crypto libraries). But logging, really?
Software development culture today is too quick to adopt a huge tree of dependencies of unknown quality, rather than thinking about how to minimize dependencies to only those truly necessary. The leftpad fiasco was but an extreme example of this, but I see it all the time, and it seems probable that there are hundreds (maybe even thousands) of similarly severe problems out there in widely used dependencies that we just don't know about yet.
yea literally any library/dependency can introduce risk, including stuff coded internally for those purposes
There is also lots of stuff you might want a logging library to do that while not very hard would be annoying to do yourself for every project like log file rotation, switching log level while the application is running, optimizations (like in c/c++/rust you might compile for INFO level and drop all the DEBUG and TRACE level logging from the binary), different log levels for different parts of the app (including dependencies), etc
Transitive dependencies for one thing. I rather enjoy having my DB driver, connection pool, kafka client, and web framework all emit their own logs, which is not possible if I wrote my own logging library.
> I get that the tradeoff it is worth it for complicated things (e.g. crypto libraries). But logging, really?
Logging at scale is more complicated than people give it credit for. Personally I've had my services suffer noticeable performance degredations due to poor logging framework config. We've ended up with a fairly complicated logging config that allows us to carefully balance collecting as much developer useful information as possible against the need to drop excessive logs during periods of heavy activity. Oh, and it's gotta be converted into JSON for aggregation into our observability tools, and decorated appropriately so that our logs can be correlated with our application traces, because that is incredibly useful during incidents.
Oh, and it all needs to work no matter what threading configuration you throw at it. That matters a lot too.
Do we need all those things? Arguably no, but then again they are helpful tools to have. Without them, the team would have a harder time developing on our systems. You have to understand the use case that people need to fulfill before dismissing the entire thing as unnecessary and pointless.
> Software development culture today is too quick to adopt a huge tree of dependencies of unknown quality, rather than thinking about how to minimize dependencies to only those truly necessary. The leftpad fiasco was but an extreme example of this, but I see it all the time, and it seems probable that there are hundreds (maybe even thousands) of similarly severe problems out there in widely used dependencies that we just don't know about yet.
I see this argument a lot, and I think it's a great example of "grass is greener" syndrome. People don't think through the practical consequences of drastically reducing their library usage. I personally worked on a project that had basically no external libraries due to the language choice, and we wasted inordinate amounts of time trying to chase down bugs we wrote in things like our database library, localization library, and yes, our logging library. Looking back I can say with very strong confidence that what we were doing was a massive waste of the company's resources, and rewriting it into Java or C# would have been the right call, even with the complexities those ecosystems bring to the table.
Yes, using external libraries comes with tradeoffs. Fixing bugs like this is one unpleasant possibility. But let's be clear that the alternative is a massive drop in developer productivity to write basic functionality that already is implemented elsewhere. Having been on both sides of this equation, I can confidently say that you'll spend way less time fixing security issues as they arise than you will reinventing the metaphorical wheel.
(I do agree that leftpad is an example of going way too far with libraries, but I think the library culture for NPM and the Java ecosystem are quite different).
Logging facades like slf4j help a lot, but then you still need to implement the required interfaces. For something like the specific Log4J CVE, things would be significantly worse if it weren't centralized in a library. They have a bug in deserialization when talking to LDAP servers, allowing a compromised LDAP server to execute code remotely. It's likely that a subset of libraries would likely have rolled this functionality independently, and then you have many libraries that all need to be patched and updated instead of one.
In terms of features, one that comes to mind log rotation.
I dunno, I suppose all together I’d answer your third paragraph to the affirmative — I’m more productive and the libraries improve significantly within crucibles like this.
The larger issue is in javas dependency system are built with dependencies on these libraries, so if there is no commons logging (or bridge, ala slf4j), then the program won't run.
There are a few other techinical traps with the jvm, like using system.out logging can cause issues with threaded code, or having one giant static logger being imported all over the place conveniently (this is before inversion-of-control became the dominant paradigm) had a few shortcomings, especially as programs grew. Ultimately, though, if you were going to use some third party libraries or app servers, you were almost certainly going to get pulled into the commons-logging vortex, which means you're now configuring that AND your homespun logger.
There is also the case for things like SOLR, where it might be just being used out of the box, but the distribution includes the affected JARs, even if you're just posting and retrieving documents with a python script doing print-logging.
Honestly, I didn't even know log4j2 was a thing, I've moved as much over to slf4j/logback as I could ages ago because of the madness of the JDK logging ecosystem.
When Java finally added a logging library, they added one that no-one had ever heard of before, and which had fundamental flaws. That's why no-one uses it. At the time, log4j existed and was in wide use, but was not represented on the committee.
The author of log4j eventually decided to go a different way and created slf4j / logback. This offering is compatible with log4j and considerably simpler in many ways.
Having not been in the Java world for many years, I was surprised log4j was still in widespread use. But maybe I shouldn't be surprised - whenever we have to interact with Java teams, we always have to tamp down on the tendency towards complexity.
It's Java; the ecosystem is, politely, rather conservative and slow moving.
Note: In case someone is tempted to say that it is trivial to plug slf4j over logback, it is for greenfield work. Replacing an existing logging library, or the use of multiple logging libraries in favor of slf4j over one lib is a confusing matter, mostly because the error thrown at runtime are misleading.
(1) Coordinating with dependencies requires at a minimum a common API (which, in a statically typed language, means a library providing at least the interface).
(2) Configurability to log to different destinations and with different formats, which you eventually tend to want on major or long-lived projects, can be fiddly, as can the specifics of different formats and targets. Each of the bits is simple, but added up it's a lot over time, and reinventing the wheel, the axle, the tire, and the air pump for it just doesn't make a lot of sense. Especially given the need for a library for #1, it's best to get the core functionality together with it, and the peripheral pieces either together or as plugins, and only solving unique problems, if you have them.
Sir, you might have just extended all of our careers by at least a decade. We salute you.
I agree that developers are too quick to put themselves at the mercy of 3rd parties, but it's a risk-dollar trade off we'll continue to make because reinventing the wheel each time would sap our margins.
A lot of these libs separate logging into layouts and parameters. If you log input parameters as layouts. You might break the logging. I have seen this at every company I worked with.
The only really surprising thing about log4j is the remote code execution.
Log4J has been around 20 years. Log4J inspired and was lifted into the JDK. JDK logging is essentially an inspired copy of Log4J. The Sun coders were no better or worse than the Apache coders. I use JDK logging personally simply because it's one less dependency. JDK logging hasn't change since JDK 1.4, and is weaker than, well, pretty much anything else. But, with a simple wrapper I've used for 15 years, it does most everything I want from a logger.
That said.
Log4J has been, and is still, a boon to the Java community. Arguably, Log4J is the root of a tree of vast array of logging frameworks, across languages. Java server developers essentially live and die by their logs. It's routine for developers in dark rooms with screen lit faces to pouring through logs with endless stack traces. Thank heavens for Java stack traces.
Logging is part and parcel to the Java server side experience, and even the client side, and we can place much of that on the shoulders of giants like Log4J because it made logging easy and set the stage. It's so helpful, so useful, so flexible (obviously, perhaps, a bit too flexible), and so powerful.
Because Log4J inspired the other loggers like the JDK Logger, we have logging shims. Shims like Commons Logging, that act as intermediaries that can have adapters written so that we can use other logging frameworks. Log4J was one of the first, was, and is still, dominant in the community, but it's not alone, and thus a bad choice for things like libraries. Instead, those choose the shims like Commons Logging that developers can use to configure to route through Log4J or JDK logging or any of the others. That said, even programs that use Log4J directly can be routed to other loggers.
This is all entrenched. It's part of the flavor of server side Java, configuring the different logging shims to write to your logger of choice on your system. Just the way it is. But, that's what happens when you don't live in a mono-culture. Feature, not a bug, and it helps empower the vast array of software that millions of developers and applications rely on everyday. Nobody designed it this way. It didn't start this way, it just evolved this way. It's a very "Java" thing.
I watch Stack Overflow questions and 90+% of the time when someone asks "How do I" what they mean is "What library do I need", not "How can I write this". This is the sign of the times, and Java is not alone. The beauty of Java is that it made this kind of sharing REALLY easy. REALLY REALLY easy.
Apache has a solid reputation for good projects and good code and good stewardship. It's not a back alley transaction to grab an Apache Java jar file and shove it in your project. Is it all perfect? No, but what is?
The problem with rolling your own is that, for a non trivial technology like a producionised logging framework or some encryption library, you're highly likely to write a massive security hole yourself. The great thing about battle-tested open source software, used by 1000s, is that years of real world use, development, bug fixes and general scrutiny from a large number of people makes for a high level of reliability normally.
Obviously here we have a problem with this approach but avoiding third party dependencies carries its own risks and obvious costs too. Let's focus on how reliable so many of these libraries seem to be and how rare problems of this scale are.
I think a bigger problem here is the fact that so many people are using so few libraries. Maybe we need more open source libraries to counter single points of failure. As with how we deploy our code on the same clouds and so on, single points of failure abound these days. Things fail however well they're engineered and therefore we need diversification of everything.
Rolling your own logging library to reimplement it is not trivial, even if you ignore the more esoteric output formats.
Sure, in retrospect they went too far with some of the dynamic logging features in log4j2, but it is no exaggeration to say that log4j revolutionized logging in the Java world when it first came out, and heavily shaped later solutions like logback and java.util.logging.
Not to mention a wide range of destinations, too. Want to log to a database table? You’re covered. What about log messages to a chat server (xmpp, slack) or email but only fatal errors? You’re covered.
Now roll your own logging system that supports those destinations safely and get back to me to do pen testing. I bet we’ll find some vulns in your code.
Most of us have log4j in our system because it's used by some various dependency of one type or another, and up until now it has generally just worked.
> It’s dynamically configurable
It's a logging library. You don't reconfigure logging, you just make another logger and use it. And by the way, you take the configurations from your global configuration system, "dynamic configurable" is a feature of the configuration system, not of the logging.
> supports a wide range of output formats and destinations
Yeah, like network aware templates. (But the destinations are a feature, multiple destinations are one of the few things a logging library should support.)
Despite what people claim, the Java culture of loving complexity never did go away.
Prior company had a shared plugin that let you turn up the logging for a given logger remotely for 30 minutes (it self reset). Extremely helpful for incident debugging.
Now we can have an interesting discussion about whether that capability is worth the complexity, but that’s a very different discussion than unilaterally declaring that other people’s use cases just don’t exist.
> Despite what people claim, the Java culture of loving complexity never did go away.
Do you develop in Java? Because I do, and your takes don’t match my daily experience one bit. My experience going from Spring 3 to Spring Boot 2 is in fact one of reducing complexity (goodbye XML bean configuration!) and to more “it just works” situations than before.
And you really believe this should be a feature of you logging library, and not of you configuration system?
> How much log data do you think cloudwatch and the other majors filter per hour?
How much do you think the major providers pay for that infrastructure per hour, and are you willing to shoulder that cost too?
(Hint: log ingestion is $0.50 per gigabyte in cloudwatch. So a terabyte of logs an hour is $512/hr)
Only if the emissions of said logs is doing more than stdout or fs writing and it’s not done in another thread.
> Log emission can often have a negative impact on performance, especially since peak logs and peak traffic tend to coincide. I’ve seen P95 latencies suffer just because of logs.
Yes and you need to fix that terabyte per hour garbage as that is well beyond typical for a single service. We’re also talking about java so 7000 page stack traces is where i’d start…
The only way to make a "reconfigurable" logging system is by having a stable facade that restarts everything behind it. But that extra layer does not belong on a logging system, there are plenty of other things you need to make "reconfigurable", and now you are adding an extra layer into each of them, instead of having only one on your main entry point.
Pushing that feature down into the logging system is the choice that maximizes complexity and minimizes functionality.
Also, I noticed you didn’t answer my question about whether you’re a Java developer.
I was once, but haven't used it professionally for a while. It's on my "I'd rather not" list, but not strongly so.
Oh, and now that you described it with so many details, the attachment to complexity does look more like addiction than love.
This is why I ask. Java has changed a lot since ~1.8, and it is super common to see ex or non java developers making grandiose and sweeping statements about a community and language that they are no longer part of. Having developed in Java on and off since ~1.7 and Spring 3, I can certainly say that I have seen a strong movement towards less complexity and more convention. It's not perfect by any stretch, but I think saying that the space is "in love with complexity" is incorrect.
You of course don't have to go and re-use Java on my suggestion. We all have our preferences, and that's ok! But I would generally recommend against such sweeping and dismissive general statements about a culture that you are no longer part of. Things change, and your experience back then is probably not representative of things today.
> Oh, and now that you described it with so many details, the attachment to complexity does look more like addiction than love.
There's a fine line between disagreement and insult, and you are over it.
That said, I do agree that there is a problem with large dependency trees. It seems building a simple hello-world app with Maven ends up downloading about 1/2 the internet. :)
Okay, more seriously... this is why I am currently staying away from Rust. Because sometimes I don't even know if the random number generation crate I imported for my banking app came from Rust team or some random dude in zanzibar.
(Sorry, I know HN loves Rust, so do I. But I currently that the cargo/crate system is a bit too easy to use and misuse).
Fortunately our log4j was too old and JVM too young to be vulnerable to the first exploit, but it seemed best to rip out log4j and replace it with something simpler anyway.
java.util.logging seemed the obvious choice, but the API is different, so would have been a lot of s/debug/fine and so on.
In the end, we:
- wrote a new class with the same 'api' as the part of log4j that we use
- globally replaced 'import org.apache.log4j.Logger' with our new class
- fixed any compile errors by adding new methods to our class
- wrote a test to make sure it rolled correctly and could handle lots of threads etc.
Here's the code if you want to do something similar (don't know how long it will be accessible from there): <https://ideone.com/XKg5M9>
Once you get to that stage there can be a follow up to gather more information, and that follow up then usually results in some advice or no further action. In rare cases - typically the ones where gross negligence or willful transgression of the rules was established - there will be a fine and if it is a repeat occurrence that fine can be quite substantial.
Also you don't have to report 'potential leaks', only actual leaks.
So I think that in the case of your hypothetical end-point the logging isn't a GDPR requirement, but if your endpoint ends up leaking data then the log can help you to establish if and if so how much data was exfiltrated. But that does not mean that the GDPR requires you to have logs, though, having access logs for your endpoints is a fairly standard thing and not having them is going to raise a few eyebrows, especially if you are also reporting a breach.
Where there are logging requirements: data retention laws, SOX, fintech, tax regulation, AML.
Then there's the user-agent header, which also can be pretty valuable to debugging and other operations. Although this one is easier not to log.
In my current logs, I'm seeing attempts to compromise log4j vulnerability in both URL query string and user-agent. (Which of course I would not know about if I weren't logging them). This particular application has no Java involved, so is not vulnerable.
The fundamental problem is that, pretty surprisingly, log4j was running the same parsing and logic on user-entered strings as they did format strings. That's essentially the root cause of all this - log4j shouldn't be attempting to parse this data at all.
So many severe vulnerabilities are due to the complexities of parsing potentially malicious user input - just look at all the severe iMessage bugs. While many other apps do have a fundamental requirement to parse this input, there is no reason a logging library needs to do it.
yes, but there can be (and have been) vulnerabilities in terminal emulators, terminal multiplexers, text editors, databases, web renderers and all manner of tools that can be used to store the contents of or view that file. your argument makes sense that the logging library should not be attempting to parse the log strings, but think about all the other millions of lines of code that could potentially try to parse that data as well. patch log4j one day and the next you find a bug in some colorizer javascript library or in the sixel support for the terminal emulator. people add parsing everywhere.
i think ideally that all user supplied data should be sanitized in the strictest possible manner before anything is done with it, including logging.
web browsers and all of their components get audited. because they process potentially malicious data as their core function, they see A LOT of attention in terms of hardening.
the point that i'm trying to make is: the amount of software that doesn't immediately and apparently touch potentially malicious data is absolutely enormous and the number of paths that data can take to touch that software is even more enormous. yes, one can put forth good principles for "building codes" for software and yes, one should... but there's an awful lot of software out there that is not up to code (where the craft is so young that the idea of a code isn't even close to being finalized) and it won't be for a very long time.
that being the actual on the ground situation, what is one simple thing that individual developers can do to help prevent catastrophes? always sanitize inputs. moreover, it's much easier to find and audit every place where data is accepted from untrusted sources than it is to find every place where data may be parsed.
Oh come on, how would you even do that if you don’t know how the data is going to be interpreted by parsers that nobody knows they’re using? WAFs have been trying to do this since forever, but it’s always been security theater that separates gullible companies from their money and randomly blocks users that use “special” characters like slash or quotation marks.
I guess you could base64 encode everything, but that wouldn’t save you from the terminal bugs you mentioned. At some point you’re gonna have to undo the “sanitizing” and look at the text.
We use a product that uses JCE. The log4j vulnerabilities worry me about the security of software in Java ecosystem in general.
Log4j is not part of the JDK and is an open source project with only few regular contributors. There is really no reason to infer any correlation to other parts of the Java ecosystem.
(Unless some hobbits have 3 breakfasts?)
No, it’s not.