Logging as a code smell
dave.autonoma.ca
dave.autonoma.ca
(2) Not all kinds of software can afford to depend on an event-publishing infrastructure. One obvious example is the event-publishing infrastructure itself. Less obviously, all of its transitive dependencies down to storage, networking, and operating systems.
(3) After dismissing logging and explaining a more complex event model, the author says in their own project he stripped away a lot of that infrastructure and ... reinvented logging. Similarly, they admit that AST-rewriting magic is a "footgun" and then recommends it. Why keep arguing for one thing and then doing the opposite?
(4) Tagging everything you don't like with buzzwords like "code smell" is a bit of a writing smell. It's "considered harmful" for millennials.
(5) The swipe at log4j seems to be just a misguided attempt to make the author's opinions seem more topical. The recent log4j vulnerability was not a problem with logging itself, but with a particular implementation that embedded some rather crazy functionality. That functionality could just as easily have been embedded in an event-oriented system, and I'll bet someone somewhere has already done exactly that.
Typically I only resort to logging in situations where I don't want to propagate an error but I do want to log that an error happened. I also sometimes get lazy and debug using logs. I'm trying to avoid this as much as possible and instead debug by attaching to a process and using breakpoints. However this article doesn't touch on either of these scenarios.
Also, log4j's vulnerability had nothing to do with the actual concept of logging. AFAIK the vulnerability came from log4j's JNDI lookup functionality, which was not doing any kind of input sanitization which allowed attackers to perform remote code execution.
I was really looking forward to a post on logging as a code smell, mostly geared towards inexperienced developers. Pretty disappointed with the actual post.
Not a problem but rather a fact of life.
This is an organisational weakness not to have figured out a way to make the people responsible for production, also responsible for reproduction and bug fixing.
I've seen companies with a full team of developers who never add new features but only reproduce and fix the errors of the other feature teams, companies with read-only access to obfuscated informations good enough to reproduce without any client reference, companies who localize enough people in each jurisdiction the production is regulated in rather than have everyone out and unable to reach it.
It's an organisational problem, fix it instead of saying it's impossible to be creative with it. And if my experience if anything to go by, regulators and auditors are either misunderstood in wayyy too conservative ways or are themselves misunderstanding in conservative ways their own rules: there always are ways to explain that it's more risky not to maintain production rather than let if rot slowly because "I can't ssh to the box".
While it's always easier to be conservative with vague rules, it's also always possible to be clear faced with absurdity and it's rare people embrace the absurdity: usually they tell you: "I'll lose my job if I don't follow the rule X put in place 20 years ago" and not "I want the production to be offline the entire day during heavy traffic because I think clients prefer when they can't use our software". So, easy: "20 years ago, X. made a mistake - let's show him", and the answer from X is almost always "Why the fuck are you even asking me, ofc exceptions are possible, fix now we do the process after". And "doing the process" is often to can the old rule, or make it change so much it loses all fear factor and start being actually useful.
Give up and rot, give feedback and soar :)
Would you let a random coder come in an attach their laptop to your network and start digging around for a problem in their software? Especially if it's your ass on the line if any regulated data is seen by the wrong pair of eyes?
Or would you just ship them the logs from their own software and tell them to figure it out.
No. But I would a proper Engineer. The same way I don’t go seek treatment in a mall, but rather in a legitimate medical clinic by real MDs.
Does this clarifies things?
Figure out how to write a control and prove it to the auditors that it's followed and adheres to both of these, then I guess you're golden. Probably easier said than done.
FFIEC audit handbook excerpts for "Segregation of Duties" [1] and "Principle of least privilege" [2]
[1] https://ithandbook.ffiec.gov/it-booklets/information-securit...
[2] https://ithandbook.ffiec.gov/it-booklets/information-securit...
It's easier to tell the customers to send the logs from the relevant timeframe and narrow down the problem. Maybe instruct them to up the log level a bit and see if the problem surfaces again.
Always log enough to see where the bug is just by looking at the logs. Getting permission to deploy a new version with additional logging might be a bit of a pain.
> Pretty disappointed with the actual post.
I am as well. What the author describes isn't new or novel, it's been in Windows in the form of ETW since the early 90's. And Windows wasn't the first OS to get it. Had the author taken a serious System/OS class he would have known.
You can't build high performance applications without using some scheme similar to what the author describe.
>Had the author taken a serious System/OS class he would have known.
I'm guessing the author probably did so - this comes across as fairly ascerbic btw.
While I do agree with the need to simplify code, moving to an event bus transfers the complexity to where we don’t necessarily want it - to system level. Now there’s one more moving thing to be deployed, monitored, scaled, networked. And, by the way, that moving thing is mission critical, as without it you are flying blind, including in terms of infosec. Is it worth it? Maybe in some scenarios. But definitely not to the level of a code smell, that should be something to be avoided universally.
What I would rather see is a new keyword that generalizes the concept of logging in it’s entirety — “emit $object.”
For discrete events it would be emit ThingHappened(), for logs it would be emit Log(), and for traces it would be emit OpenSpan() and emit CloseSpan(). You could collapse whole classes of libraries into a few emit calls and a single listener.
And at the end of the day, few things are more versatile than a big wall of text with timestamps and request ids attached. Saying that logging is a code smell is a little like saying all tech debt is bad.
I go even lower though. Just
print(“XXXXXXXXXXX”)
print(“XXXXXXXXXXX”)
print(amount)
So is a sewer though, and I wouldn’t call that an event bus either.
Yes, you still want logs (or events or traces, etc) in production to identify when edge cases or heisenbugs have occurred. A debugger doesn't replace those but it does replace 'println` style debugging when developing new code.
Frankly the best developers I have met did not rely on debuggers, or logs for that matter. Using neither forces them to keep the program flow in their head and completely understand the code and system before pressing run.
For debugging production systems of course logs are useful.
It's a shame the author didn't actually look at logging as a code smell. For example, is a high density of logging statements in code an indication of places which are brittle and hard to figure out and therefore candidates for reworking? That would be an interesting topic I think.
> eases the transition to internationalized log messages
You don't internationalize debug traces any more than you internationalize your function names or programming language words. You barely even fix spelling mistakes in trace messages. Often, you build software without them at all: "compile them out".
One excellent reason not to use events for trace messages is that you may sometimes want every last trace message leading up to a crash to be recorded, which won't happen if they go through an event system where they sit in a queue which is no longer being serviced since everything died.
Oh, and your shiny event system could have issues, debugging which could benefit from trace messages.
The reasons in the post did not convince me on why my log messages should be structured. For example, why would I want to internationalize those messages?
Is it? What am I missing here?
(1) Some code that does a thing
(2) Another line of code that logs that the thing happened
And, like in-lined documentation, the code can diverge from the text that says what it does.
userRoles.addMember( p ); RoleAssignedEvent.publish( p, roleName );
Seems to have the exact same risk of divergence.
Depending on the language you can (and I have) used decorators/higher order functions to log the invocation of a function with its name and arguments. Zero risk of divergence there, but it was still logging not some kind of event system. The problem of divergence seems to me to be orthogonal.
(1) Some code that does a thing
(2) Another line of code that sends an event that the thing happened
TL;DR: Log impure actions. If you minimize and group side effects, you don't need to log as much. Pure functions can be recomputed any time if you know their inputs.
Event buses can be useful, but it's hard to look at an event being published and know what it does and whether it's important. Whereas the importance of a logger call is usually pretty clear. Logs are rarely load-bearing, but an event can mean anything.
vs
> NodeMarkingInactiveEvent.publish( localPause );
The second statement is totally not logging because it doesn't have the word "logging" in it. ;-)
That being said... conceptually, I do like the generality of publishing an event that describes internal system state, rather than logging a string on some level. OTOH the logging is supposed to be human, rather than machine-readable. But it's also gotta be a heap more work to do it this way and, maybe you'll never know if the error occurred cause the service lost network for some period of time (because it can't write to the event bus either!)
Wish there was effort to establish a pattern for this in C.
Replies extremely heavy on POSIX api and threads. All the cool functionality is with sockets. There is an inproc:// but it doesn’t do much.
Lighter to just work it out using RTOS objects like queues and streams.
Wrap it in a class of your own so the logger (or whatever) can be easily swapped out if you want to use a different logging system, events, or whatever.
In Java that would be BTrace.