Structured logs are the way to start
shippingbytes.com
shippingbytes.com
This works with any repeatable task and identifier, like runs of a cron job, and user ids.
You should still pseudonymize your PII wherever you can, but protection measures are required.
You should have a short term log retention period (i.e. not something like 2 years, unless you are legally required to store all the data for this period). If the log retention is too long, even if you can prove you need this data, you would still have to handle data erasure requests (a.k.a. "the right to be forgotten"). Setting up a way to delete all log entries for a certain user ID could be an alternative to a short retention period.
The other feature you want to have is encryption (both in-transit and at-rest). Encryption is not a hard requirement in GDPR, but almost no "technical" measure is. GDPR Article 32 mentions various technical measures (including encryption and pseudonymization) and requires controllers to implement them while "[t]aking into account the state of the art, the costs of implementation and the nature, scope, context and purposes of processing as well as the risk of varying likelihood and severity for the rights and freedoms of natural persons".
Here we have to do some interpretation, but generally the accepted interpretation[2] is: "state-of-the-art" refers to methods and techniques which are widely available. Cost of implementation is strongly tied to risk: high-cost technical measures are not required unless the risk is equally high.
Since encryption, pseudonymization are simple, cheap and have widely available open source implementations, it's a good idea to make sure everything is encrypted.
Restricting access control to log data is another measure that is trivial enough to implement, that it becomes a practical requirement. I've seen a lot of cases where the entire company or org had unrestricted access to logs containing PII in the past. This probably won't fly with GDPR.
Lastly, I would make sure the logging infrastructure and everything connected to it, follows industry standard security measures (at least OWASP Top 10). It sounds obvious, but I've seen a lot of cases where logs have been treated as "non-critical" part of the system ("they're just logs!") and had not been reviewed or tested for security. If you suffer from a data breach and investigation reveals your logs were not properly secured, you'll likely be fined.
So in short, if you want to play it safe, all of the below:
- Pseudonymous IDs
- Short retention period (in accordance with other laws)
- Encryption
- Restricted access
- Industry-standard security best-practices
[1] https://www.dataprotection.ie/en/dpc-guidance/anonymisation-...
(Though perhaps one can meet compliance needs by keeping these logs only for a fixed maximum period of time, e.g. 30 days, and keeping only appropriately anonymized data longer.)
If your internal compliance people don't like it you can also rephrase it as "we are removing the data starting right now, the procedure takes 30 days". You have one month to even respond to removal requests, and can stretch that by another two. As long as you are not intentionally causing delays these are perfectly reasonable time frames.
Of course you still have to do all the other stuff for GDPR compliance, like making sure you have rules who gets access to the log system instead of just giving it to the entire company, making sure you store to an encrypted drive, etc.
Efficient or fast is not a requirement for GDPR, so it can happen slowly and in the background just fine.
In reality you need multiple different steps here: anonymous IDs, well-defined reasonable retention periods, strong access control and audit logging, and a privacy policy that says why the data is collected (for service quality typically) and how/when it will be deleted.
There's no one-clever-trick to GDPR, the law was intentionally designed to require businesses to apply holistic best practice. Whether it has done that well or not is another matter, but that was at least the aim.
GDPR request comes in, just delete the record the ID refers to and you're done.
First, as another reply above has mentioned, other data in the logs (such as IP address, list of friends, browser fingerprint) can be used to de-anonymize the pseudonymous ID.
Second, GDPR makes it quite clear (for the reasons above) that pseudonymized data, is still considered personal data. Pseudonymization reduces the risks, but does not remove them entirely. It should generally be combined with other measures such as encryption.
Much better is a well thought of error handling. This shows exactly when and where something went wrong. If your error handler supports it, even context information like all the variables of the current stack is reported.
Add managable background jobs to the recipe which you can restart after fixing the code...
This helps in 99.99% of all cases.
Error handling as well can be very helpful to communicate what your system is doing, but errors are not the only state you want to look for.
In theory, but it is something that I didn't see used too much a logging library can be wrapped into an abstraction where useful to enforce consistency. For example if wrap your library in something that conventionally sounds like "ThisLoggerIsCriticalDontMoveItAsYouWillDoWithOtherLogs(logger)" you are communicating something more about how that.
- Add OTel based instrumentation to generate traces
- Do salted hash of PII (injected in plain text by API Gateway in each request) like userid, etc to propagate internally to other downstream services via Baggage
- Inject all this context like trace-id and hashed PIIs into log
- Have Log4j and Logback Layout implementations to structure logs in JSON format
Logs are compressed and ingested to AWS S3 so it is also not expensive to store so much logs to S3.AWS provides a tool called S3Select to search structured logs/info in S3. We built a Golang Cobra based cli tool, which is aware of the structure we have defined and allows us to search for logs in all possible ways, even with PII info even without saving.
In just 2 months, with 2 people we were able to build this stack and integrate to 100+ microservices and get rid of Cloudwatch. This not just saved us a lots of money on Cloudwatch side but also improved our capability to search to logs with a lot of context when issues happens.
WordPress is fantastic for its WYSIWYG features, but it's probably one of the worst in terms of performance (especially w/o caching).
This article could have been an HTML (or even Markdown) page...
Then in my standard error logs I always just include this event ID and an actual description of the error and it's context from the call site. These logs are usually very small and easy to analyze to spot the error and every log line includes the event ID that was being processed when it was generated.
For instance, you could maintain an in-memory copy of the log DB schema for each http/logical request context and then conditionally back it up to disk if an exception occurs. The request trace SQLite db path could then be recorded in a metadata SQLite db that tracks exceptions. This gets you away from all clients serializing through the same WAL on the happy path and also minimizes disk IO.
I hope 2024 is the year where we realize that if we make the log levels dynamically update-able we can have our cake and eat it too. We feel stuck in a world where all logging is either useless bc it's off or on and expensive. All you need is a way to easily modify log level off without restarting and this gets a lot better.
Almost like a local breakpoint debugger on crash, but for prod.
In C or C++, you just run 'ulimit -c unlimited' in your shell before running your program. When it crashes, a GDB-friendly core dump is generated. Then you can load it in gdb ('gdb myexecutable mycoredump'), and it takes you to the exact line where it crashed, including showing you the stack trace, letting you view local variables at every frame of the stack, etc. Every C++ IDE supports loading a core file, so it's literally an interactive debugger at the time you most need it. It's a life-saver.
Keep in mind you have to compile with debug symbols enabled to be able to make sense of the coredump. However, you can then strip your binary, as long as you keep an unstripped copy around to help you with debugging.