PAGNIs: Probably Are Gonna Need Its
simonwillison.net
simonwillison.net
I am a big fan of CI (continuous integration) and automated tests, but please don't do this:
> Even better: continuous deployment! When the tests pass, deploy.
Continuous deployment (CD) means deploying to your production environment. This is dangerous and should always be a conscious decision by someone (preferably even multiple persons). Normally, a new release either fixes a bug or adds new functionality.
If your using CD and your software has a bug, that means that your original tests were not sufficient. The fact that you now have new software and new tests does not mean that you have solved the original bug.
If you want to add new features, you will also have to add tests that these features work as intended.
In all these cases expert human judgement is needed, it is not wise to automatically deploy changes by developers.
Also, the great thing about CD is that mistakes can be fixed really quickly: push another commit.
Not using CD is no guarantee that you will ship less bugs. In fact I'd argue that it makes shipping bugs more likely, because the friction involved in deploying means you'll bunch up multiple changes in a single release - and that's a great way to discover bad interactions between changes too late.
(Even if you don't use CD to production, having CD to a staging environment is a no-brainer in my opinion: giving everyone on the team access to a current demo of the main branch is another big productivity boost.)
If the change messes up your database and you don't discover it immediately, it is not always possible to simply rollback the database.
CD to a staging environment is fine, but I would regard it more as a form of CI.
Isn't this code review? This can and should be enforced before it goes to automatic deployment. If not, what other expert human judgement is needed?
Yes, you need to reach a certain quality bar of your automated testing, and you may choose to implement a canary system before full deployment, but I see no reason to avoid automatic deployment once these safeguards are in place.
It works fine as long as developers are used to working that way.
A job queue is slightly more debatable but most sites with some kind “app” like behaviour will have stuff that can benefit from out of band processing.
Not that you necessarily need to “build” those things yourself of course, but they each require some setup/integration for your workflow/app/deployments.
Job queues for me fall into the same category as feature flags: once you need them they add a huge amount of value, but many projects can survive for a while (or forever) without them and adding them to an existing project usually isn't too painful.
I'm not sure about this one. I explicitly avoid logging bodies because they can contain sensitive or proprietary information. But for the past few years I've been working at companies where it was not ok to look at customer data without customer's approval, that probably skews my ways.
When I implement this pattern I make sure not to log the incoming authentication token - I decode it first and log the user ID. I don't want logs that include secret tokens that could be used to impersonate an end-user.
There are patterns you can use to help with automated redaction: one that I particularly like is using a prefix on fields that should be redacted before being logged - "private_password": "..." for example.
You can also redact things on a specific deny-list, but it's relatively easy for someone to forget to update that. I like to make the redaction mechanism as obvious as possible.
There are cases where the bug is account specific.
I think the compromise here is on shorter retention policy if they contain sensitive info, or granular log access permissions.
Anyone informed better on approaching these scenarios please?
(I fully expect the list I've provided to be considered highly debatable)
Also, w.r.t. "API pagination", one also need not necessarily alter the API payloads with those previous-next parts -- RFC 8288 defined "Link:" HTTP headers (used by GitHub for pagination) meaning it's cheap to carry that kind of metadata without having to have the metadata inside the payload https://datatracker.ietf.org/doc/html/rfc8288#section-3.5 and https://docs.github.com/en/rest/guides/traversing-with-pagin...
Related to that API payload modification: another consideration is whether one turns on "Ignore Unknown properties" (ala Postel's Law) to allow extending the payload without fear of blowing up the older versions. That one cuts both ways, however, since one would want `{"resutls": []}` to blow up, so maybe a middle ground of saying that any fields beginning with "x-" or "$" are ignored
I considered using Link headers for pagination in Datasette and got some useful pushback: https://github.com/simonw/datasette/issues/782#issuecomment-...
> It has a lot of strikes against it: poor discoverability, new developers often don’t know how to use them, makes CORS harder, makes it hard to use eg with JQ, needs ad hoc specification for each bit of metadata, etc.
Always return objects from a JSON api. (never an array, never just a boolean, etc.)
You are going to need to add something even if it isn't pagination.