Some items from my “reliability list”
rachelbythebay.com
rachelbythebay.com
> Item: Rollbacks need to be possible
This is the dirty secret to keeping life as an SRE unexciting. If you can't roll it back, re-engineer it with the dev team until you can. When there's no alternative, you find one anyway.
(When you really and truly cross-my-heart-and-hope-to-die can't re-engineer around it fully, isolate the the non-rollbackable pieces, break them into small pieces, and deploy them in isolation. That way if you're going to break something, you break as little as possible and you know exactly where the problem is.)
Try having a postmortem, even informal, for every rollback. If you were confident enough to push to prod, but didn't work, figure out why that happened and what you can do to avoid it next time.
> Item: New states (enums) need to be forward compatible
Our internal Protobuf style guides strongly encourage this. In face, some of the most backward-compatible-breaking features of protobuf v2 were changed for v3.
> Item: more than one person should be able to ship a given binary.
Easy to take this one for granted when it's true, but it 100% needs to be true. Also includes:
* ACLs need to be wide enough that multiple people can perform every canonical step.
* Release logic/scripts needs to be accessible. That includes "that one" script "that guy" runs during the push that "is kind of a hack, I'll fix it later". Check. It. In. Anyway.
* Release process needs to be understood by multiple people. Doesn't matter if they can perform the release if they don't know how to do it.
> Item: if one of our systems emits something, and another one of our systems can't consume it, that's an error, even if the consumer is running on someone's cell phone far far away.
Easy first step is to monitor 4xx response codes (or some RPC equivalent). I've rolled back releases because of an uptick in 4xxs. Even better is to get feedback from the clients. Having a client->server logging endpoint is one option.
And if a release broke a client, rollback and see the first point. Postmortem should include why it wasn't caught in smoke/integration testing.
She was both a long-time SRE at Google and a long-time PE at Facebook.
Couldn't find a resume on her site to confirm it.
Amen
More than accessible: all scripts must be in git just like source code because that is also source code.
Building releases should be done using a fresh VM. Creating and configuring that VM should only use a script which is also in git.
Everything needs to be automated using scripts. If you type "apt-get" on the command line you have lost reliability. When multiple people are involved, such manual setups will be a problem at some point: people make mistakes. Manual steps also means you have lost good testing of the build process.
I always feel like people who write these never faced SQL schema changes or dataset updates. I wonder what rollback plans are in place for complete MySQL replication chains, for example.
First make sure your code handles both old and new schemas
Then introduce the schema change
Then introduce the code which depends on the schema
Each of these steps is performed separately and monitored for problems, with rollback possible at each stage. The most painful thing to rollback is the database, but it is possible, though rarely necessary if done in isolation and tested against the old code first.
There may be a few more steps after if you want to tidy up the schema by removing old data etc. It's trading off complexity in development/deploy for reliability.
Writing code that can handle both old and new schema is admittedly annoying. But it's safe and forces a thoughtful rollout.
Even data mutations can be rolled back. Dual storage writes, snapshots, etc.
The goal isn't to eliminate risk, it's to reduce the risk and make it calculated and bounded. Heck, having a rollback plan that says "we'll restore a snapshot and lose X minutes worth of data mutations" is fine in some cases. It's tradeoffs.
(I've seen a case where literally adding a column caused an outage, because the presence of the column triggered a new code path. Rollback was to delete the column.)
The typical scenario I’ve had to help with: MySQL database that’s been running in prod for a while. Because of the lax validation checks early in the project, weird data ends up in a column for just a couple rows. Later on you go to make a change to that column in a migration, or use the data from that column as part of a migration. Migration blows up in prod (it worked fine in dev and staging due to that particular odd data not being present), and due to the lack of transactional DDL, your migration is half applied and has to be either manually rolled back or rolled forward.
The former case is infrequent in real life. Users with smaller tables don't encounter it, because DDL is fast with small tables. Meanwhile users with larger tables tend to use online schema change tools (e.g. Percona's pt-osc, GitHub's gh-ost, Facebook's fb-osc). These tools can be interrupted without harming the original table.
Additionally, DDL in MySQL 8.0 is now atomic. No risk of partial application if the server is shut down in the middle of a schema change.
The latter case (multiple schema changes bundled together) tends to be problematic only if you're using foreign keys, or if your application is immediately using new columns without a separate code push occurring. It's avoidable with operational practice. I do see your point regarding it being painful on smaller projects though.
ALTER TABLE user ADD COLUMN country; -- oops
CREATE TABLE user_email (blah blah);
-- process here to move data from user.email column to user_email table. This blows up because someone has a 256-byte long email address and the user_email table has a smaller varchar()
ALTER TABLE user ADD state_province;
Yes, this is sloppy and should be cleaner. The net result is still the same though; you have the user_email table and country column created, and due to MySQL DDL auto-commit, they're persisted even though the data copy process failed. The state_province table does not exist, and now if you want to re-run this migration after fixing the problem, you need to go drop the user_email table and country column.With e.g. Postgres, you wrap the whole thing in a transaction and be done with it. It gets committed if it succeeds, or gets rolled back if it fails.
Otherwise, if the table is quite large, the DML step to copy row data would result in a huge long-running transaction. This tends to be painful in all MVCC databases (MySQL/InnoDB and iiuc Postgres even more so) if there are other concurrent UPDATE operations... old row versions will accumulate and then even simple reads slow down considerably.
A common solution is to have the application double-write to the old and new locations, while a backfill process copies row data in chunks, one transaction per chunk. But this inherently means the DDL and DML are separate, and cannot be atomically rolled back anyway.
It's admittedly been a long time since I've used MySQL for anything significant, but I feel like teams in the past have run into issues where DDL operations on their own have succeeded in dev/staging and failed in prod, even though they're running on the same schema. Simply due to there being "weird" data in the rows. I don't know that for sure though. If I'm remembering right, one of those had something to do with the default "latin1_swedish_ci" vs utf8 thing...
They planed changes to the system years in advance in lock step with changes to the physical plant and incrementally introduced changes.
eg release 55 and 56 would introduce stuff that would be used 6 months later in release 58 "Caller ID works on every phone in the country"
Apart from that you can do stupid SRE tricks like copying tables around in-memory before schema changes.
It's consistently a source of sorrow to me how many bugs exist and how much inefficiency there is because people don't want to be embarrassed by the code they quickly wrote.
I just had a discussion about this yesterday where we have an internal JSON API that auths a credit card, and if the card is declined it returns a status and a message. Another developer wanted it to return a 4xx error, but that made me uneasy. I think you could make a good argument either way, but to me that isn't a failure you'd present at the HTTP layer. 4xx is better than 5xx, but I was still worried how intermediate devices would interfere. (E.g. an AWS ELB will take your node out of service if it gives too many 5xxs, and IIS can do some crazy things if your app returns a 401.) Also I don't want declined cards to show up in system-level monitoring. But what do other folks think? I believe smart people can make a case either way.
EDIT: Btw based on these Stack Overflow answers I'm in the minority: https://stackoverflow.com/questions/9381520/what-is-the-appr...
>The 400 (Bad Request) status code indicates that the server cannot or will not process the request due to something that is perceived to be a client error (e.g., malformed request syntax, invalid request message framing, or deceptive request routing).
Of course, a lot of things could be said to be a client error.
But at the end of the day, your app must target a subset of available APIs. If you use every available API the surface will be too big and you will have snowballing complexity. So if you are already using another API which can signal what you’re trying to signal then feel free to institute a “everything we can expect returns 200” policy. That kind of decision is critical for putting an upper bound on your system complexity.
I want to institute a “no default exports” policy in the JavaScript app I’m working on. It’s not that there aren’t places a default export makes sense—there are. It’s that prohibiting them will decrease the number of things our coders have to think about.
This is a line I hear a lot in regards to this general situation, and the interesting part is that the 400 code really does not signal the thing the developer is trying to signal, not at all. But a lot of devs have a weird hang-up about this for some reason, and will go to any lengths to not send a 200 response. I blame the REST cultists.
HTTP status codes can be used to good effect if done thoughtfully, because there's often already a lot of logic in the useragent to handle some cases (such as redirects, etc). Not using anything except for 200 is limiting. Using a bunch of status codes to refer to things that make no sense is confusing. Defining a sane standard for how your application will choose between them that allows for the benefits they offer without confusing clients as to what is actually being communicated is the best of both though.
You seem to be using a strange definition of "application", but assuming you mean "transport" — I mean, this is obviously not true, right?
If you sent a perfectly semantically valid HTTP request for a resource that doesn't exist on the remote, that's a 404. But by your logic, you're saying it should be a 200, because there wasn't a failure in (essentially) nginx or the HTTP library of your application server?
HTTP is a transport designed for a specific type of application (document-oriented resources) and its status codes (and lots of other things) reflect that initial coupling, so it's always gonna be a pain when you swap out the application (to a credit card processing service or whatever) and have to think about how the status code mapping will work. And you can certainly "cheat" and say that every request that doesn't crash the server is 200 OK, I guess, but that's pretty clearly not what the Hyper Text Transfer Protocol wants you to do. Its set of status codes clearly reflect details of the underlying application.
>If you sent a perfectly semantically valid HTTP request for a resource that doesn't exist on the remote
Requesting a path (via the application layer) that doesn't exist is a client error, hence the 4XX. But, say, if your website has a search bar and a client searches for "blahblah-thisdoesntexist" and gets no results, the webpage should return a 200 (the request was valid: the client sent form data to an existing resource and it was processed successfully) with something like "no results found" in the body -- because the search function itself is above the application layer. This is what I mean about "is the CC is authorized or not" being above the application layer.
>Its set of status codes clearly reflect details of the underlying application.
This is where I think you're wrong: HTTP status codes reflect the result of the HTTP request itself, not the next application in the chain, if one exists.
This can also be expressed if we look lower down the chain. If my HTTP request gave me a 404, should an error be shown in the TCP stack? Well no, because TCP doesn't know or care about HTTP, as far as it's concerned the TCP connection is established and happy. Likewise, in this example, the API HTTP stack doesn't care about whether the data was a CC or something else or what the result was: as far as the HTTP stack is concerned, some JSON got passed in and parsed properly, and some output got sent back, and everything went well so "200".
I understand OSI, but it's not a useful model in this circumstance, because it draws no distinction between the HTTP layer and the business logic served over HTTP. The complexities of that interface is what we're talking about.
> HTTP status codes reflect the result of the HTTP request itself, not the next application in the chain, if one exists.
Let's say I request /robotz.txt and I get 404 Not Found.
How do you maintain that 404 is a result of "the HTTP request itself" and not the HTTP server handing the request to the component that looks on the filesystem (or in the cache, or in MongoDB, or...) and not finding the resource?
At that point, you rely on the body content to tell you what the service correctly determined for you. A result that the user doesn't like is way different than a result that comes about because something was done wrong at the client side (4xx) or a failure on the server side (5xx).
It can get a little fuzzy for 401 and 403 errors, though; those seem to be returned by APIs pretty often. Not entirely sure how I feel about that, but I think those are a bit more sensible.
> The first digit of the Status-Code defines the class of response. The last two digits do not have any categorization role. There are 5 values for the first digit:
> - 4xx: Client Error - The request contains bad syntax or cannot be fulfilled
What do you think status codes like 404, 405, 406 are for? You say they shouldn't be for "did the application successfully validate the user's input data" but status code 400 is explicitly for bad requests. In your view should a HTTP server ever return 4xx?
Obviously if you send a TAIL method request, you should get a 405, and if you send Accept: eggs/*, you should get a 406. If a route doesn't exist, you should get a 404. If you fail HTTP basic auth, you should get a 403 (but why are you using HTTP basic auth?). If you want certain paths to never be accessed for some reason, you should return a 401.
But a provided credit card number not existing is not success, it is unambiguously a failure.
HTTP is a transport but its response codes are clearly designed to map somehow to the "business logic" of the thing underneath it. We have to twist ourselves up in knots to map most of our business systems to the set of status codes that make sense for the default HTTP-fronted application (Fields' thesis stuff) but that's what we sign up for when we decide to expose our business services via HTTP. Other transports require different contortions.
404 means "this API doesn't exist", not "the API exists but it returns an error"
The 404 (Not Found) status code indicates that the origin
server did not find a current representation for the target
resource or is not willing to disclose that one exists.
Resource here is a transport-level (HTTP) representation of a business-logic concept, in most cases a document, but if your HTTP server happens to front a credit card validation service, potentially a credit card. Or a user. Or a recipe.The payoff of using 400 is you can watch your 400 rates with almost no effort (HTTP is well established and there's many many tools out there.) If you somehow start accidentally munging the card number sometimes or if your card processor starts doing wacky stuff you'll see a spike in 400 rates.
If it was really that troubling that declined cards are expected, I would personally at least want to see 200 come from the internal API and 400 go out to the client.
And if your "intermediate devices" start doing goofy stuff to 400's then you've got bigger problems... 4xx's shouldn't be taking nodes out of prod. That's wack.
400 is "bad request"; you might use it if the request body was not valid JSON.
Welcome to AWS, where this is actually part of standard procedure (at least Elastic Beanstalk does this, not sure if it’s actually Elastic Loadbalancer under the hood).
But the failure is ambiguous. If I get a 404 back I can't rely on that to mean the credit card doesn't exist. It could mean that I have the wrong URL or that the service application screwed up their routing code.
And who's to say you can't put the reason in the body and still keep the code? What are you hurting by sending back 400? Unless you have lb's taking out nodes because of excessive 4xx's (which sounds like insanity) I don't see a reason _not_ to send 4xx's. At the very least it's a useful heuristic tool.
Often systems will have application wide error handling to catch that and handle it in a systemic way. It can be a pain to short circuit that in a customizable off the shelf applications like Salesforce.
Philosophically, 4XX means the client did something wrong. Sending invalid data to a validation service is not doing something wrong.
The precise definition of 400 Bad Request is
The 400 (Bad Request) status code indicates that the server cannot or
will not process the request due to something that is perceived to be
a client error (e.g., malformed request syntax, invalid request
message framing, or deceptive request routing).
There's some room for interpretation for the phrase "process the request". If the only job of the service is to validate the request and return Valid or Invalid, then I guess yeah even an invalid request will have been processed, and 400 Bad Request may not be totally right. I think you could make a case for it! But I think you could also make a case for 200 OK with a more detailed error message and code in the body.But if the service both validates the request and _then_ does something more substantial with it, I think 400 Bad Request is probably the most appropriate response to something that fails at step one.
I had this discussion recently about 'security' with regard to X-Header versus ?query=param. Either it's http all plaintext on the network or it's http with tls all cyphertext on the network. Every bit in the http request and response is equivalent - verb, path, headers, body, etc - agree?
You could represent the card number as cyphertext in the request body, that's a good practice regardless of tls, but of course don't roll your own crypto. You could put that cyphertext in the path as well but if the cyphertext isn't stable that makes for a huge mess of paths.
You could make a case for trad 'combined' access logs situation with the path disclosed in log files. I can appreciate keeping uris 'clean' makes it safe to integrate a world of http monitoring tools, I would make this argument. In the case of the card represented in a stable cyphertext it's kinda cool to expose it safely to those tools.
Anything else?
And let me say it's totally valid to just ignore that advice, and say that here, in this company, we use 200 OK for everything that isn't a parse failure. And that's not totally wrong! It won't break anything. You're not really leveraging any of the power of your transport but maybe that's fine for your situation.
Doesn't this depend on the nature of the REST API? (Though putting a credit card number on the path doesn't seem like a good practice.)
Personally I prefer to think of HTTP as a transport mechanism, with business logic at a higher level. The question "If I replaced this http transport with, say, FTP or a dvd drop, would it require changes at the business logic level?" feels important to me.
That preference puts me on the wrong side of history at the moment though, and Im going with history. Lets express our business logic through the error system provided by our transport mechanism instead....
> A "HTTP 400 bad request" is only the sort of thing you get to ignore when it's coming from the great unwashed masses who are lobbing random crazy crap at you, trying to break you. Internally, that's something entirely different.
If this is an internal-only API then most people look at 400s as their bug. Sounds like you are in the same boat.
If it's an external API, it's not their bug.
Specifically, you want to be able to easily distinguish between "your url is wrong" or "your authentication credentials are wrong" or "the API endpoint threw an exception" type errors and "credit card processing failed" errors.
It's easier long run to put business logic errors somewhere separate from protocol/routing layer (i.e. even a http header would be better) so that you can tell what is Rails/Flask/whatever failing vs. logic failing. This also gives you more flexibility to do stuff at the hardware layer (another commenter mentioned ELB) without interfering with the application layer.
Generally, 4xx errors means 'Client screwed up'. There's probably no point sending this message again, it isn't going to work until something is fixed. That might mean it will work if, for example, the account is funded or card unlocked. But something needs to be done on behalf of the client.
5xx errors mean 'Server screwed up'. It probably _is_ worth having another attempt at sending the request. Maybe it was a temporary glitch, or maybe a new release of the server has to happen. Regardless, there was probably nothing inherently wrong with the actual request.
You could argue either way. What's important is being consistent, at the very least throughout an API but preferably throughout the whole organization.
(Personally I'd probably lean towards a 40x of some kind, just make sure it doesn't clash with something that you care about.)
Along the same lines, and arguably more important, is how to log the operations where a transaction completed successfully but with a negative answer. If you log expected negatives as errors you can get error blindness.
It's very easy to get absorbed into the awareness of the high level change you're making and miss the details of the process. Even just sitting down together and outlining what you think is actually going to go on (and then breaking those down into what they each are comprised of) can make it really clear that you don't have to run as many giant risks. I'm occasionally amazed how brilliant people (including some with big names in devops) can forget it's an option.
It's like taking small steps from stable to stable when you're going across a steep scree slope and only jumping when you have to - sometimes it feels riskier to take lots of small steps, but if you start to slide it can be a lot easier to recover from. Your chance of dying taking a big leap isn't the sum of the equivalent small steps. Perhaps complex computer systems have the equivalent of an Angle of Repose?
"if you only need 53 bits of your 64 bit numbers"
JSON numbers are arbitrary precision.
"blowing CPU on ridiculously inefficient marshaling and unmarshaling steps"
On the other hand I am not blowing dev and qa time on learning/developing tools to replace curl/jq/browser/text editor.
This specification allows implementations to set limits on the range and precision of numbers accepted.
https://tools.ietf.org/html/rfc7159#section-6Other implementations can have other limits, but that does not mean that JSON itself or all implementations have the specific limits of javascript.
>JSON.stringify(1n)
TypeError: BigInt value can't be serialized in JSON
Chrome: >JSON.stringify(1n)
Uncaught TypeError: Do not know how to serialize a BigInt
at JSON.stringify (<anonymous>)
at <anonymous>:1:6
I don't think we're quite there yet.Javascript can do arbitrary precision integers now. So the last part is not true. You can generate your JSON strings with string concatenation in JS if you like.
Just because JSON.stringify has some limit doesn't mean JS has.
For everything else, it is really a non-issue.
> 9007199254740991n * 9007199254740991n
> 81129638414606663681390495662081nThe fact that JS has 53 bit precision will be a JS problem whether you use protobuf or anything else. On the other hand, if you are not using JS, it will almost certainly parse numbers of the precision that your language offers.
That's impossible with text-based formats like XML and JSON, because everything is fractally variable-length.
At smaller scales, I think the (machine) benefits of using something like protobuf don't nearly outweigh the human benefits of just using JSON.
On the other hand, if you have many developers, and the problem space starts converging toward scale, then certainly the problem space is different. Encoding of data cannot follow the language with the weakest type zoo, and efficiency starts to matter.
The hard part is to know when to make the switch, or when to anticipate growth in advance, such that you pick the right tool for the job.
It isn't just a question of the machine. For complex messaging, the added value of a well-defined message typing, provided by protobuf, will help. It will also remove a lot of problems if you have multiple different languages in the stack, talking to each other.
Just kidding, but I suspect it would be a lot faster.
This is really, really bad thinking and will always cause you a timebomb that will 100% explode on you in the long run.
"just a quick POC" will always end up being the actual product, and once you start down this path you end up with "let's just use mongo, it's schema-less, it'll be faster for devs and we can use a proper db later"... "let's just use nodejs for the POC, all the devs already know JS so it'll be great"... a few years later you end up with some gargantuan monstrosity that nobody wants to maintain, your "quick" language/db-du-jour is dead and unmaintained, you're EOL on four different software fronts and you now get to explain to the bosses why you need to rewrite the last four years of dev work.
Problems started somewhere in early '00s when "software architect" became a dirty word, and one hit wonders like Paul Graham pontificated that "young is smarter".
And here you are.
1. Single source of truth for certain types, enums.
2. (assuming you’re using gRPC too) A separate, minimalist, and language-agnostic definition of your service interface, which you can document with comments to your heart’s content, and which (unlike other documentation) can NOT be out of sync with the actual service. I don’t have to read C++ to understand your C++ service should work.
3. Protobufs encourage message type / enum reuse, by allowing you to import other definitions. This might seem trivial, but it’s super important in mediumish orgs that everyone is using the same definition of time, geography, etc. It all adds up to less surprises when you open up a new .proto file.
The kicker is not that you can’t somehow get these things with JSON-over-HTTP, too. It’s that protobufs-over-gRPC won’t work without them. The trade off is that you can’t inspect raw requests unless you have built some tooling around it.
So now there is a point in your codebase where you are receiving data in a JSON form. At this point having protobuf elsewhere is not chosing between JSON and protobuf, it's chosing between "json+protobuf" and "json".
gRPC-web isn't quite feature complete relative to normal gRPC, but it is getting pretty close, and the gains of avoiding JSON (de-)serialization would be big. I think once the protobuf story has a complete chapter for the front-end, bigger engineering orgs will roll it out much like they're rolling out typescript today.
If you're talking about actual developer/public-facing APIs, those will probably remain in JSON land for a while.
IMO they're readable enough to serve #2 too.
Before someone mentions "JSON schema" - I've tried to use them, and they're difficult to write and impossible to read. Reading/writing TypeScript types are a dream in comparison.
Protobuf - or, more specifically, proto files - gives you a central place where you can define and also document your formats. You can throw an ASCII-field UML sequence diagram in there, if you need to. And it's right there in the single file that everyone will use to communicate the protocol, and the protocol at least can't change in any structural way without editing that file, so it's got a much higher chance of being kept up-to-date, and of being read by the people who need to read it, than any of the available options for documenting JSON-based protocols.
All JSON gives you is human readability, and browsers can read it without a library. The 2nd, I don't care about with back-end services. The 1st I don't really care about at all, because command-line utilities and library functions for dumping protobuf datagrams to a text format are a dime a dozen.
The more parts of your stack you tighten up, the fewer errors you'll hit and the more flexibility you'll have to use less-strict tools when it really matters. That's true at all but maybe the smallest scale. Worrying about the costs of having to document what you intend in a way your machines can verify—I mean, shouldn't you be doing that anyway?—is baffling to me. You don't need anything like Google scale to see the benefits of it. It's basic communication AFAI am concerned.
Because learning how to use correct tool for the job is so difficult and so outside of dev/qa job descriptions, that it's much better to waste compute on your servers and deal with performance and scalability problems.
It's a pain, but it's a short term pain, and necessary.
Your comment is unnecessarily inflammatory.
> AWS bills can easily eclipse engineer salaries if you use it inefficiently.
If this happens, it is more due to architectural and algorithms problems than to the runtime of the language. There are many large applications written in Python, Ruby and PHP that run just fine.
What a pleasant surprise it will be for you when you find out that jq silently corrupts integers with more than 53 bits.
But still, given that easy consumption from JavaScript is the ultimate primordial reason for choosing JSON over other formats, it seems like trying to transmit integers with more than 53 bits of precision over JSON is asking for trouble. Because it's only a matter of time until someone will want to do something like write a new service in Node, and the JavaScript parsers for other formats are at least somewhat more likely to guide people toward using BigInt for large integers.
[1]: https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
I explicitly asked for and achieved step 2. in the modified SerializeJSONProperty algorithm[1] so that users could decide and opt-in to serializing BigInts as strings if they so choose, with or without some sigil that could be interpreted by a reviver function. e.g.:
> JSON.stringify(BigInt(1))
TypeError: Do not know how to serialize a BigInt
...
> BigInt.prototype.toJSON = function() { return this.toString(); }
> JSON.stringify(BigInt(1))
'"1"'
[1]: https://tc39.es/proposal-bigint/#sec-serializejsonpropertyNo, it's typically IDs. And my observation is that inexperienced devs start out serializing 64bit IDs as ints and get burned first before they move to text.
There are several reasons why you end up with IDs > 2^53 even if 2^53 is good for enumerating ~100k things per earthling:
Many cloud providers will give you IDs that are specified to be 64 bit uints, and you'd better not corrupt some fraction of them. Generating unique integers strictly serially does not scale, so you might end up with something like Twitter's snowflake. Or you might want to tag additional info (reserve a few of the low bits for the type of id, or reserve particular ranges for particular types of acccounts etc etc).
I can promise you I have run into the issue in real life more than once and it's not just javascript and jq that like to silently corrupt ids by truncating them to double precision. For example pandas is really great at it as well.
What makes it especially fun that often only a very small fraction of IDs will be affected (< ~1 in 2000 if uniformly distributed).
Is your turnover so high that learning a single tool is a significant portion of employee costs?
I think recent experience should tell us that the easiest thing plainly wins out. JSON can be terrible, Javascript can be terrible, but limitations on large integers isn't the worst hurdle to take.
I haven't heard enough people discuss the deployment management of growing enums or state machine evolution. This is a problem more particular to software than hardware, as once hardware is shipped it's usually set in silicon, but growing of the state garden is an expectation in many software architectures.
Some customers will skip releases altogether making strategies like add a new column, back populate it online, then the next release uses the new value impossible.
I guess that point is slightly moot when it'd take 2-3 releases to achieve the end goal and each release cycle is about a month.
But your biggest reliability improvement would come from getting this system moved to continuous deployment without downtime. Now you can make one change at a time and roll back if it doesn't work.
There's still a lot of small software companies maintaining, and selling, on-premises installations of systems that use, e.g. Microsoft Access as a front-end client. Even then, continuous deployment is possible, and all-else-equal, a huge improvement for the developers and support staff, but also something that lots of management or owners may be (reasonably) averse to committing to implementing.
In proto3 all fields are optional, and have default values, so it becomes impossible to detect the absence of data unless you explicitly encode an empty/null state in your values.
There are other ways:
- Check if the value different than the default, e.g., empty string
- If your data is repeated, then check number of data elements != 0
which is a worse version of the GP's
> explicitly encode an empty/null state in your values
Some migration tools do support rollback scripts for schema changes, but unless you're actually testing these before release (deploy the new version in staging, accumulate representative data in the new schema, roll back the schema, deploy an old version of the app, test that it is doing the right thing), then they aren't really something you can rely on in production.
Lets say you want to add a new non nullable column with foreign keys, to replace an old non nullable column of foreign keys to a different table that’s obsolete and needs to be deleted.
1) update the code to be ok with a new nullable column. Rollback: deploy previous version of code.
2) create the new column in DB with it’s desired constraint, but make it nullable. Roll back: delete the column.
3) have the code start populating the new column as well as the old. Rollback: deploy previous version of code.
4) start backfilling historical entries with the new column. Rollback: you can’t roll this back!
5) make new column non nullable. Rollback: make it nullable.
6) update code to read from new column, continuing to write to both. Rollback: deploy previous version of code.
7) make old column nullable. Rollback: you can’t roll this back!
8) stop writing to old column. Rollback: deploy previous version of code.
9) once you’re satisfied the old column is no longer used and the version of code from the previous step will never be deployed again, drop the old column.
Rollback: you can’t roll this back!
10) rename the obsoleted table and see if anything breaks. Rollback: rename it back to its original name.
11) delete the obsoleted renamed table.
Rollback: you can’t roll this back!
* Use external online schema change tooling which operates on a shadow table, so the tooling can be interrupted without affecting the original table. (Generally all of the open source MySQL online schema change tools work this way.)
* Use declarative schema management (e.g. tool operates on repo of CREATE TABLE statements), so that humans never need to bother writing "down" migrations. Want to roll something back? Use `git revert` to create a new commit restoring the CREATE TABLE to its previous state, and then push that out in your schema management system. (Full disclosure, I spend my time developing an open source system in this area, https://skeema.io)
* Ensure that your ORMs / data access layers don't interact with brand new columns until a code push occurs or a feature flag is flipped.
Can't those first 2 steps be combined together? Why do they need to be shipped separately?
1. Have the code recognize the new value. Get that shipped everywhere.
2. Have the code do something reasonable when the new value appears. Get that shipped everywhere.
3. Start emitting the new value.
My proposal is
1. Have the code recognize the new value. Have the code do something reasonable when the new value appears. Get that shipped everywhere.
2. Start emitting the new value.
If the user has to manually update, and might refuse to do so, the problem exists under both the 3 step and the 2 step process. So nothing is gained by using the 3 step process.
If the client and backend are updated independently, that should still be fine with the 2 step process. The 2 step process says "Get that shipped everywhere", meaning shipped to both the client and the backend.
1. Support the new value in your schema
2. Support the new value in the client
3. Emit the new value from the server
Combining 1 and 2 is probably possible, but not a great idea. Imagine if you end up rolling back 2 and 3, but there are still potentially new values in flight. If 1 is still there, you're good. If 1 and 2 are combined, you rollback all 3, but the new value is still in flight and your client crashes.
You have to
* Ship the new version
* Wait until you are sure you don't need a rollback to an old version
* flip the switch
If you don't wait, you might need to roll back to an old version that doesn't support the new state, which then blows up.
The third can be a flag flip instead of a code push, but it still needs to be a discrete event to start generating the new value that triggers new behavior.
But also there could have been a bug in 1/3 that wasn't exposed until step 3/3. Meaning even rolling back to 1/3 won't solve the problem.
The 2 step rollout is safe if we assume there are no bugs in the code, but problematic if there are bugs. But the 3 step rollout is also problematic if there are bugs. I guess the 3 step rollout has the advantage that it partitions off a section that might be buggier than other sections so that it can be individually disabled. But that'll only help sometimes, and I'm not sure if the additional complexity is worth it.
Flag flip vs code push doesn't seem to make a difference to me. All 3 rollouts could be flag flips enabling code that was written much earlier but hidden behind flags.
Another problem is doing a double rollback like that seems a bit risky to me, because other features that were being deployed simultaneously might not have been designed to handle a double rollback, only a single rollback, so they could break. If we want to allow double rollbacks, we must require that all development not just handle single rollbacks smoothly but also double rollbacks.
Are there any good books that are full of more rules of thumb like these?
that seems like a non-trivial point of friction when it comes to "just using solid storage/RPC formats" or whatever.
However, I think the point in the article is that you need to have well defined schemas for inter-service messaging. Something like protobuf or thrift or flatbuffers. Whether you layer gRPC on top of that is a separate concern. For example I have used Protobufs extensively at work but never gRPC, since we mostly have point-to-point connections. We checked in the message schemas into their own repo and all users across the company pull from it. We have Python and C++ codebases sharing the same schemas, it’s quite wonderful.
The biggest impacts for our code health are
- Typescript
- Automated Testing through Gitlab CI