I just want to serve 5 terabytes [video]
youtube.com
youtube.com
-- Rob Pike, Notes on Programming in C, 1989[0]
Generally speaking, I feel that the bureaucracy involved in a programming project should be proportional to the scale of the project itself. If the 'getting started' tutorial for your programming language demands that I choose a package name for my Hello World program, you fucked up.
The problem is that if you build quick n dirty, but then end up reaching success; you’re in for a bad time or a rewrite. This is where the above example of a single algorithm departs from a whole stack.
Further, one bad “hello world” on a google/apple/amazon domain can make the front page of the NYTimes
Well, whatever. I try to download his videos from the web UI. Select all files in directory, wait about 20 seconds, download zip, wait another 20 seconds, finally the file download dialog shows up. It gives me a 2.3GB zip file. I open it, it's just a few files and not the complete directory. It doesn't give me the contents of the whole directory, just files until it reaches 2.3GB, and then silently fails.
Great jaaab, Google!
For large binaries better to use Google Drive for Desktop (their official client), which is kinda like rsync mount but bulletproof.
Google Takeout converts them to actual, solid files.
Google drive has a zip size limit of around 2GB while downloading folders. Usually when you try to download a folder larger than 2GB, it will split these into multiple zip files which all start downloading parallely.
Most browsers block these behind this permission prompt near the address bar "Allow site to download multiple files". If you blink you'll miss it.
So someone at Google was unsatisfied with the system right click options, so they remade them. Then someone else at Google probably thought that was dumb and put the popup to remind people that Ctrl + C / V is better? I don't know, all I want to do is move data. But now I have to confront this issue, which takes me away from my actual problem and makes me process some new information that is completely irrelevant. And besides, I already pressed the button. Why didn't you respect that button press? Why did you give me the option to press that button if you didn't want me to use it?
Thanks Google.
They use a single table with columns named PK1, PK2, PK3, PK4, and so on, dumping all kinds of data into it. users, orders, addresses, and even blob data. For example, PK1 might be the user ID if the row is a user entry, or PK2 could be the address ID if it’s an address object. And if PK1 happens to contain a #, that might mean it’s a user profile object because apparently that’s how they chose to distinguish it.
As more entities were added, the table just kept morphing into this unstructured, unreadable mess.
There’s absolutely no good reason for this kind of design. It’s just a lack of understanding of how NoSQL systems are meant to be used. The team was told to “utilize NoSQL solutions,” and this is the chaos they produced.
I wish I could unsee it, but the damage is already done. It still gives me nightmares.
How did somebody so clearly incompetent get so much political capital that such a system could not only be created but perpetuated and even praised as a good idea.
And then you start to wonder… do the people who praise it secretly know it’s hot garbage or do they genuinely believe their praise? And the answer to that invites even more questions.
And then you get to watch everything else because such a monstrosity doesn’t get built in a vacuum. It’s probably the tip of the iceberg in terms of organizational dysfunction. Who are they hiring (and what does that say about you?). What process enables such things? Why do people tolerate it? Do they know any better? Is it fixable (probably not!) and if so, is it worth it (also no!)
…and of course how the fuck do I get out before I start drinking the koolaid too?
Especially the "only if you think your users are scum" part.
Back in the days, even the asshat in green seem to have tried to uphold some kind of guiding principles.
Half of it is about setting up a bigtable for hosting the data. I don't know of any team that's set up a new bigtable in the past 5 years; if you're really going all-in on the cloud ecosystem (per AppEngine) you'll stick it in Google Cloud Storage (where you'll note the file size limit is... 5 TiB) and call it a day. (Also: PCR zones are basically dead.)
The other half is about setting up monitoring, where the mentioned choices are Diplomat (which I don't think anyone's used in about a decade) and Borgmon (which is barely staffed and strongly, strongly discouraged for any new use cases). Borgmon readability hasn't been enforced in years. And again, if you're spinning this up in GCP, just set up some cloud monitoring.
Is 'monarch' still a thing? It was newish around the time that I left.
Further, the entire point of automon is to automatically generate common monitoring dashboards, which you should expect to be sufficient if you're creating a bog-standard setup.
> It didn't matter if the backend was Monarch or not.
It totally does. Borgmon has a totally different data model, a custom query language, its own UI, and various quirks along the way. To add insult to injury, it was a very real thing where if you wanted to set up new monitoring, you needed to get someone with Borgmon readability to approve your change (that requirement has since been lifted). Meanwhile, today, you don't need anyone with Python readability to ever look at your GMon code.
You can have Automon graphs that fetch from Borgmon under the hood, but everything else that you've described is 100% Monarch-specific.