HNHacker News
TopNewBestAskShowJobs

dub

384 karma · joined July 2, 2015

submissionscomments
dub··on How Pinterest scaled
Typically the price of not having horizontal scaling is felt more by the engineers than the users, at first:

- Data migrations, schema changes, backfills, backups & restores, etc., take so long that they can either cause or risk outages or just waste a ton of engineer time waiting around for operations to complete. If you have serious service level objectives regarding time to restore from backup, that alone could be a forcing function for horizontal sharding (doing a point-in-time backup of a 40TB database while dropping some unwanted bad DELETE transaction or something like that from the transaction log is going to be very slow and cause a long outage).

- The lack of fault isolation means that any rogue user or process making expensive queries impacts performance and availability for all users, vs being able to limit unavailability to a single shard

- When people don't have horizontal scalability, I've seen them normalize things like not using transactions and not using consistent reads even when both would substantially improve developer and end-user experience, with the explanation being a need to protect the primary/write database. It's kind of like being in an abusive relationship: you internalize the fear of overloading the primary/write server and start to think it's normal not to be able to consistently read back the data you just wrote at scale or not to be able to use transactions that span multiple tables or seconds as appropriate.

dub··on GCP CloudSQL Vulnerability Leads to Internal Container Access and Data Exposure
As the article says, the vulnerability was fixed in April and the people who discovered it have already been rewarded under Google's Vulnerability Reward Program. Google also proactively detected the problem before being notified by the researchers.
dub··on Scaling Databases at Activision [pdf]
The reasons Vitess didn't have foreign key support historically actually weren't a bold performance tradeoff or anything like that. It was more of a classic, boring, backlog prioritization thing: everyone heavily using Vitess was using gh-ost for schema changes, and gh-ost didn't support foreign keys.

Now that Vitess has native schema change tools it's more reasonable to revisit user-friendly, out-of-the-box foreign key support.

dub··on Remote code execution vulnerability in Google they are not willing to fix
When I worked at Google back in the day, we used to make dollar bets all the time. You'd tape the signed dollars you won to your monitor.

A willingness to take pride in your work and to not take it too seriously when smart, well-intentioned people make mistakes (e.g. blameless postmortems) is part of the culture difference that led to Google's engineering becoming so exceptional and innovative vs the more corporate, don't-rock-the-boat, fear-driven culture that the traditional businesses had at the time.

dub··on Remote code execution vulnerability in Google they are not willing to fix
If you're not even willing to make a bet for a single signed dollar, that doesn't speak highly to your confidence in your work.

It's fine to not be confident, but when professional security teams at large companies are afraid to express confidence that their systems are non-trivial for a random engineer to hack in their free time, that seems at odds with the claim that it's "obvious" that permission escalation is hard

dub··on Remote code execution vulnerability in Google they are not willing to fix
> Obviously not true, in fact none of the companies I worked in that was the case

I once offered a bet to the large security team at a well-known decacorn tech company I worked at: I offered to make a personal, reasonable-sized cash bet with any member of the security team that I would win if I could deploy malicious, unreviewed code to any service or machine of their choice without it being prevented or proactively noticed by them.

The members of the security team all declined my bet. We're talking about a team of probably at least a dozen people, many of who had been working at the company far longer than I and who had been shaping and reviewing the company's security design for years.

They knew perfectly well that I would be able to win the bet. Not because their security was unusually bad, but because it was bad in the common, usual ways. Securing the supply chain is hard, and real security is almost impossibly expensive to add to a system late in the game if you didn't design it in from the beginning.

dub··on You cannot have exactly-once delivery (2015)
> A surprising number of systems exhibit this behavior, sadly.

I noticed [0, ∞] delivery semantics in a widely-used, internal/homegrown message delivery system at a big tech company once. The bug was easy to spot in the source code (which I was looking at for unrelated reasons), but the catch-22 is that engineers with the skills to notice these sorts of subtle but significant infra bugs are the same engineers who would've advised against building (or continuing to use) your own message delivery system in the first place when there are perfectly serviceable open source and SaaS options.

dub··on Ask HN: What happened to flatbuffers? Are they being used?
JSON is the default serialization format that most JS developers use for most things, not because it's good but because it's simple (or at least seems simple until you start running into trouble) and it's readily available.

Large values are by no means the only footguns in JSON. Another unfortunately-common gotcha is attempting to encode an int64 from a database (often an ID field) into a JSON number rather than a JSON string, since a JS number type can lead to silent loss of precision.

A more thoughtful serialization format like proto3 binary encoding would avoid both the memory spike issue and the silent loss of numeric precision issue, with the tradeoff that the raw encoded value is not human readable.

dub··on Ask HN: What happened to flatbuffers? Are they being used?
While I haven't benchmarked JSON vs protobuf, I've observed that JSON.stringify() can be shockingly inefficient when you have something like a multi-megabyte binary object that's been serialized to a base64 and dropped with an object. As in, multiple hundreds of megabytes of memory needed to run JSON.stringify({"content": <4-megabyte Buffer that's been base64-encoded>}) in node
dub··on Log4Shell Still Has Sting in the Tail
Open Office losing popularity and having a shortage of developers makes some sense to me given all the progress in web-based document editors.

Something I have a harder time understanding is how it came to be that Apache Thrift and Facebook Thrift both exist as competing implementations of the same software originated by the same company.

dub··on Log4Shell Still Has Sting in the Tail
> What kind of brave soul wants to trudge through and maintain log4j in their spare time for zero compensation?

It's not clear to me as an outsider what exactly the Apache foundation is doing for these projects. It feels like Apache is willing to accept code donations from anyone and is willing to attach the foundation's name to code that isn't widely used, actively maintained, or may just be abandonware.

I have soooo much more confidence in CNCF projects. The conditions for graduating as a CNCF project include criteria like that your project must be in use by multiple real companies, have maintainers who are (paid) employees of multiple different companies, and get a professional security audit.

dub··on GPT based tool that writes the commit message for you
I'd be more excited to use GPT to draft a summary of release notes by scanning all the new PRs in a release, summarizing what they are, and dividing them up into categories (bug fix, feature, breaking changes, etc.)
dub··on The Twitter Advertiser Exodus
> Anything that isn't in the "happy path" of the AdsUI probably gets handled by some engineer making some API calls to a prod API

Prior to going private, Twitter would have had recurring Sarbanes-Oxley audits. Auditors understand the need for occasional emergency break-glass methods of making manual database queries or API calls, but they are less tolerant about that being a normal way of operating.

Plus, if you use emergency access often you'll eventually waste more time explaining each individual access to auditors at the end of the quarter than it would have taken to just implement a UI for the feature in a code-reviewed and audited internal admin console or user-facing UI.

dub··on FTX to file for U.S. bankruptcy, CEO resigns
It's been half a year since terra/luna crashed but the name lives on at Nationals Park: https://www.mlb.com/nationals/tickets/premium/nightly/terra-...
dub··on Tesla engineers were on-site to evaluate the Twitter staff’s code, workers said
With TSLA in the S&P 500 it's difficult to avoid being a shareholder
dub··on Elon Musk owns Twitter: The story so far
Class A players don't work for controlling managers who overload them with near impossible targets. Class A players have plenty of options in tech, and have no need to tolerate having their skills called into question or disrespected by pointy-haired middle managers.

When Bill Coughran was asked how he was able to successfully manage so many famously high performers (engineers like Jeff Dean), Bill explained how that caliber of engineer tends to have strong opinions and you basically have to build a whole team around them and give them the freedom to do their thing.

dub··on Office vacancy rate in San Francisco just hit a new high
SF was never a late night city, but it's become even sleepier in recent years. Some examples:

Orphan Andy's, Safeway on Market, and Pinecrest diner are no longer 24/7. Sparky's closed forever (a 24/7 diner). It's Tops Coffee Shop stopped staying open late and then eventually closed forever. Corner/liquor stores that used to be open late are closing well before midnight. Kowloon Tong Dessert Cafe used to be open until 2am most nights, but now they always close at midnight.

dub··on How boring should your team's codebases be
Figuring out how to reward simplicity, reliability, and maintainability feels like one of the most important unsolved social/human/economic issues in the software industry.

Seems there's only incentive to simplify at small companies where the employees feel they can save their own time or increase the value of their equity by delivering value to customers more efficiently. At large companies employees work 40-hour weeks regardless of their output and they're trying to impress a performance review committee, not customers.

dub··on California man fined for selling maps of property boundaries without a license
"It may be proved that no society can make a perpetual constitution, or even a perpetual law.... Every constitution then, & every law, naturally expires at the end of 19 years. If it be enforced longer, it is an act of force, & not of right"

- Thomas Jefferson to James Madison in 1789, expressing feelings on why laws should automatically expire

Full letter: https://founders.archives.gov/documents/Madison/01-12-02-024...

dub··on How the Apple AirTag became a stalker’s gift
Note that the product you linked to is larger, much more expensive (it requires a $20-per-month monthly subscription), and has much shorter battery life
dub··on Understanding Google’s File System (2020)
I think what people really want is:

- SaaS solution, not just open source (improvements to s3, for example)

- A global hierarchy of files rather than an assortment of different buckets like s3

- Per-subdirectory permissions and chargeback so anyone can create project-specific subdirectories and have their organization billed for storage costs

- Filesystem drivers allowing every production and development computer and coding environment to access remote files in the same way and with the same ease as remote files. The difference between accessing remote files and local files in Google code is literally just the file prefix. In all other ways, remote files are _better_ than local files (they have all the same features, plus more) so there's very little reason to directly use local disk most of the time.

dub··on Removing HTTP/2 Server Push from Chrome
There's nothing stopping developers from doing server-side validation before sending markup to clients.

It seems like the majority of developers have never really wanted to make that cost-benefit tradeoff, though. I ran tons of websites declaring an xhtml doctype through https://validator.w3.org/ back before HTML5 when xhtml was still trendy: almost all of the "xhtml" websites failed validation.

dub··on Phrases in computing that might need retiring
The trouble with the term "technical debt" is that people tend to use it retroactively to describe messes which may not have been necessary or the result of an intentional technical choice.

It's uncomfortable to say, "In retrospect that was an unfortunate, unnecessary, and shortsighted choice that didn't really save any justifiable amount of time or money", so instead people say, "We've accumulated some technical debt" to cover up mistakes instead of acknowledging and learning from them.

dub··on An ex-Googler's guide to dev tools (2020)
I don't particularly want to "mess with sandboxes", but I do want my builds to be relatively fast, correct, reproducible, extendable/customizable, with bonus points for being secure (meaning a compiler shouldn't be able to tamper with parts of the output it has no business tampering with) and more bonus points for supporting distributed builds and/or distributed caching

If someone wanted to make a new build system to compete with bazel and have those kinds of features, it's probably a safe bet the competing system would use some kind of sandboxing as well

Even if you ignore everything else, just the security part is a big deal: supply chain attacks are an increasingly big concern for companies of all sizes. If your build system allows any script invoked during any part of build process to secretly read or modify any input or output file, hackers are going to love it.

Almost all tech companies (even the multi-billion dollars ones) that aren't doing something in the spirit of `bazel build` to generate their binaries have wide open, planet-sized security holes in their build systems where if you get one foot in the door you can pretty much do anything.

dub··on An ex-Googler's guide to dev tools (2020)
make(1) has no native support for giving each build rule its own sandboxed view of the filesystem like bazel does.

If I could have a wish to upgrade file transformation with make(1), I'd probably want a widely-available, standard, simple command to make a rule-specific virtual filesystem that overlays a configurable read-only/copy-on-write view of selected existing files or directories behind a writable rule-specific output directory.

dub··on An ex-Googler's guide to dev tools (2020)
Being familiar with npm doesn't save the time it loses you on its clunkiness. I see for myself and other engineers around me that we waste real time every week clearing our yarn caches, waiting for "yarn install", realizing we forgot to run "run install" after syncing and we have to re-merge and re-upload, etc.

Here's a real-word example from last week: I spent several hours doing creative hackery trying to convince a particular compiler to compile the same file using different dependencies depending on context. That would have been a few minutes of work with bazel, but when your build process consists of "run this off the shelf compiler" and the compiler doesn't have any native support for building two different versions of the same thing with slightly different dependencies then you're in trouble.

Teaching new hires bazel is a one-time cost, and the only thing they'll need to get started on their first day is "bazel build <target>" and "bazel test <target>". When you don't have bazel, on your first day every new hire has to read through a gigantic wiki page explaining how to set up their dev environment (with different sections for mac and linux and special asides describing what to do if you're on a slightly outdated version of the OS, and then the wiki gets out of date and you have new engineers and the infra team all wasting their time debugging why the instructions suddenly stopped working for people on slightly newer machines, etc.)

dub··on An ex-Googler's guide to dev tools (2020)
Bazel can be clunky, but not having some bazel equivalent can have very significant costs that are easy to get accustomed to or overlook.

Things like engineers losing time wondering why their node dependencies weren't correctly installed, or dealing with a pre-commit check that reminds them they didn't manually regenerate the generated files, or having humans write machine-friendly configuration that's not actually human-friendly because there's no easy way to introduce custom file transformations during the build.

Bazel doesn't spark joy for me and I wouldn't say I look forward to using it, but personally I would still always choose it for a codebase that's going to have multiple developers and last a long time. It's vastly easier to go with bazel from the beginning than to wish you could switch to it and realize people have already introduced a million circular dependencies and it's going to be a multi-month or multi-year process to migrate to it.

dub··on An ex-Googler's guide to dev tools (2020)
For those of us stuck on AWS it's sad not having BigQuery, but the thing that really gets me is not having Dataflow

Most of industry still seems unaware that no-knobs data query and pipeline systems even exist. If I only had a dollar for every time I saw a PR tweaking the memory settings of some Spark job or hive query that stopped running as the input data grew....

I'd love to see more people write their workflows using the Apache Beam API so they'll have the option to switch to a no-knobs, scalable pipeline engine in the future even if they're not using one today.

dub··on PostgreSQL upgrades are hard
Hypothetically it should be possible to make an entirely new RDS cluster as a replica at a new version and fail over to it, with a similar error rate to a normal replica failover.

Setting up the infrastructure to manually manage your own cluster failover would kinda go against the spirit of using RDS and letting AWS manage infrastructure for you, though.

dub··on ZeroStableCoin is out from stealth mode
A challenge for the crypto world is that the values of being decentralized and trustworthy seem a bit at odds with the value of rapidly evolving.

If you compare the Amazon or Google of 1999 vs today, they're practically different businesses. Most well-funded, centralized businesses have a hard time evolving at all, let alone to that degree.

Seems like the challenges facing a decentralized organization wanting to evolve would be even larger.

Page 1 of 4Next →