HNHacker News
TopNewBestAskShowJobs

zbentley

6,040 karma · joined June 30, 2017

Zac Bentley

he/him

github.com/zbentley

blog.zacbentley.com

submissionscomments
zbentley··on Unknown number of Texas voter registrations went unprocessed due to DPS error
> Broken implementations were the desired outcome, because the constant doubt allows for things like disenfranchising people via voter roll purging.

I have trouble believing that, given that the "design", such as it is, took place over many decades of inter-agency regulatory/technical wrangling and organic growth.

At different times and points, the people that built what we think of as "the voting system" had totally different goals. Sometimes those goals were "make sure everyone eligible can vote as easily as possible"--and that's laudable! But more often the goals were "build $niche_system_component to comply with $specific_regulation". The number of times, during the design phase, where the people planning actually wanted to suppress votes and were empowered to make a part of the system actually do that was probably a rounding error compared to times when people fully accidentally moved the system in that direction due to bureaucratic myopia and shitty implementation.

I'm all for blaming political actors for things they intentionally do that are bad! Defunding the EPA? Bad! Getting into stupid wars? Bad! Banning mail-in voting? Bad! People with power and pre-meditated goals did those things, and many more, and should be opposed for it.

But "voting system is broken due to bureaucratic inter-operation failures, and state government staff have to do a postmortem and data recovery (ballot reprocessing)" isn't one of those times.

zbentley··on Unknown number of Texas voter registrations went unprocessed due to DPS error
I know this was mostly glib, but it really doesn't work that way. There aren't centralized data repositories that even can be joined, in many cases. Even when they could be (by batch/paper-form request/extremely slow system integration) joined conceptually, many decades of inertia, law, regulation, and political opposition often prevent those interconnections from existing. Even when they are permitted to exist, they're usually implemented in extremely broken ways.

And plenty of that headwind is not coming from the "party of small government"! The refrain of "the government should not have centralized data it can use to infringe on my rights" often transcends party, depending on what issue is at hand. That's sometimes a reasonable position to hold, but it definitely has a cost, especially when taken to extremes, or when make-citizen-services-work-better initiatives are framed as spying or whatnot.

zbentley··on Unknown number of Texas voter registrations went unprocessed due to DPS error
> Surely the government already knows who is a citizen

No. There's a lot more nuance there than you'd think, in the "broken regulatory and technical processes" domain, not political issues.

Example questions with unsatisfactory (not "impossible to do well", just "broken due to decades of technical and regulatory debt") answers: what government agencies, plural, are the authority of who is a citizen? Do they all agree? Do they provide ways for different agencies to check identities against their registries? If so, how do they match information in a citizenship-check request against their internal systems? By name? If so, how do they handle different name spellings (marriages, name changes, present/omitted middle names, non-ASCII)? By some other identifier? If so, how do they validate uniqueness of that identifier (SSNs are not secure/are widely leaked for tens of millions of people, addresses change and can have multiple residents, not everyone has a passport number, state-issued IDs/drivers licenses aren't stored centrally by the federal government)? If a check-requested identifier is invalid, can they tell whoever requested the check why it's invalid? Do these validation requests happen in real time/one at a time, or periodically in batches? Do the people sending the validation requests have any contact at all with the people providing information about why unsuccessful requests failed?

I've worked on these systems--human and technical. It's hard even for a functioning, outcome-focused, centralized government. In this problem domain, the US federal government is none of those things, and has not been for many decades (or ever, in some areas).

Sure, some of the issues are politically-caused (one party has an incentive to not fix problems that privilege their agenda), but most of them are accumulated failures with no malicious agenda as the cause.

zbentley··on Unknown number of Texas voter registrations went unprocessed due to DPS error
> The whole process feels like it’s designed to discourage people from voting.

Possible, but I doubt it. I've worked on the systems that do things like ID correlation/verification on the backend. They're ... imagine the most broken possible implementation you've ever worked on, then imagine it worse, then imagine that it's maintained by entirely different contractors than the ones that built it, with minimal incentive for improvement.

That's most of government IT implementations, and has been for decades.

Even when specific systems in this area are implemented well, the regulatory/bureaucratic complexity of the government prevents them from being more than marginally useful at best. I expanded more on some of the concerns/issues in this area in an adjacent comment here: https://news.ycombinator.com/item?id=49832393

zbentley··on Making Tailscale Faster
No idea if this is what they meant, but it’s plausible that client battery usage goes up when connected to Tailscale using an exit node. Suddenly, a lot more traffic has to be processed in userland, on the CPU. Ordinary internet transit might be offloaded to the kernel or hardware such that it’s more battery-efficient.

Then again, that begs the question of what you’re heavily using exit-node-routed internet links for that Tailscale’s battery draw is noticeable. Most network-intensive internet tasks draw way more power to drive whatever the application is, such that VPN overhead is a rounding error.

zbentley··on C++26: Trivial infinite loops are no longer undefined behaviour
Yes. Loops that manipulate data are extremely common. Combining memory accesses and increasing cache locality are extremely beneficial to performance. Transformations that rewrite loops to increase the chances of locality/combining are therefore likely to be worthwhile.
zbentley··on UTF-8000: Unlimited UTF-8
I think the DoS has pretty common amplification vectors in the form of APIs that split or otherwise copy (e.g. materializing code points for Unicode regex searching).

Additionally, OOM inside a low level routine can be a troublesome attack, since OOM handling in many applications does questionable (nee vulnerable) things when crashes occur in not-known-to-be-memory-intensive code. Sure, that’s sloppy engineering, but it’s common.

zbentley··on The scourge of x86 emulation
“Emulator” is a term of art in this area which refers to emulation of a hardware (and usually machine instruction) environment. Ironically, “translation” (as in instruction translation), as proposed by a sibling comment, is an even more connoted-with-hardware-emulation term.

WINE is … well, most directly it’s just an implementation of an API (the Windows APIs). In webdev parlance it might be called a “polyfill”. Perhaps a “compatibility shim”?

zbentley··on PyPy v8.0.0 Release
I do think Go and Java are strong examples of this. Those languages have native code support, but it’s clunky to use and, more importantly, very rare.

If you removed JNA/JNI/CGo, plenty of people would complain…but the vast majority of uses of those languages would still work, unaltered, because most common tasks on the JVM and Go runtime don’t require external native deps.

zbentley··on What Zig felt like, coming from Rust
> Arenas make lots of sense for video games, but their applicability for these other applications is much more dubious.

It's a fairly common pattern in non-video contexts to preallocate arenas of different sizes for different workload pools (e.g. different routes for a server).

It's also fairly common to allocate per-work-item (e.g. request/session/batch job) arenas for "general-purpose" scratch allocations that are small (your header parsing, auth context retrieval, etc.) and then default to a global manual/refcounting allocator for one-off large data actions like your holiday change. Since most business applications are returning summary/small aggregates over the large data action (in this example, something like "holiday updated successfully", not the entire history of the PTO table or the database connection's internal state), the copy cost of moving data between the global dynamic allocator and the arena for result transmission tends to be small.

That's not a terrible approach in some situations. If the large majority of code only needs the scratch arena and/or some per-handle allocators stored on e.g. the database connection pool, it can work well. But if, over time, the amount of bookkeeping required to maintain data tagging/movement between arenas/allocators becomes severe, it's worth stopping, stepping back, and considering that you've kind of walked backwards into inventing a shitty generational GC.

zbentley··on Warez: The Infrastructure and Aesthetics of Piracy (2021)
Does that extend to medical systems? If someone hacks my pacemaker and destabilizes my heart, or my pharmacy/hospital and prevents me from receiving medical care, is that something that should be legally permitted on the basis that "medical device makers/pharmacies/hospitals should try harder, and should somehow compensate their patients if a hack occurs"?

How do you compensate someone who's dead?

This isn't hypothetical: https://www.politico.com/news/2022/12/28/cyberattacks-u-s-ho...

zbentley··on I don't like passkeys
All true, but there's a wrinkle with delegating auth in computer systems--the same wrinkle that comes up when thinking about digital data as property/copyrightable etc.: when you delegate auth, you copy the access; you don't loan it. So every digital delegated-auth scenario is like your "make a copy of your house keys" example, not your "loan out your debit card" example.

If we extend the metaphor, this would be like making a copy of your keys and handing those out every time someone other than you needed access to your house. Dinner guest? Key copy. Neighbor dropping off a borrowed tool? Key copy. Relative from out of town visiting? Key copy.

In the same way that I think most homeowners would look askance at passing out so many copies of their keys, delegated digital auth is troublesome. Nontechnical users are unlikely to pay attention to "what clients are using delegated credentials for which actions" dashboards. Revocation, while technically easy, isn't something that I think most casual users will be mindful of, resulting in endless growth in the list of principals with access to a resource (just like the "sharing passwords" scenario we have now). Time-based auto-revocation will be an annoyance for delegates who only need to access a resource rarely, resulting in exasperated administrators rubber-stamping new-delegate-credentials requests.

I don't know if there's a good solve here. Shared passwords might be the local maximum of convenience and security, but that feels pretty bad.

zbentley··on Bend 2 and the Vibe-Coding Trap
> That's why all your LLM requests to build something substantial should start with "run prior work research first".

Yes, but I think there are incentives to not do this for many LLM providers. Doing prior-work research is slow (web searches aren't fast, LLMs are rate-limited or blocked from plenty of pages, etc.), and sometimes contradictory which annoys LLM users, many of whom like faster gratification cycles from the agent slot machine handle.

Also, writing a bunch of bespoke code instead of leveraging prior art makes a lot of users feel like they own something novel/big/important, and also poses a larger maintenance surface for the LLM to make future changes (which costs tokens).

I don't think there's, like, a conspiracy at LLM providers to set up system prompts/RAG/etc. to discourage research-and-use-prior-art-by-default approaches. Rather, OpenAI/Anthropic/Google/etc. are optimizing for real but sometimes misleading success metrics which often lead away from a research-first approach.

zbentley··on You can run Git on object storage if you re-make packfiles
I posted this because it seemed timely given that 'CGamesPlay and I were discussing this exact kind of system two weeks ago on the thread about why kernel.org's cgit (git web UI) hosting system was becoming expensive to operate due to LLM scraper load: https://news.ycombinator.com/item?id=49504674
zbentley··on GitHub is having trouble counting things
Ideally written as:

5) multit

6) database synchreadinghronization

zbentley··on Libraries Run Rust Inside Python (With PyO3)
I'm with you on the broader sentiment that Rust isn't making nearly as many inroads many people think.

But this example is silly. Do we say "FORTRAN is eating R" because R uses BLAS/LAPACK/whatever via gfortran? Nah. A zillion languages use LLVM.

zbentley··on The case against JPEG XL
I ... what? I've never heard anyone conflate the two--not teenagers, non-technical adults, or software people. Most folks know that "GIF" means "moving image that can be easily shared but has no sound and usually crap quality", not a video/audio clip.
zbentley··on Mullenweg has returned as CEO after attempted board ouster
Off topic, but the fact that that's a Wordpress link just makes it funnier.
zbentley··on Mullenweg has returned as CEO after attempted board ouster
> It would set a terrible precedent for them to do so unilaterally.

Speaking generally (not about Automattic), big SaaS have law enforcement desks that work on exactly this type of thing, and the "terrible precedent" is already widely set. Plenty of law enforcement outreach (which includes lawyers, courts, and actual law enforcement officials) results in pre-emptive compliance by SaaS companies. I would be massively surprised if Slack has not already done this in many cases, because most huge companies routinely do.

That's neither generally good nor generally bad; whether it's the right move depends on the charge, requested actions by law enforcement, status of legal proceedings, and the values/diligence by which the SaaS business assesses the legitimacy and likely cost/benefit of a law enforcement request. Note that "pre-emptive compliance" doesn't always mean an email saying "hey, the FBI said you suck so we terminated your account". There's a broad spectrum of tools available to a SaaS ranging from sending that email, to holding bespoke contract re-negotiations (which are functionally always in process between a SaaS and a huge customer) hostage to endless redlining rounds, to enforcing ToS violations that the SaaS previously turned a blind eye towards due to customer size.

> Whatever the legal process this battle follows, it will be a year or two before it's even possible for a final ruling + court order

Preliminary injunctions can be issued in days to weeks, not months to years, in all sorts of civil and criminal cases in all sorts of jurisdictions. Those can take the form of "don't change stuff with your admin access" or "grant admin control to someone else"-type orders. In cases where a service administrator is materially involved, injunctions are also easy to get on the basis of evidence preservation.

That's a pretty sharp tool. Failure to comply with those opens individuals and businesses up to way more legal penalties and tighter timeframes. Even if an injunction is later vacated/dismissed/modified, the legal argument that you violated it because you knew that would happen is an extremely tough sell.

zbentley··on Linux Zoom client proactively reading everything written to X11 clipboard
Narrowly, I don't think there's any technical reason why a better clipboard couldn't require the user to specify where to send text ("paste content into application X field Y") and then handle transmitting that data on the user's behalf. That way, the clipboard wouldn't be readable global state without intent.

The paste intent doesn't seem antithetical to accessibility. A keyboard combination states intent just as effectively as a voice command, mouse input, or other source.

zbentley··on Linux Zoom client proactively reading everything written to X11 clipboard
> How hard would it be to have something like this?

Very. Both because of the technical issues others in this thread mentioned, and because Android and iOS had the unutterably massive advantages of starting from zero pre-existing software and having a single controlling authority.

The controlling authority allowed them to avoid problems related to consensus or people working on other priorities. Starting from zero allowed them to not worry about existing software (the closest thing to a controlling authority that Linux has cares a lot about backwards compatibility).

zbentley··on Why is Google still serving dodgy ads?
Yep.

And I suspect (would love some real data from scammy advertisers) that this is a bigger factor than most folks appreciate.

If you're a regular business that sells clothes or food or SaaS or insurance or whatever, online ads are a) not your whole customer acquisition system (you have regular marketing, word of mouth, maybe sales people, maybe brick-and-mortar brand presence, physical advertising, etc.) and b) not part of your expense sheet that you want to spend tons of time on.

The "b)" there leads companies to either overspend on stupid (but unlikely to be flagged) ads, or to try to keep their ad spend as low as possible for the highest reward, because they think of it like opex. After all, they have all the normal expenses of a normal business to worry about in addition to ad spend. Sure, there are outliers here who spend almost all their non-product money on online ads, but not most ordinary companies from Verizon to Dave's Local Auto Insurance to Safeway.

So "good"/ordinary advertisers aren't generating super high flag rates, but they're probably either a) spending lots of money without churn/support-needs risk or b) aren't spending that much ... compared to scam-ads-only businesses. Sure, google would prefer more upstanding customers in category "a)", but they're a near-monopoly and some of that growth is expensive to chase, while the scammers are beating a path to their door.

If you're a scam-ad business that produces malware or LLM-written gacha games or whatever, your product expenses are minimal and, if your business is lucrative, you can spend very large amounts of your balance sheet on ads for scam delivery channels. Tiny businesses in this area are likely to have whale-sized ad spends and a strong likelihood of growth (vertical as they get bigger/greedier, or horizontal as they keep spinning up new businesses to keep ahead of the law and bad press). I bet they punch above their weight ad-spend-wise, in other words.

You thought the unit economics of a SaaS business were good? Well, wait until you see the unit economics of a Play Store whitelabeled gacha game app with a credential stealer inside!

zbentley··on Homebrew 7.0.0
True, but Python is the one with an env management system (virtual environments) which is the most prone to breakage for projects that depend on system Python.

Uv is far superior to both Mise and Homebrew for Python work, and I find that it removes the vast majority of pain preventing me from using Homebrew by default for most things, and Mise only occasionally for specific dev envs. Mise is a great tool though!

zbentley··on Nvidia is the central bank of AI
Related but in specific ways. Stock is often priced in anticipation of growth. If NVDA could meet its credit obligations while its real profit stayed flat, the two would diverge, at least for awhile. A large amount of NVDA’s current cash flow is likely purchase contracts with a fixed multi-year term, which further smooths out the impact of, say, a stock crash following a couple of quarters of terrible earnings.

Now, whether many things NVDA has invested in with expectation of repayment or earnings would be able to repay or appreciate in a market environment where Nvidia’s stock was crashing? That’s another question entirely.

zbentley··on Stop making swap partitions—use swap files instead
Anecdotal, but I've been happily using it under Linux desktops for years, and it works quite well. Workloads include: development, VM hosting, steam gaming, web browsing, multimedia playback. OSes include Debian, Proxmox, vanilla Arch, CachyOS, and others. Daily-driver hardware (ignoring servers and less general-purpose desktop stuff) included 2019 chromeboxes, 2015 (!) laptops, current-gen gaming laptops, and desktop towers with handfuls of spinning rust and solid state drives.

It seems to work well in a variety of situations: 4GB/single-slow-SSD ancient systems work just as well as spinning rust bulk storage pools with NVME ARC/ZIL caches for my gaming/server/database datasets, and all-SATA-SSD pools can get to near-NVME performance with bonus redundancy for boot volumes and latency-critical stuff. For personal desktop use, I haven't found dedup worth the squeeze in RAM costs and tuning (it works, but it's generally easier to solve most dedup-compatible problems at a layer closer to the cause).

ZFSBootMenu and the ability to roll back to snapshots and restore/maintenance disks from outside of the primary operating system, without having to think about fallback boot drives or physical backup volumes, is a godsend in the "try random sketchy commands that might trash my installation in order to get a low-level driver problem resolved" and "I could take the time to understand what this curl | bash invocation does, but I have better things to do; I want to be able to reverse it if it breaks stuff" departments.

In general, I strongly recommend ZFS for daily-driver use. Its core primitives are quite flexible, it makes redundancy/backups/drive addition/replacement easy, and it works fine on old and under-resourced systems; the mythos of "it requires ECC and enterprise-grade hardware and tons of RAM/CPU to work at all" was always bunk. The enterprise/SAN features are there if you want them, but are off by default, and the core FS capabilities are widely useful. Even casual desktop Linux users would do well to set it up, since there are a lot of rare-but-real ordinary user needs that, if they come up and you're not running something like ZFS, can't be done at all unless you connect purpose-specific hard drives or reinstall your OS.

Especially now that NVMEs are so expensive, ZFS should be considered for its ability to make RAIDing up a set of slower drives (or mostly slow drives with an NVME cache) very easy. That way, you can make your existing disks into something that performs well enough that you don't need to spend money on new hardware.

Just don't install it via DKMS; get a distro that ships it compiled into the kernel or as an installable kernel-paired module. Many such distros exist. The DKMS edition won't eat your data, but you'll get real tired of failed system updates because the kernel changed some source and the compile failed. That happens often; turns out that the volume of the kernel API surface used by something as massive as the ZFS codebase is quite large.

Edit: upon reading back through this, I'm a bit sheepish that I sound like such a breathless shill. I promise I'm not in the ~pocket~ zpool of big filesystem. I just like it.

zbentley··on New York thoracic surgeon: "For many patients 9/11 is not over"
GP didn't say that, though.

Agency isn't some magical axiomatic property where if you exercise it you're the sole contributing factor to your actions. Exercising agency, like making any other choice, is a decision whose inputs include your circumstances, history, etc.

Consider someone who quit smoking cigarettes in the 1950s. That was a time when cigarette smoking was common and not widely considered unhealthy. That person exercised agency. Now consider someone who quit in the 2020s, when cigarette smoking is widely warned against. That person exercised agency, too. More people quit smoking in the 2020s than the 1950s, by any measure. Does that mean people had less agency in the 1950s? No, it means the inputs to their decision-making on how to exercise their agency were different.

zbentley··on I've operated petabyte-scale ClickHouse clusters for 5 years
No apology needed; I understand what you're getting at, and I broadly agree. It's a spectrum between "extremely easy-to-hire people that operate in such a narrow niche that they're an operational liability with limited capabilities" and "expect everyone to be an expert at every level of the stack". The right point on that spectrum is different depending on context, but I do think that a majority of software shops would be better served by moving their required skillset more towards the generalist end of that spectrum, because the default is often far too niche (driven by poor tradeoffs and short-termism in service of growth/hiring, usually).

I wanna re-emphasize that this is not a new problem. It's not because of DevOps culture or cloud complexity or scale or whatever. Very limited-specialty people were always operational liabilities and had limited positive impact on feature delivery once you accounted for the help they needed to do anything that spanned multiple levels of the stack. There are just more engineers working on more systems with tighter timeline expectations now, so it seems like the complexity incumbent on the engineering role went up in general. It didn't (it went up in some situations and down in some situations), we just started noticing operational pain more often.

I definitely do agree that there's widespread ignorance of the velocity and difficulty-of-work tradeoffs that arise from requiring a wider range of specialties from engineers, and a similarly widespread failure to adjust compensation and timeline expectations accordingly.

zbentley··on I've operated petabyte-scale ClickHouse clusters for 5 years
I think a DBA/ops/infrastructure person as an imposed bottleneck is a useful capability in some environments.

But I won't follow you as far as "expecting developers to have expertise in how and where their software runs is unreasonable".

Like, yeah, it sucks that added DevOps responsibilities etc. don't come with adjusted compensation/time allocation expectations. I'm with you there.

But it's simultaneously true that a ton of "just regular developer" people are significant liabilities because they don't understand anything about the environment where their software runs. That liability manifests operationally (if someone's just running integration tests on Windows for their Java business logic changes and don't have any familiarity with e.g. the Linux, container, or cloud environments where their code runs, they're going to be useless when their code breaks in production and operations staff needs context), and it also makes them less effective when writing code--this culture of "developers should just live in business logic and not have to context-switch or fill their brains with other levels of the stack" is what leads to full table scans, lack of awareness of memory use, N+1 query hell, looping microservice dependencies, misunderstanding of what HTTP fields are set on requests that are mutated by load balancers, mistaken assumptions about how many instances of code can run and what concurrency/thread/coroutine behaviors are present, and so on. Those are very common problems, and it's incumbent on developers in every specialty to gain familiarity with how and where their code runs in order to write and maintain that code effectively.

If your code runs on Linux in Kubernetes, all of your developers should know how to read Linux system logs, check database sessions/queries issued by parts of the application, ls/grep/cat/strace/ps their way around, interpret k8s/application dashboards, check application logs both in log storage and as they're emitted from a process, exec into a container, restart pods, check deployment liveness, etc. Even if they don't have permission to do those things in production.

That was true in 2005 when they deployed their code to IIS on Windows Server/MSSQL, too--just with different operational specifics.

That's a low bar that's often unmet, and all sorts of teams suffer from that failure. Those skills can be trained, kept up to date, and hired for; I don't think there's a great excuse for not expecting them.

zbentley··on I've operated petabyte-scale ClickHouse clusters for 5 years
It's one thing to be able to scale database compute/storage; it's another thing to be able to partition it. It's extremely common for bad queries/access patterns to cause noisy-neighbor effects on other simultaneous accesses to the database, to the extreme of knocking the whole database over with timeouts/OOMs/etc.

Scaling out DB compute can only help with that to a (expensive) point; eventually, you end up wanting to either prevent the bad queries from being added to the system (DBA culture) or ensure that the bad query runs on database infrastructure that doesn't affect other queries. That's why partitioning DB compute (and storage: noisy-neighbor effects from a bad query at the storage layer don't require storage to be running e.g. a BookKeeper or whatever on a server; they can manifest as hot S3 keys or cloud object/block store rate limiting) is a necessary capability if your plan for dealing with a culture of "anyone can add any access pattern they want" is to scale the DB.

zbentley··on Proof of Capture: Apple Reference Image, but open source and using steganography
> a case where a normal person had genuine photos leaked and then just baldly denied everything

"Genuine" is doing a lot of work in that sentence. A big part of the threat model for image provenance/signing/similarity diffing is identifying when images aren't genuine--if they're from elsewhere than they're claimed to be from, or have been modified.

You're right that there are privacy/security costs to attributability, and that it's not always the right thing to do. I hope that keeping provenance information either entirely cryptographic in nature (okay, the image has a signature--you can't determine anything about that signature other than "signed with this key y/n" when you present a key) or reducing identifying or fingerprintable information presence in provenance metadata is sufficient to mitigate some of those concerns.

Dr. Neal Krawetz has written and researched a lot about this topic:

https://hackerfactor.com/blog/index.php?/archives/1069-The-B...

https://hackerfactor.com/blog/index.php?/archives/1098-Metas...

https://www.hackerfactor.com/blog/?/archives/529-Kind-of-Lik...

← PreviousPage 2 of 34Next →