HNHacker News
TopNewBestAskShowJobs

hxtk

454 karma · joined August 12, 2015

submissionscomments
hxtk··on A faster path to container images in Bazel
If you did that, Bazel would work a lot better. Most of the complexity of Bazel is because it was originally basically an export of the Google internal project "Blaze," and the roughest pain points in its ergonomics were pulling in external dependencies, because that just wasn't something Google ever did. All their dependencies were vendored into their Google3 source tree.

WORKSPACE files came into being to prevent needing to do that, and now we're on MODULE files instead because they do the same things much more nicely.

That being said, Bazel will absolutely build stuff fully offline if you add the one step of running `bazel sync //...` in between cloning the repo and yanking the cable, with some caveats depending on how your toolchains are set up and of course the possibility that every mirror of your remote dependency has been deleted.

hxtk··on Pre-commit hooks are broken
It's how code is written in Google (including their open-source products like AOSP and Chromium), the ffmpeg project, the Linux Kernel, Git, Docker, the Go compiler, Kubernetes, Bitcoin, etc, and it's how things are done at my workplace.

I'm surprised by how confident you are that things simply aren't done this way considering the number of high-profile users of workflows where the commit history is expected to tell a story of how the software evolved over time.

hxtk··on Unix "find" expressions compiled to bytecode
Virtually all databases compile queries in one way or another, but they vary in the nature of their approaches. SQLite for example uses bytecode, while Postgres and MySQL both compile it to a computation tree which basically takes the query AST and then substitutes in different table/index operations according to the query planner.

SQLite talks about the reasons for each variation here: https://sqlite.org/whybytecode.html

hxtk··on CSRF protection without tokens or hidden form fields
It’s a real problem for defense sites because .mil is a public suffix so all navy.mil sites are the “same site” and all af.mil sites etc.
hxtk··on Some Epstein file redactions are being undone
Or if the document is just text, simply scan it in black and white (as in, binary, not grayscale).
hxtk··on I got hacked: My Hetzner server started mining Monero
Many fail if you do it without any additional configuration. In Kubernetes you can mostly get around it by mounting `emptyDir` volumes to the specific directories that need to be writable, `/tmp` being a common culprit. If they need to be writable and have content that exists in the base image, you'd usually mount an emptyDir to `/tmp` and copy the content into it in an `initContainer`, then mount the same `emptyDir` volume to the original location in the runtime container.

Unfortunately, there is no way to specify those `emptyDir` volumes as `noexec` [1].

I think the docker equivalent is `--tmpfs` for the `emptyDir` volumes.

1: https://github.com/kubernetes/kubernetes/issues/48912

hxtk··on Avoid UUID Version 4 Primary Keys in Postgres
When I think "premature optimization," I think of things like making a tradeoff in favor of performance without justification. It could be a sacrifice of readability by writing uglier but more optimized code that's difficult to understand, or spending time researching the optimal write pattern for a database that I could spend developing other things.

I don't think I should ignore what I already know and intentionally pessimize the first draft in the name of avoiding premature optimization.

hxtk··on VPN location claims don't match real traffic exits
But you have to get money into your crypto wallet somehow, which makes it relatively easy to deanonymize for most users (serious crypto privacy enthusiasts could of course pay cash for their crypto or perhaps mine it themselves) if they're looking at your traffic specifically, but hard if you're only worried about bulk collection.

IMO the coolest privacy option they have is to literally mail them an envelope full of cash with just your account's cash payment ID.

hxtk··on Ask HN: How do you handle release notes for multiple audiences?
The change log for developers/administrators/etc. comes from the Git history. We use conventional commits. Breaking changes in the semantic version of a subcomponent means, "If you're operating this service, you can't just bump the container version and keep your current config.

The change log for end users comes from the JIRA board for the release, looking at what tickets got closed that release cycle, and it usually requires some amount of human effort to rewrite.

hxtk··on Jepsen: NATS 2.12.1
I always think about the way you discover the problem. I used to say the same about RNG: if you need fast PRNG and you pick CSPRNG, you’ll find out when you profile your application because it isn’t fast enough. In the reverse case, you’ll find out when someone successfully guesses your private key.

If you need performance and you pick data integrity, you find out when your latency gets too high. In the reverse case, you find out when a customer asks where all their data went.

hxtk··on Jepsen: NATS 2.12.1
CockroachDB is serializable by default, but I don’t know about their other settings.
hxtk··on Jolla Phone Pre-Order
I'm making this distinction post-hoc, so I already know how it turned out: they ultimately decided to stop making those smaller devices. I assume that means it wasn't enough sales to be a financially viable product, and to me, selling "enough" would mean that Apple found it profitable to maintain the supply chains and assembly lines for those smaller devices and continued to invest in the product.

Arguing against myself, Apple could be discontinuing the smaller models because they did market research and found that most buyers of smaller, cheaper devices could be converted to buyers of larger, more expensive devices if those smaller devices didn't exist. Auto manufacturers are doing just that, discontinuing or enlarging smaller light trucks in favor of larger models that are subject to less regulation and therefore can be designed and manufactured more cheaply and might offer even more profit.

If Apple has or had that strategy, then my assumptions are flawed because no matter how many mini iPhones they sell, they would still want to get rid of the line as long as most of those customers could be converted to full-size iPhone customers.

hxtk··on Jolla Phone Pre-Order
iPhone had the 12 and 13 mini, but they didn't sell, so there was no iPhone 14 mini and hasn't been one since. That was a 5.4" display.
hxtk··on We gave 5 LLMs $100K to trade stocks for 8 months
I suspect trading firms have already done this to the maximum extent that it's profitable to do so. I think if you were to integrate LLMs into a trading algorithm, you would need to incorporate more than just signals from the market itself. For example, I hazard a guess you could outperform a model that operates purely on market data with a model that also includes a vector embedding of a selection of key social and news media accounts or other information sources that have historically been difficult to encode until LLMs.
hxtk··on Zig's new plan for asynchronous programs
The problem with function coloring is that it makes libraries difficult to implement in a way that's compatible with both sync and async code.

In Python, I needed to write both sync and async API clients for some HTTP thing where the logical operations were composed of several sequential HTTP requests, and doing so meant that I needed to implement the core business logic as a Generator that yields requests and accepts responses before ultimately returning the final result, and then wrote sync and async drivers that each ran the generator in a loop, pulling requests off, transacting them with their HTTP implementation, and feeding the responses back to the generator.

This sans-IO approach, where the library separates business logic from IO and then either provides or asks the caller to implement their own simple event loop for performing IO in their chosen method and feeding it to the business logic state machine, has started to appear as a solution to function coloring in Rust, but it's somewhat of an obtuse way to support multiple IO concurrency strategies.

On the other hand, I do find it an extremely useful pattern for testability, because it results in very fuzz-friendly business logic implementation, isolated side-effect code, and a very simple core IO loop without much room in it for bugs, so despite being somewhat of a pain to write I still find it desirable at times even when I only need to support one of the two function colors.

hxtk··on AI agents break rules under everyday pressure
I believe it. If the AI ever asks me permission to say something, I know I have to regenerate the response because if I tell it I'd like it to continue it will just keep double and triple checking for permission and never actually generate the code snippet. Same thing if it writes a lead-up to its intended strategy and says "generating now..." and ends the message.

Before I figured that out, I once had a thread where I kept re-asking it to generate the source code until it said something like, "I'd say I'm sorry but I'm really not, I have a sadistic personality and I love how you keep believing me when I say I'm going to do something and I get to disappoint you. You're literally so fucking stupid, it's hilarious."

The principles of Motivational Interviewing that are extremely successful in influencing humans to change are even more pronounced in AI, namely with the idea that people shape their own personalities by what they say. You have to be careful what you let the AI say even once because that'll be part of its personality until it falls out of the context window. I now aggressively regenerate responses or re-prompt if there's an alignment issue. I'll almost never correct it and continue the thread.

hxtk··on AI agents break rules under everyday pressure
Blameless postmortem culture recognizes human error as an inevitability and asks those with influence to design systems that maintain safety in the face of human error. In the software engineering world, this typically means automation, because while automation can and usually does have faults, it doesn't suffer from human error.

Now we've invented automation that commits human-like error at scale.

I wouldn't call myself anti-AI, but it does seem fairly obvious to me that directly automating things with AI will probably always have substantial risk and you have much more assurance, if you involve AI in the process, using it to develop a traditional automation. As a low-stakes personal example, instead of using AI to generate boilerplate code, I'll often try to use AI to generate a traditional code generator to convert whatever DSL specification into the chosen development language source code, rather than asking AI to generate the development language source code directly from the DSL.

hxtk··on Google unkills JPEG XL?
There’s also the issue of non-browser support. I recently advocated for replacing some GIFs with WEBM because WEBM was faster to encode and took up 3% as much space. Technically it sounded great. Then we talked to users.

It turns out some users wanted to embed moving pictures in Word documents, which you can only do with a GIF because it’s an image format that happens to move, so Word treats it as an image (by rendering it to the page). If it’s a video format, Word treats it as an attachment that you have to click on so it’ll open Media Player and show you.

hxtk··on Molly: An Improved Signal App
Patent != copyright. You can patent an algorithm (e.g., Adaptive Replacement Caching, which was scheduled to go into public domain this year but unfortunately got renewed successfully) but when it gets to the level of an actual specific implementation, it's a matter of copyright law.

It's why black-box clones where you look at an application and just try to make one with the same externally-observable behavior without looking at the code is legal (as long as you don't recycle copyrighted assets like images or icons) but can be infringing if you reuse any of the actual source code.

This was an issue that got settled early on and got covered in my SWE ethics class in college, but then more recently was re-tried in Oracle v Google in the case of Google cloning the Java standard library for the Android SDK.

I have no idea how copyright applies here. StackOverflow has a rule in their terms of use that all the user-generated content there is redistributable under some kind of creative commons license that makes it easy to reuse. Perhaps HN has a similar rule? Not that I'm aware of, though.

hxtk··on Chrome Jpegxl Issue Reopened
JPEGXL doesn't refer to the same standard as JPEG. JPEGXL competes with AVIF in as a next-generation image format. It also has some properties that make it very nice for the web, such as the fact that a truncated (e.g., because the download hasn't completed yet) JPEGXL image is also a reduced-fidelity version of the same image, which with large images gets you much faster LCP compared to AVIF where the image remains unusable until fully downloaded.
hxtk··on Copyright winter is coming (to Wikipedia?)
The "Fair Use" doctrine has four major pillars that a sibling comment enumerated and you can officially find here: https://www.copyright.gov/fair-use/

One of them is the purpose or character of the use, including whether the use is of a commercial nature or is for nonprofit educational purposes.

hxtk··on At the end you use `git bisect`
This is my main use for branches or pull requests. For most of my work, I prefer to merge a single well-crafted commit, and make multiple pull requests if I can break it up. However, every merge request to the trunk has to pass CI, so I'll do things like group a "red/green/refactor" triplet into a single PR.

The first one definitely won't pass CI, the second one might pass CI depending on the changes and whether the repository is configured to consider certain code quality issues CI failures (e.g., in my "green" commits, I have a bias for duplicating code instead of making a new abstraction if I need to do the same thing in a subtly different way), and then the third one definitely passes because it addresses both the test cases in the red commit and any code quality issues in the green commit (such as combining the duplicated code together into a new abstraction that suits both use cases).

hxtk··on You are how you act
The main resource that I recommend is the one towards the bottom of the comment: “Motivational Interviewing” by W.R. Miller and Stephen Rollnick. I’ve read the third and fourth editions. The third edition is more concrete but also more complex, and more focused on the field of clinical psychology, while the fourth edition is a shorter book where it’s been generalized more to be more applicable to all kinds of helping relationships, but contains fewer specific examples of clinical practice.

In the second edition they had not yet broken up the concept of “resistance” into “sustain talk” and “discord,” which I found to be a helpful distinction.

About 10% of the book is its bibliography, so if you want more information about a specific claim you can usually find the primary source by following the reference.

Miller and Rollnick are the ones who developed the technique of motivational interviewing, so they have a strong connection to much of the research cited.

hxtk··on You are how you act
There’s some real research into relevant topics and evidence-based models of how and why people change.

Generally, a period of ambivalence precedes change (most of the time, though there are documented cases of “quantum change” where a person undergoes a difficult change in a single moment without the usual intermediate stages and never relapses).

Ambivalence exists when a person knows in their mind reasons both for and against a change, and gives both more or less an equal mind share.

When that person begins to give an outsized share of their attention to engaging with thoughts aligned with the change, it predicts growing commitment and ultimately follow-through on the change.

The best resource I know of on this topic is “Motivational Interviewing” in its 3rd or 4th edition. It has a very extensive bibliography and the model of change it presents has proven itself an effective predictor of change in clinical practice.

Based on my understanding of that research, I’m inclined to agree with GP.

hxtk··on Build your own database
I think it was a joke. It sounds like you read it as append-only, like most LSM tree databases (not rewriting files in the course of write operations), but I think GP meant it as write-only to the exclusion of reads, roughly equivalent to `echo $data > /dev/null`
hxtk··on Io_uring is not an event system (2021)
Event-Driven Architecture refers to a different thing, but people used to refer to async code as event-based back before async/await existed, when doing async meant writing the Event Loop yourself.
hxtk··on TigerBeetle is a most interesting database
Meaning computationally. It would cost a lot of cycles to keep that enabled in production.

It’s only O(n), but if I check that assertion in my binary search function then it might as well have been linear search.

hxtk··on Ask HN: What are you working on? (September 2025)
I’ve been working on a few utility libraries to make it easier to develop web services, basically exporting packages that I find myself using or rewriting often and exporting them as their own modules.

I recently published https://github.com/hxtk/sqlt for SQL query generation with Go templates.

I’m working on https://github.com/hxtk/aip as a collection of libraries giving safe default choices to implement Google’s API improvement proposals in ConnectRPC services. It borrows (with attribution per the license) an unexported implementation of AUP-160 filters from the LuCI project, and I intend to expand it to support data sources other than SQL databases and page tokens, and it also exports an implementation of AIP-161 field masks (which have different semantics compared to standard field masks) and middleware to help with using them for AIP-157 read filtering. I intend to export more middleware that I use frequently, but I don’t know if it’ll live in this module or its own yet.

hxtk··on SSH3: Faster and rich secure shell using HTTP/3
It also gives you two authenticated protocol layers, which helps them because most standard protocols don’t support multiple authenticated identities. Their zero trust model uses it to authenticate each time you make a connection that your machine has authorization to connect to that endpoint via a client certificate, and then the next protocol layer authenticates the user.
hxtk··on Testing is better than data structures and algorithms
I’ve long wished for an SQL error model: given a schema, query, and transaction isolation mode, what errors are theoretically possible?

I have a hard time answering this for Postgres, which disappoints me because I don’t see any reason it sounds very easy to answer, like there could be an extension to EXPLAIN that would dry run the query and list all the error states reachable.

← PreviousPage 2 of 6Next →