HNHacker News
TopNewBestAskShowJobs

fdr

1,331 karma · joined October 18, 2010

submissionscomments
fdr··on Develop Cross-Platform CLI and GUI Tools with Tcl/Tk
A fun one: PostgreSQL uses an adopted version of Tcl's regular expression engine, written by Henry Spencer. Tom Lane seems to think it was an impressive piece of work, and patched it up from time to time.

https://en.wikipedia.org/wiki/Henry_Spencer https://en.wikipedia.org/wiki/Tom_Lane_%28computer_scientist...

fdr··on PostgreSQL and the OOM killer: Why we use strict memory overcommit
The key thing is Postgres does handle enomem well and does a nice rollback rather than crashing the server and entering crash recovery. It’s one of the few programs that does. Exceptions for the exception.

Even a revised heuristic that only spots large, individual allocations is not going to do the job.

Oom score adjust also doesn’t do the job: because the only interesting workload is Postgres, if a backend does a page fault that needs memory, who dies? Another sibling Postgres, almost certainly. Then postmaster does crash recovery, which most would rather avoid. High performance databases with distant checkpoints can take a while to come back up.

fdr··on Flat Datacenter Networks at Scale at Amazon
I always like these randomized/semi-randomized network papers. Here's a little known one you might enjoy if you liked this one. https://repositorio.unican.es/xmlui/handle/10902/23594
fdr··on Ian's Secure Shoelace Knot
I've been tying this for years. Good knot. I only have failures if I hit a snag (and no easy-release knot is going to be able to get around that)
fdr··on Simply Scheme: Introducing Computer Science (1999)
I enjoyed my work with Scheme, having received my instruction in the early 2000s. I'm not a functional language or lisp language advocate in any way, and I don't even dislike Python for professional work, but I do regret that it is not taught anymore: Python's management of scopes is not as good for the instruction.
fdr··on Silver plunges 30% in worst day since 1980, gold tumbles
Fun fact about silver, besides its heavy industrial footprint, which you mentioned: the supply is dominated by Mexico. There have been some, uh, erratic words about Mexico from the people in the position to affect trade policy and foreign policy.
fdr··on IPv6 is not insecure because it lacks a NAT
For those of you with this handy technology, the mobile phone, in the United States: you have an IPv6 address without NAT. Some of you even exist on a network using 464XLAT to tunnel IPv4 in IPV6, because it's a pure IPV6 network (T-Mobile). These mobile phone providers do not let the gazillion consumer smartphones act as servers for obvious reasons.

This is all to underscore the author's point: NAT may necessitate stateful tracking, but firewalls without translation has been deployed at massive scale for one of the most numerous types of device in existence.

fdr··on Gas Town Decoded
the main area I'd like to see some departure from beads is to use markdown files (or something) to be able to see the issue context/comments better in a diff generated by git.

The other area I'd like to see some software engineering thinking that's more open ended is on regression testing: ways of storing or referencing old versions of texts to see if the agent can complete old transformations properly even with a context change that patches up a weakness in a transformation that is desirable. This is tricky as it interacts with something essential in software engineering, the ability to run test suites and responding to the outcome. I don't think we know yet when to apply what fidelity of testing, e.g. one-shot on snippets versus a more realistic test based on git worktrees.

This is not something you'd want for every context, but a lot of my effort is spent building up prompt fragments to normalize and clean up the code coming out of a model that did some ad-hoc work that meets the test coverage bar, which constrains it decently into having achieved "something." Kind of like a prototype. But often, a lot of ungratifying massaging is required to even cover the annoying but not dangerous tics of the LLM, to bring clarity to where it wrote, well, very bad and unprincipled code...as it does sometimes.

fdr··on Gas Town Decoded
that's one reason I am less worried about him than some, although, I don't want to say that only to have something bad happen to him, that is, a form of complacency. Just because (say) Boltzmann and Cantor had useful insights along the way didn't mean people shouldn't have been looking to support them.
fdr··on Gas Town Decoded
I use beads quite a bit, but not as steve intended. And definitely the opposite of "Gas Town," where I use the note-taking capability and integration with Git (that is, as something of a glorified Makefile and database) to debug contexts, to close the loop and increase accuracy over time. Nevertheless, it has been useful for large batch runs over my code base: the record has been processing for thirty hours straight while getting something useful, and enough trace data to make further improvements.

Steve has gone "a bit" loopy, in a (so far) self aware manner, but he has some kind of insight into the software engineering process, I think. Yet, I predict beads will break under the weight of no-supervision eventually if he keeps churning it, but some others will pick up where he left off, with more modest goals. He did, to his credit, kill off several generations of project before this one in a similar category.

fdr··on S&P500 Priced in Gold
That's a lot of financial devices painted with a broad brush, and I think the charge that so may central banks are knuckled under with fiscal dominance is simply not sustainable. The ones that are, we tend to hear about.

Because there's a lot one could write about each of: equities, real estate, gold, silver, platinum (which have very different industrial exposures), and bitcoin, which have many price drivers.

So let's try something more parsimonious: what do you make of people, institutions, etc that bid on short and even long-dated sovereign debt around the globe, and come up the collective discovered price of, say...3.5%, annualized, for maturity in a month? https://www.treasurydirect.gov/auctions/announcements-data-r...

fdr··on S&P500 Priced in Gold
You don't even need to trust CPI alone when looking in history, where things have evened out a bit: we have historical short-term bond yield data, even the yield curve: people bidding on short periods with the safest debtor expecting changes in nominal value.

Not to suggest CPI is redundant, there's a reason why central bankers read it after all. For one, it's the most timely data they have. But it's impossible to nudge it year after year -- accumulative error -- without it become obviously decoupled from other data, including the long-term bond market data. It just so happens commodities are the wrong yardstick.

fdr··on S&P500 Priced in Gold
It's not very convincing, though: there's a huge runup in gold prices (as is often the case) between 2023 and the present, and a long do-nothing period before that (also often the case). The major consumers of gold are about: 50% jewelry, 10% industrial, 20% central banks, a large run-up from about 10% in the 2010s.

I like to think about the inherent contradictions of goldbugs going long on central bank portfolio policy: they both tend to distrust the central bank but in a way the central bank activities partially endorse their habits, and are the source of recent appreciation and thus accusations of "hidden" inflation. But central banks operate in an anarchic world system where they need something even independent of reserves held in other sovereign currencies, I presume most gold bugs are holding ETFs in an existing financial system (which is non-orthogonal: if you assume a financial system, why not avail yourself of the superior alternatives?) or have it in a safe in their house which has some other obvious problems.

I hold no gold, if I want hydraulic and non-volatile inflation compensation, it's quite simple: short-dated sovereign debt, aka the humble money market fund, which can be seen as the lower-fee version of the checking account. Nobody likes being a sucker, holding debt for below the time value of money, including changes in nominal value. It has immense price discovery pressure, and it finds its level nicely. If I were to hold gold, I would need some viable theory about how much I should hold to be de-correlated from other assets to be worthwhile. Maybe if I was exposed to jewelry costs and wanted to hedge them.

See https://www.jpmorgan.com/insights/markets-and-economy/market..., https://www.ecb.europa.eu/press/other-publications/ire/focus...

fdr··on Times New American: A Tale of Two Fonts
Public Sans seems like a good candidate for a new "web safe" font. Perhaps one new web safe font per twenty five years is not too much. From there, it can percolate to the word processor and pdfs, and finally: government standard for government workers who just want to open their word processor and get to work, where sourcing even a free font to meet standard is just a snag to annoy.
fdr··on OpenAI's cash burn will be one of the big bubble questions of 2026
The biggest run classified nuclear stockpile loads, at least in the US. They cost about half a billion apiece. And are 30 (carefully cooled and cabled) megawatts. https://en.wikipedia.org/wiki/El_Capitan_(supercomputer)

No chance they're going to take risks to share that hardware with anyone given what it does.

The scaled down version of El Capitan is used for non-classified workloads, some of which are proprietary, like drug simulation. It is called Tuolumne. Not long ago, it was nevertheless still a top ten supercomputer.

Like OP, I also don't see why a government supercomputer does it better than hyperscalers, coreweave, neoclouds, et al, who have put in a ton of capital as even compared to government. For loads where institutional continuity is extremely important, like weather -- and maybe one day, a public LLM model or three -- maybe. But we're not there yet, and there's so much competition in LLM infrastructure that it's quite likely some of these entrants will be bag holders, not a world of juicy margins at all...rather, playing chicken with negative gross margins.

fdr··on Cursed Knowledge
that also popped out at me: binding that many parameters is cursed. You really gotta use COPY (in most cases).

I'll give you a real cursed Postgres one: prepared statement names are silently truncated to NAMEDATALEN-1. NAMEDATALEN is 64. This goes back to 2001...or rather, that's when NAMEDATALEN was increased in size from 32. The truncation behavior itself is older still. It's something ORMs need to know about it -- few humans are preparing statement names of sixty-plus characters.

fdr··on A.I. researchers are negotiating $250M pay packages
I think it's pretty funny because, for example, Katalin Karikó was thought to be working in some backwater, on this "mRNA" thing, that could barely get published before COVID...and, the original LLM/transformer people were well qualified but not pulling quarter billion dollars kicking around trying to improve machine translation of languages, a time-honored AI endeavor going back to the 1950s. The came upon something with outstanding empirical properties.

For whatever reason, remuneration seems more concentrated than fundamentals. I don't begrudge those involved their good luck, though: I've had more than my fair share of good luck in my life, it wouldn't be me with the standing to complain.

fdr··on The world could run on older hardware if software optimization was a priority
One of the things I think about sometimes, a specific example rather than a rebuttal to Carmack.

The Electron Application is somewhere between tolerated and reviled by consumers, often on grounds of performance, but it's probably the single innovation that made using my Linux laptop in the workplace tractable. And it is genuinely useful to, for example, drop into a MS Teams meeting without installing.

So, everyone laments that nothing is as tightly coded as Winamp anymore, without remembering the first three characters.

fdr··on I use zip bombs to protect my server
Seems like an exponential backoff rule would do the job: I'm sure crashes happen for all sorts of reasons, some of which are bugs in the bot, even on non-adversarial input.
fdr··on Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode
GP is more or less correct.

Building and owning an institution that finances, racks, services, networks, and disposes of servers, both takes time and increases the commitment level. Hetzner is month to month, with a fixed overhead for fresh leasing of servers: the set-up fee.

This is a lot to administer when also building a software institution, and a business. It was not certain at the outset, for example, that the GitHub Actions Runner product would be as popular as it became. In its earliest form, it was partially an engineering test for our virtual machines, and we went around asking friendly contacts that we knew would report abnormalities to use it. There's another universe where it only went as far as an engineering test, and our utilization and revenue pattern (that is, utility to other people) is different.

fdr··on Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode
Ubicloud does not have an OpenStack dependency.
fdr··on Debugging Hetzner: Uncovering failures with powerstat, sensors, and dmidecode
It varies by system. As the legendary (to some) Kelly Johnson of the Skunk Works had as one of his main rules:

> The inspection system as currently used by the Skunk Works, which has been approved by both the Air Force and the Navy, meets the intent of existing military requirements and should be used on new projects. Push more basic inspection responsibility back to the subcontractors and vendors. Don't duplicate so much inspection.

But this will be the only and last time Ubicloud does not burn in a new model, or even tranches of purchases (I also work there...and am a founder).

fdr··on IPv6 Is Hard
There is such a prefix, though, but the problem is the end user devices themselves (or a few applications) are not always modern enough to have decent operation with a IPv6 stack. Less so these days, though. See https://en.wikipedia.org/wiki/IPv6_transition_mechanism, ::ffff:0:0:0/96.

That said, a lot of posts here don't seem to reckon with the fact that a slim majority of www.google.com connections in the United States are via IPv6, and a super-majority from India, Germany, and France. Comcast, T-Mobile, Verizon, as far as I have experienced, these all default to IPv6. While dropping IPv4 support is both a worthy, distant goal and sometimes used in goal-post moving rhetoric, it's not like nobody uses IPv6...rather, mobile broadband networks have depended on it for over a decade (see T-Mobile's deployment of 464XLAT)

fdr··on Cloud Virtualization: Red Hat, AWS Firecracker, and Ubicloud internals
The Ruby we write is quite strange by the normal reckoning. In particular, it has 100% branch coverage. Being able to do this affordably is one reason I use it.

A fair number of the dependencies we have also have 100% branch coverage, because I copied the practice, starting about ten years ago, from Jeremy Evans, who maintains a huge number of libraries under that principle. That includes "Sequel," the ORM that I've used for many years and originally copied the practice from, around 2015. You can see the libraries he maintains in this way: http://code.jeremyevans.net/ruby.html. He has joined Ubicloud somewhat recently, so I look forward to getting a sense of how he completes the rest of his rather singular & extraordinary maintenance regime.

To have Ubicloud rest at this standard is my objective. My tendentious claim is as follows: this is higher than any other constellation of libraries I have seen in any programming language. If anyone knows of any constellation of libraries that is more capably and rigorously maintained in any language, let me know. The bar as roughly as follows they need to release every month, or something like that (yes really: https://rubygems.org/gems/sequel/versions/), have a wide interface with your program, and break it no more than once every five years.

It's also not a Rails program, and I have never written or maintained a Rails program in any seriousness, which makes me an odd duck among longtime Ruby programmers.

fdr··on 13 Years of Building Infrastructure Control Planes in Ruby
Sure, it reduces costs quite dramatically to be able to do stuff like this:

    upgrade_check_ssh = ->(vmh) do
      p [vmh.ubid, vmh.created_at, vmh.sshable.host]
      vmh.sshable.cmd(<<BASH)
    set -xeuo pipefail
    sudo apt-get update -qq && sudo apt -qq -y satisfy 'openssh-server (>= 1:8.9p1-3ubuntu0.10)' && sudo systemctl restart ssh.service
    BASH
      vmh.sshable.cmd(<<VERIFY)
    set -xeuo pipefail
    dpkg-query --showformat='${Version}\n' --show openssh-server
    ssh_pid="$(systemctl show -p MainPID ssh.service | cut -d= -f2)"
    (set +e && sudo grep -F deleted "/proc/$ssh_pid/maps" ; [ $? -eq 1 ])
    VERIFY
    end
    
    cohort_draining = VmHost.where(allocation_state: 'draining').order_by(:created_at)
    
    cohort_draining.map { upgrade_check_ssh.call(_1).tap { sleep 3 } }
This is me upgrading OpenSSH on July 1st to account for the RCEs reported at that time on some low impact servers.

I then wrote many minor variants, to change the cohort (eventually targeting all servers), as well as a verification pass. The methodology and output is recorded, along with the time, in Slack. That's how I'm able to roll the tape for you now with precision, almost two months later, in late August. This kind of precision in recall and methodology is important for efficient operations...especially when things go wrong. A common thing we do, upon seeing, say, a broken VM Host, is paste its identifier into slack, to see if it's something of a troublemaker. From people's other code-and-output pastes, we can see what they ascertained, and how, and what was done.

I would not consider a language without a robust REPL for this kind of work. It is connected with an integrated develop-operate model, where the people writing the programs in these symbols every day are also assaying the problems. This unification is key.

And, somewhat related to that, I have not seen JVM nor BEAM libraries as high quality as Sequel, Roda, and Rodauth in their respective functions, and roughly in that order of importance, descending. These dependencies are invasive to how my code is written: above, you see some Sequel. We rely on other libraries being high quality (e.g. the pg driver gem, or net-ssh), but they are less invasive in this crucial way.

I did, at various points, consider applying this methodology to Python (the grammer's whitespace sensitivity is a serious problem, consider "cpaste"), TypeScript, Elixir, Scala, Julia, and even Swift. Although these rather conspicuously have REPLs, none have a Sequel.

I think people could make other REPL-enabled choices that work for them. But in my evaluation, some of the features of these runtimes did not overcome the consideration of a handful of key libraries.

fdr··on Unconditional Cash Study: first findings available
UBI is generally not metered by "%" but some flat quantity of money, whether nominal or real. That is, like a "head tax," but...negative.

In that common formulation, it would compress consumption by the entire tax+benefit base, that is, everyone would move towards median consumption by some amount, keyed to the magnitude of the UBI, if funded by any kind of proportional taxation (including a nominally regressive proportional tax, like consumption tax/VAT).

Politically, it has tough problems: 18% of the population [over age 65] already has a "MeBI" in the form of Social Security that they can vote to increase, and 22% of the population is below the age of 18, and can't vote. So that's 40% right there. Of the remaining 60% in their working years that produce the output split among themselves and that 40%, quite a few would rather not be compressed towards median consumption: the voting population is shifted higher in the consumption deciles, and people are not often so disposed to think they might find themselves luckless in the future. There's a thicket of "tax expenditures" that can form a "MeBI" for the electorate at the upper-half, like the mortgage interest deduction.

If we look at the difficulty in gaining electoral support in splitting consumption to the benefit of minors (thus, future labor) to even things out a bit, in the form of the semi-recently expired expanded child tax credit, we see the magnitude of the political problem.

Personally, I prefer to see UBI as tax reform to avoid crazy wiggling in effective marginal tax rate. But there are many reasons why it's unlikely that the electorate would see it that way, or approve of it even if they did.

fdr··on Unconditional Cash Study: first findings available
You are looking for the term "effective marginal tax rate." https://en.wikipedia.org/wiki/Effective_marginal_tax_rate

To give a sense how much benefits code and tax code have in common, see this worksheet for SNAP eligibility, which resembles a second tax return: https://www.fns.usda.gov/snap/recipient/eligibility. You get to do something similar, again(!), for Medicaid.

The American benefits code is a patchwork of conflicting sensibilities of the electorate: the smallest possible tax, paternalism and suspicion against the poor, plus a few policy analysis trying to obtain the maximum poverty reduction within those constraints. The result is a thicket of means tested programs with extremely steep phase-outs and a lot of paperwork. The all-in EMTR for an American with income between 0-40K a year is chaotic beyond reason as a result as they roll up the income spectrum.

This person who gave the presentation is indeed in one of the worst cases for the code: a single parent with multiple children.

fdr··on Doomsday Prepping: Reactionary Behavior or Inherited Instinct?
I think of it as something of a hobby. A little bit camping-adjacent. But just a little bit.
fdr··on A write-ahead log is not a universal part of durability
yeah, that one is a path dependency issue, re: postgres. I'm not sure if it's well reasoned past inertia after the initial "it's new, let's not throw the entire user base into it at once", but at the very least, it will complicate pg_upgrade somewhat. WiredTiger didn't really have that path dependency, being itself new to Mongo to shore up the storage situation.

It's probably about time to swap the default, but some people's once-working pg_upgrade programs that they haven't looked at in a while might break. Probably okay; those things need to happen...once in a while. I suppose some people that resent the overhead of Postgres checksumming atop their ZFS/btrfs/dm-integrity/whatever stacks, but they are somewhat rarer.

fdr··on Large Language Models are not a search engine
I do like Perplexity.ai. But interpreting how it works, and portrays itself as, is that the LLM component of it is, in fact, not a search engine.

How I interpret it: it is a more powerful version of stemming and synonym expansion of information retrieval classics when generating the queries it feeds into traditional information systems (such as the Bing search engine via API, or other index).

After retrieval, it's a selector and summarizer of repetition seen in the results to give you something of a blended outcome, pertinent to the prompt you gave it. Like any other tool, you get a feel for when it has is having problems, and some of those problems can be assessed by at least glancing at the sources it consulted. You get all sorts of weird stuff when your sources don't include relevant results or biased results

The first problem happens when the documents you are searching for do not exist, or something about your prompt -- it's usually obvious what it is -- is not sourcing documents you know to exist.

The second, bias, I've seen when researching something like the design conceits of Infiniband. While it has its genuine virtues, almost nobody talks about it...and many of those things that discuss it are Infiniband marketing materials that are both a bit too fluffy and sometimes stretch the truth, as marketing materials are wont to do. But you can spot this in the sources panel immediately.

I never found "disembodied" LLMs very useful.

Page 1 of 14Next →