HNHacker News
TopNewBestAskShowJobs

electroly

5,206 karma · joined September 23, 2014

he/him

brian@electroly.com

https://github.com/electroly

submissionscomments
electroly··on Sparse File LRU Cache
I think I did a poor job of explaining. SQLite is dealing with cached filesystem blocks here, and has nothing to do with their query engine. They aren't migrating their query engine to SQLite, they're migrating their sparse file cache to SQLite. The SQLite blobs will be holding ranges of RocksDB file data.

RocksDB has a pluggable filesystem layer (similar to SQLite virtual filesystems), they can read blocks from the SQLite cache layer directly without needing to fake a RocksDB file at all. This is how my solution (I've implemented this before) works. Mine is SQLite both places: one SQLite file (normal) holds cached blocks and another SQLite file (with virtual filesystem) runs queries against the cache layer. They can do this with SQLite holding the cache and RocksDB running the queries.

IMO, a little more effort would have given them a better solution.

electroly··on Sparse File LRU Cache
I simply use SQLite for this. You can store the cache blocks in the SQLite database as blobs. One file, no sparse files. I don't think the "sparse file with separate metadata" approach is necessary here, and sparse files have hidden performance costs that grow with the number of populated extents. A sparse file is not all that different than a directory full of files. It might look like you're avoiding a filesystem lookup, but you're not; you've just moved it into the sparse extent lookup which you'll pay for every seek/read/write, not just once on open. You can simply use a regular file and let SQLite manage it entirely at the application level; this is no worse in performance and better for ops in a bunch of ways. Sparse files have a habit of becoming dense when they leave the filesystem they were created on.
electroly··on Stop using low DNS TTLs
Being specific: AWS load balancers use a 60 second DNS TTL. I think the burden of proof is on TFA to explain why AWS is following an "urban legend" (to use TFA's words). I'm not convinced by what is written here. This seems like a reasonable use case by AWS.
electroly··on RIP Low-Code 2014-2025
Not one of the downvoters, but I'd guess it's because this is only true with HATEOAS which is the part that 99% of teams ignore when implementing "REST" APIs. The downvoters may not have even known that's what you were talking about. When people say REST they almost never mean HATEOAS even though they were explicitly intended to go together. Today "REST" just means "we'll occasionally use a verb other than GET and POST, and sometimes we'll put an argument in the path instead of the query string" and sometimes not even that much. If you're really doing RPC and calling it REST, then you need something to document all the endpoints because the endpoints are no longer self-documenting.
electroly··on RIP Low-Code 2014-2025
A lot of negative responses so I'll provide my own personal corroborating anecdote. I am intending to replace my low-code solutions with AI-written code this year. I have two small internal CRUD apps using Budibase. It was a nice dream and I still really like Budibase. I just find it even easier yet to use AI to do it, with the resulting app built on standard components instead of an unusual one (Budibase itself). I'm a programmer so I can debug and fix that code.
electroly··on Vibe coding kills open source
Reviewing is the easier task: it only has to point me in the right direction. It's also easy to ignore incorrect review suggestions.
electroly··on Vibe coding kills open source
LLMs are great at reviewing. This is not stupid at all if it's what you want; you can still derive benefit from LLMs this way. I like to have them review at the design level where I write a spec document, and the LLM reviews and advises. I don't like having the LLM actually write the document, even though they are capable of it. I do like them writing the code, but I totally get it; it's no different than me and the spec documents.
electroly··on The Holy Grail of Linux Binary Compatibility: Musl and Dlopen
Steel-manning the idea, perhaps they would ship object files (.o/.a) and the apt-get equivalent would link the system? I believe this arrangement was common in the days before dynamic linking. You don't have to redownload everything, but you do have to relink everything.
electroly··on Show HN: A small programming language where everything is pass-by-value
My hobby language[1] also has no reference semantics, very similar to Herd. I think this is a really interesting point in the design space. A lot of complexity goes away when it's only values, and there are real languages like classic APL that work this way. But there are some serious downsides.

In practice I have found that it's very painful to thread state through your program. I ended up offering global variables, which provide something similar to but worse than generalized reference semantics. My language aims for simplicity so I think this may still be a good tradeoff, but it's tricky to imagine this working well in a larger user codebase.

I like that having only value semantics allows us, internally, to use reference counted immutable objects to cut down on copying; we both pass-by-reference internally and present it as pass-by-value to the programmer. No cycle detection needed because it's not possible to construct cycles. I use an immutable data structures library[2] so that modifications are reasonably efficient. I recommend trying that in Herd; it's almost always better than copy-on-write. Think about the Big-O of modifying a single element in an array, or building up a list by repeatedly appending to it. With pure COW it's hard to have a large array at all--it takes too long to do anything with it!

For the programmer, missing reference semantics can be a negative. Sometimes people want circular linked lists, or to implement custom data structures. It's tough to build new data structures in a language without reference semantics. For the most part, the programmer has to simulate them with arrays. This works for APL because it's an array language, but my BASIC has less of an excuse.

I was able to avoid nearly all reference counting overhead by being single threaded only. My reference counts aren't atomic so I don't pay anything but the inc/dec. For a simple language like TMBASIC this was sensible, but in a language with multithreading that has to pay for atomic refcounts, it's a tough performance pill to swallow. You may want to consider a tracing GC for Herd.

[1] https://tmbasic.com

[2] https://github.com/arximboldi/immer

electroly··on Alex Honnold completes Taipei 101 skyscraper climb without ropes or safety net
How do I square "he has debunked that" with the article about his brain fMRI and the results about his amygdala, linked above in this subthread? It's full of direct quotes from both Honnold and the doctors. Where did he debunk it... and how? He's got a more accurate analysis than the fMRI? Do you have a link?
electroly··on Internet Archive's Storage
There's a mention on Wikipedia [1] that the Internet Archive maintains international mirror sites in Egypt and the Netherlands, in addition to several domestic sites within North America.

[1] https://en.wikipedia.org/wiki/Internet_Archive#Operations

electroly··on Show HN: CleanCloud – Cloud cleanup that can't delete anything
My feedback: it seems like this tool isn't really like aws-nuke, but the copy keeps comparing it to aws-nuke, extending further into this HN post. aws-nuke doesn't need delete permissions (you just can't do the "delete" step, obviously), aws-nuke makes you decide what to delete, aws-nuke doesn't need confidence scoring since it shows you everything in the account, and aws-nuke is open source. From your list of key differences, the only one that aws-nuke doesn't already do is the one that doesn't make sense for aws-nuke. This is, IMO, a problem with your list and not with the app: there are differentiating things CleanCloud does that you can focus on instead.

IMO, don't mention aws-nuke at all. This isn't the same kind of product as aws-nuke, which is explicitly the "One-click cleanup workflows" category in your "Not designed for" box. Your tool is for accounts that I'm not trying to nuke. So why invite the comparison? These tools are not intended for the same use case.

Spitballing here, I'd think you would want to lean into the cost savings aspect of deleting orphaned resources. aws-nuke is about cleaning out disposable AWS accounts. CleanCloud is about cloud cost optimization on real production/staging accounts.

A final note: it seems like the name CleanCloud is already used by a laundry service provider. You still have time to pick a different name for which you can take the top Google spot.

electroly··on Linux boxes via SSH: suspended when disconected
EC2 instances can hibernate, too. You stop paying for the instance while it's hibernated; you pay the EBS storage cost only.

https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/Hibernat...

electroly··on Xfce is great
XFCE is X11-only, isn't it? Wayland support is still in development/experimental. I personally use XFCE with X11 to this day.
electroly··on ASCII-Driven Development
Can you explain tmux's contribution here? I'm confused why this process wouldn't work just the same if CC directly executed the program rather than involving tmux. Are you just using tmux to trick the program under test into running its TUI instead of operating in a dumb-stdout mode?
electroly··on Ask HN: Is it time for HN to implement a form of captcha?
My website's contact form has a reCAPTCHA and it still gets spam sent through it (though vastly less). They pass the reCAPTCHA somehow. My contact form literally only emails me and they still do it.
electroly··on Global Memory Shortage Crisis: Market Analysis
Note: the scale does go further than "r" on the high-memory end with some specialty "x" families.

    c*: 2GB per vCPU
    m*: 4GB per vCPU
    r*: 8GB per vCPU
    x2idn/x8g: 16GB per vCPU (!)
    x2iedn/x2iezn/x8aedz: 32GB per vCPU (!)
electroly··on Rob Pike goes nuclear over GenAI
I'm not sure any humans were behind the email at all (i.e. "do that yourself"). This seems to be some bizarre experiment where someone has strapped an LLM to an email client and let it go nuts. Even being optimistic, it's tough to see what good this was supposed to do for the world.
electroly··on It's Always TCP_NODELAY
This hasn't mattered in 20 years for me personally, but in 2003 I killed connectivity to a bunch of Siemens 505-CP2572 PLC ethernet cards by switching a hub from 10Mbps to 100Mbps mode. The button was right there, and even back then I assumed there wouldn't be anything requiring 10Mbps any more. The computers were fine but the PLCs were not. These things are still in use in production manufacturing facilities out there.
electroly··on Log level 'error' should mean that something needs to be fixed
I think OP is making two separate but related points, a general point and a specific point. Both involve guessing something that the error-handling code, on the spot, might not know.

1. When I personally see database timeouts at work it's rarely the database's fault, 99 times out of 100 it's the caller's fault for their crappy query; they should have looked at the query plan before deploying it. How is the error-handling code supposed to know? I log timeouts (that still fail after retry) as errors so someone looks at it and we get a stack trace leading me to the bad query. The database itself tracks timeout metrics but the log is much more immediately useful: it takes me straight to the scene of the crime. I think this is OP's primary point: in some cases, investigation is required to determine whether it's your service's fault or not, and the error-handling code doesn't have the information to know that.

2. As with exceptions vs. return values in code, the low-level code often doesn't know how the higher-level caller will classify a particular error. A low-level error may or may not be a high-level error; the low-level code can't know that, but the low-level code is the one doing the logging. The low-level logging might even be a third party library. This is particularly tricky when code reuse enters the picture: the same error might be "page the on-call immediately" level for one consumer, but "ignore, this is expected" for another consumer.

I think the more general point (that you should avoid logging errors for things that aren't your service's fault) stands. It's just tricky in some cases.

electroly··on 8M users' AI conversations sold for profit by "privacy" extensions
Google allows minified extensions and doesn't require you to provide the original unminified source. I've never provided Google the real source code to my extension and they rubber-stamp every release. The Chrome Web Store is the wild west--you're on your own.

Mozilla allows minification but you're required to provide the original buildable source. Mozilla actually looks at the code and they reject updates all the time.

electroly··on GNU Unifont
I'm curious, have you had this support checked by a native speaker of Chinese/Japanese/Korean? Due to Han unification (Unicode's biggest mistake) many code points require different glyphs depending on the target language, but Unifont only has a single glyph per code point--it's impossible for the one glyph to be correct in all CJK languages. It seems more likely that this "kinda works" for CJK but renders incorrect glyphs some of the time.
electroly··on It seems that OpenAI is scraping [certificate transparency] logs
The CT log tells you about new websites as soon as they come online. Good if you're intending to scrape the web.
electroly··on Wall Street sees AI bubble coming and is betting on what pops it
Just knowing that the bubble will pop at some point in the future isn't enough to trade on. You'll get trounced if this is the only piece of information you have. To a first approximation, everybody knows the bubble will pop. The question is: when and how?
electroly··on Rust Coreutils 0.5.0 Release: 87.75% compatibility with GNU Coreutils
The GNU authors almost certainly did have access to the AT&T UNIX source code, and they had to be reminded not to refer to UNIX source code when writing GNU replacements. GNU made intentional efforts to design their programs along completely different lines to avoid similarity to the originals. This is described at https://www.gnu.org/prep/standards/standards.html#Reading-No... under "Referring to Proprietary Programs".
electroly··on Amazon EC2 M9g Instances
The latter is a 4-core machine with 8 HyperThreads. This doesn't actually matter to your price-performance metric but is worth mentioning because it's the reason why the Intel part performs so comparatively poorly. They're fast chips, they're just wildly uneconomical. If you wanted to compare equal core counts (c8i.4xlarge vs. c8g.2xlarge), then the Intel instance type wins on performance but the Graviton is 58% cheaper.
electroly··on 10 Years of Let's Encrypt
Both Let's Encrypt and 3-year certificates were introduced in 2015. We had 5+ year certificates before that. At the time you'd buy the longest certificate possible and forget about it--that's what I did. In 2013 I bought a 5-year certificate (self-service, no tickets) and didn't think about it again until 2018.
electroly··on Shrinking While Linking
This response answers too much--why did we invent and adopt tar for all other use cases, then? Why do we not still use ar for its original general archiving purpose in any other situation but this one? Obviously there is some reason! I'm curious what that reason is. You can make regular archives with ar today but we don't--"it still works" wasn't good enough reason for any other archive use case. What are the general archiving requirements that don't exist for object archives that meant ar is still fine for this one use case?
electroly··on At IT School with Apple Lisa
> in return for giving Xerox the _opportunity_ to invest in Apple, wow

You're being sarcastic but this would have been the most lucrative thing Xerox ever did in its entire corporate life, by far, if it had held onto the stock. This was a really good deal in hindsight. Indeed, it would have been better to liquidate Xerox and put all the proceeds into Apple stock; I don't think anybody argues that Xerox could have made as much hay as Apple did with the technology, even in the best of scenarios. It couldn't have known that at the time, of course.

electroly··on Japanese game devs face font dilemma as license increases from $380 to $20k
OP means it's the only one that renders at all. It could be wrong but the other two are just missing-glyph boxes which are definitely wrong. Those code points appear unusable as a practical matter.
← PreviousPage 4 of 34Next →