HNHacker News
TopNewBestAskShowJobs

emn13

3,627 karma · joined December 12, 2011

submissionscomments
emn13··on English: A vs. An
Well, I did once scrape wikipedia and classify based off simply off string suffix (storing those where longer suffixes disagreed with the shorter rule), and that's quite simple and effective - https://eamonnerbonne.github.io/a-vs-an/AvsAnDemo/ - and since it's based off actual statistics, it correctly distinguishes textually subtle stuff like "a NASA scientist" vs. "an NSA analyst".

There are structural limitations (i.e. this implementation never looks at preceding context and sometimes that matters, nor does it understand other clues like punctuation or multi-word suffixes). Nevertheless, the accuracy is high enough that I'm not sure it'd be easy for a small model to beat it, especially not without considerably more work to make sure your inputs cover more context (which in principle the plain statistics approach could likely deal with too).

If the whole point of multi-layer networks is to deal with weirdly shaped, non-obvious manifolds in a very high-dimensional space, then this problem just doesn't look that difficult and perhaps does not need that mathematical finesse: just store the prototypical examples and you're pretty much there without anything fancier.

emn13··on English: A vs. An
Sure, but note that a tiny LLM is many order of magnitudes larger than algorithms like this. It might well work better, but you're probably paying for very, very marginal improvements with a solution that's thousands of times larger and slower, unless you're you're being really creative and using a fairly non-LLM network of perceptrons or other ML technique. But it sounds like a fun challenge; if you or anybody feels like exploring that, I bet it'd get quite a few curious clicks here!
emn13··on Let's make quality the norm again
The thing is, it's a false dichotomy. Raising the quality bar makes software a lot _easier_ to deliver, in practice. Even for a not techie, the reason is probably somewhat intuitive - if you're trying to add features to something that's a hugely complicated mess, you're going to struggle and cause many bugs; but if your process allows for the creation of high quality code, you have the means to do so with confidence, and that usually means quickly and reliably.

The real issue is one of timing and tech debt - higher quality software is cheaper and better, but... that's only over time. If you want that new software from scratch _now_, and don't have the time to set up a high quality development process, you can hack something together quickly. If the software is small enough, the person or LLM writing it smart enough, and while everything from requirements to code is still fresh in memory - that works pretty well! But that process scales poorly, and yet it's what we've as an industry done far too often, and at far too pervasive a scale.

So even though we know and have examples of how to design software better and more cheaply, actually getting to that situation isn't a simple question of trying harder; it's working entirely differently. And unfortunately, we build more and more of our software on the building blocks of the past, and many of those are themselves of low quality too; the more we do that, the more entrenched low quality becomes. As an end user that kind of isn't perceivable; it just means all software is a bit worse and more expensive, but as humanity, it's absurd how much effort we waste working around fairly trivial mistakes made in the past, typically by organizations that just had no incentives to care about future impacts especially on others.

While I get the instinct to prefer something over nothing as an end user it kind of is exactly that bad habit that got us into the mess we're in today. We don't build and throw away; we copy and tweak - so all the messes are inherited, kind of like a form of pollution.

emn13··on Let's make quality the norm again
I think all three of these points, but especially the first two are rather speculative; I don't personally believe they hold, though I can't be sure.

Most quality defects I see are ultimately down to culture, not training and certainly not certification. Getting to excellent quality simply does not require great training nor rigorous certification, and I'm not even sure it really benefits from either all that much. But what it does require is a culture of chasing down all the avoidable risks, to fix not just the immediate bug, but also the process that allowed it to come to be, to pre-emptively look for whole classes of bugs or misfeatures seriously, and to always look deeper than for underlying causes or interactions (i.e. take the 5-why's game seriously).

It's perhaps almost a trope by this point, but it's still a great reminder and example - sqlite. About a month ago Richard Hipp, one of the primary authors, gave a great presentation "Reliability Lessons From SQLite" (https://youtu.be/V_qzqY1bb7I) - and notably, I'm not seeing even _echos_ of training and certification in their clearly world-class process. I don't think you can even seriously claim it's all that expensive, if you consider how tiny that team is, and yet how much they've achieved.

If larger organizations fail to achieve high quality - well, _why?_ It's not due to lack of training, certifications or smart people. Perhaps its intrinsic about larger for-profit corporations, but I suspect we could do better with better incentives, and a more serious attitude about dealing with mis-aligned incentives throughout our culture. Is it really a flaw for a corporation to prioritize short term cost reduction even over long term self-harm if that's what they're judged by? Is it a flaw for a profit seeking corporation to largely ignore costs their low quality imposes on others? Is it a flaw for a for-profit organization to take huge risks when the upside is unbounded but the downside is mere bankruptcy without any further liability?

I don't know how to solve these problems, and don't want to claim it's even remotely reasonable to throw out the capitalist baby with the corporate bathwater, but surely we can do better than the status quo even without a revolution. Misaligned incentives and tragedies of the commons aren't new issues, nor ones without potential mitigations. We're just choosing to ignore those inefficiencies, and have for decades. And at this point, even small changes in incentives might have truly wrenching consequences; the ingrained and by now deeply embedded carelessness in our huge software and other engineering stack is corroded at every level.

TLDR: I don't have the answer, but I'm pretty positive it's not an issue of cost, training or certification.

emn13··on YouTube to automatically label AI-generated videos
Note that the uploader apparently still retains control over labeling in most cases; uploaders that intentionally misrepresent AI-generated content might not be discouraged by this. Whether youtube will (and can) ban accounts that do that might determine in practice if this matters or not.
emn13··on The Impossible Prompt
Create an image that displays two seven-pointed stars, two eight-pointed stars, and two nine-pointed stars. All stars are connected to each other, except for the ones with the same number of strands. The lines connecting the stars must NOT intersect.
emn13··on Claude Skills are awesome, maybe a bigger deal than MCP
I think it's not at all a marshmellow test; quite the opposite - docs used to be written way, way in advance of their consumption. The problem that implies is twofold. Firstly, and less significantly, it's just not a great return on investment to spend tons of effort now to maybe help slightly in the far future.

But the real problem with docs is that for MOST usecases, the audience and context of the readers matter HUGELY. Most docs are bad because we can't predict those. People waste ridiculous amounts of time writing docs that nobody reads or nobody needs based on hypotheses about the future that turn out to be false.

And _that_ is completely different when you're writing context-window documents. These aren't really documents describing any codebase or context within which the codebase exists in some timeless fashion, they're better understood as part of a _current_ plan for action on a acute, real concern. They're battle-tested the way docs only rarely are. And as a bonus, sure, they're retainable and might help for the next problem too, but that's not why they work; they work because they're useful in an almost testable way right away.

The exceptions to this pattern kind of prove the rule - people for years have done better at documenting isolatable dependencies, i.e. libraries - precisely because those happen to sit at boundaries where it's both easier to make decent predictions about future usage, and often also because those docs might have far larger readership, so it's more worth it to take the risk of having an incorrect hypothesis about the future wasting effort - the cost/benefit is skewed towards the benefit by sheer numbers and the kind of code it is.

Having said that, the dust hasn't settled on the best way to distill context like this. It's be a mistake to overanalyze the current situation and conclude that documentation is certain to be the long-term answer - it's definitely helpful now, but it's certainly conceivable that more automated and structured representations might emerge, or in forms better suited for machine consumption that look a little more alien to us than conventional docs.

emn13··on Apple M5 chip
Yep. Still, I think it's a pretty decent benchmark in the sense that it's fairly short, quite repeatable, does have a quite a few subtest, and it's horribly different from the nebulous concept that is "typical workloads". It's suspiciously memory-latency bound, perhaps more than most workloads, but that's a quibble. If they'd have simply labelled it "lightly threaded" instead of "multithreaded", it would have been fine.

As it is, it's just clearly misleading to people that haven't somehow figured out that it's not really a great test of multithreaded throughput.

emn13··on Apple M5 chip
It's not trash - it's quite nice for its niche. It's just not very scalable with cores, so it's best interpreted as a benchmark of lightly threaded workloads - like lots of typical consumer workloads are (gaming, web browsing, light office work). Then again, it's not hard to find workloads that scale much better, and geekbench 6 doesn't really have a benchmark for those.

For the first 8 threads or so, it's fine. Once you hit 20 or so it's questionable, or at least that's my impression.

emn13··on Upcoming Rust language features for kernel development
I mean, reliably tracking ownership and therefore knowing that e.g. an aliased write must complete before a read is surely helpful?

It won't prevent all races, but it might help avoid mistakes in a few of em. And concurrency is such a pain; any such machine-checked guarantees are probably nice to have to those dealing with em - caveat being that I'm not such a person.

emn13··on Detect Electron apps on Mac that hasn't been updated to fix the system wide lag
There's also the alternative of announcing this breakage publicly to electron beforehand; and the alternative of having a hack and publicly announcing it will be removed in a year. There's even the alternative of just announcing the caveat at all, so your users aren't unwitting guinea pigs. If they don't want to support a million workarounds forever, they don't have to it's not all or nothing.
emn13··on Detect Electron apps on Mac that hasn't been updated to fix the system wide lag
Put it this way: if I were in charge of a major OS, and I having one of the major app frameworks used on my OS tested on for my annual upgrade, I'd feel pretty embarrassed, even if there's a figleaf excuse why it's not my fault.

This doesn't exactly instill confidence in Apple's competence.

emn13··on Solar panels + cold = A potential problem
Hyper amusing, thanks for sharing! Doesn't really improve the analogy, but fun quirk of history :-D.
emn13··on Solar panels + cold = A potential problem
So on the one hand we have a product which isn't even remotely designed for the use case (hamsters), and during normal use shows obvious behaviour (cooking) that should imply risk to said hamsters. On the other side, we have a product designed to be installed in an electrical system, and shows no signs during normal use that it's installed unsafely, and where the advertised specs are not actually safe for normal usage.

Whether or not the company in this case shares some or most of the blame with novice users - the analogy is not a great one.

emn13··on Next.js is infuriating
The author's examples of rough edges are however no better when hosted on vercel. The architecture seems... overly clever, leading to all kinds of issues.

I'm sure commercial incentives would lead issues that affect paying (hosted) customers to have better resolutions than those self-hosting, but that's not enough to explain this level of pain, especially not in issues that would affect paying customers just as much.

emn13··on Beware of Fast-Math
If you care about absolute accuracy, I'm skeptical you want floats at all. I'm sure it depends on the use case.

Whether it's the standards fault or the languages fault for following the standard in terms of preventing auto-vectorization is splitting hairs; the whole point of the standard is to have predictable and usually fairly low-error ways of performing these operations, which only works when the order of operations is defined. That very aim is the problem; to the extent the stardard is harmless when ordering guarrantees don't exist you're essentially applying some of those tricky -ffast-math suboptimizations.

But to be clear in any case: there are obviously cases whereby order-of-operations is relevant enough and accuracy altering reorderings are not valid. It's just that those are rare enough that for many of these features I'd much prefer that to be the opt-in behavior, not opt-out. There's absolutely nothing wrong with having a classic IEEE 754 mode and I expect it's an essentialy feature in some niche corner cases.

However, given the obviously huge application of massively parallel processors and algorithms that accept rounding errors (or sometimes conversely overly precise results!), clearly most software is willing to generally accept rounding errors to be able to run efficiently on modern chips. It just so happens that none of the computer languages that rely on mapping floats to IEEE 754 floats in a straitforward fashion are any good at that, which is seems like its a bad trade off.

There could be multiple types of floats instead; or code-local flags that delineate special sections that need precise ordering; or perhaps even expressions that clarify how much error the user is willing to accept and then just let the compiler do some but not all transformations; and perhaps even other solutions.

emn13··on Beware of Fast-Math
I get the feeling that the real problem here are the IEEE specs themselves. They include a huge bunch of restrictions that each individually aren't relevant to something like 99.9% of floating point code, and probably even in aggregate not a single one is relevant to a large majority of code segments out in the wild. That doesn't mean they're not important - but some of these features should have been locally opt-in, not opt out. And at the very least, standards need to evolve to support hardware realities of today.

Not being able to auto-vectorize seems like a pretty critical bug given hardware trends that have been going on for decades now; on the other hand sacrificing platform-independent determinism isn't a trivial cost to pay either.

I'm not familiar with the details of OpenCL and CUDA on this front - do they have some way to guarrantee a specific order-of-operations such that code always has a predictable result on all platforms and nevertheless parallelizes well on a GPU?

emn13··on Past, present, and future of Sorbet type syntax
Yeah, before required properties/fields, C#'s nullability story was quite weak, it's a pretty critical part of making the annotations cover enough of a codebase to really matter. (technically constructors could have done what required does, but that implies _tons_ of duplication and boilerplate if you have a non-trivial amount of such classes, records, structs and properties/fields within them; not really viable).

Typescript's partial can however do more than that - required means you can practically express a type that cannot be instantiated partially (without absurd amounts of boilerplate anyhow), but if you do, you can't _also_ express that same type but partially initialized. There are lots of really boring everyday cases where partial initialization is very practical. Any code that collects various bits of required input but has the ability to set aside and express the intermediate state of that collection of data while it's being collected or in the event that you fail to complete wants something like partial.

E.g. if you're using the most common C# web platform, asp.net core, to map inputs into a typed object, you now are forced to either expression semantically required but not type-system required via some other path. Or, if you use C# required, you must choose between unsafe code that nevertheless allows access to objects that never had those properties initialized, or safe code but then you can't access any of the rest of the input either, which is annoying for error handling.

typescript's type system could on the other hand express the notion that all or even just some of those properties are missing; it's even pretty easy to express the notion of a mapped type wherein all of the _values_ are replaces by strings - or, say, by a result type. And flow-sensitive type analysis means that sometimes you don't even need any kind of extra type checks to "convert" from such a partial type into the fully initialized flavor; that's implicitly deduced simply because once all properties are statically known to be non-null, well, at that point in the code the object _is_ of the fully initialized type.

So yeah, C#'s nullability story is pretty decent really, but that doesn't mean it's perfect either. I think it's important to mention stuff like Partial because sometimes features like this are looked at without considering the context. Most of these features sound neat in isolation, but are also quite useless in isolation. The real value is in how it allows you to express and change programs whilst simultaneously avoiding programmer error. Having a bit of unsafe code here and there isn't the end of the world, nor is a bit of boilerplate. But if your language requires tons of it all over the place, well, then you're more likely to make stupid mistakes and less likely to have the compiler catch them. So how we deal with the intentional inflexibility of non-nullable reference types matters, at least, IMHO.

Also, this isn't intended to imply that typescript is "better". That has even more common holes that are also unfixable given where it came from and the essential nature of so much interop with type-unsafe JS, and a bunch of other challenges. But in order to mitigate those challenges TS implemented various features, and then we're able to talk about what those feature bring to the table and conversely how their absence affects other languages. Nor is "MOAR FEATURE" a free lunch; I'm sure anybody that's played with almost any language with heavy generics has experienced how complicated it can get. IIRC didn't somebody implement DOOM in the TS type system? I mean, when your error messages are literally demonic, understanding the code may take a while ;-).

emn13··on Past, present, and future of Sorbet type syntax
I love building libraries, so having the chance to talk about the gotchas with things like this is a fun chance to reflect on what is and is not possible with the tools we have. I guess my favorite "feature" in C# is how willing they are to improve; and that many of the improvements really matter, especially when accumulated over the years. A C# 13 codebase can be so much nicer than a c# 3 codebase... and faster and more portable too. But nothing's perfect!
emn13··on Past, present, and future of Sorbet type syntax
"Recovered" sounds so binary.

I think it's pretty usuable now, but there is scarring. The solution would have been much nicer had it been around from day one; especially surrounding generics and constraints.

It's not _entirely_ sound, nor can it warn about most mistakes when those are in the "here-be-dragons" annotations in generic code.

The flow sensitive bit is quite nice, but not as powerful as in e.g. typescript, and sometimes the differences hurt.

It's got weird gotcha interactions with value-types, for instance but likely not limited to interaction with generics that aren't constrained to struct but _do_ allow nullable usage for ref types.

Support in reflection is present, but it's not a "real" type, and so everything works differently, and hence you'll see that code leveraging reflection that needs to deal with this kind of stuff tends to have special considerations for ref type vs. value-type nullabilty, and it often leaks out into API consumers too - not sure if that's just a practical limitation or a fundamental one, but it's very common anyhow.

There wasn't last I looked code that allowed runtime checking for incorrect nulls in non-nullable marked fields, which is particularly annoying if there's even an iota of not-yet annoted or incorrectly annotated code, including e.g. stuff like deserialization.

Related features like TS Partial<> are missing, and that means that expressing concepts like POCOs that are in the process of being initialized but aren't yet is a real pain; most code that does that in the wild is not typesafe.

Still, if you engage constructively and are willing to massage your patterns and habbits you can surely get like 99% type-checkable code, and that's still a really good help.

emn13··on Past, present, and future of Sorbet type syntax
While I'm most familiar with C#, and haven't used Ruby professionally for almost a decade now, I think we'd be better off looking at typescript, for at least 3 reasons, probably more.

1. Flowsensitivity: It's a sure thing that in a dynamic language people use coding conventions that fit naturally to the runtime-checked nature of those types. That makes flow-sensitive typing really important.

2. Duck typing: dynamic languages and certainly ruby codebases I knew often use ducktyping. That works really well in something like typescript, including really simple features such as type-intersections and unions, but those features aren't present in C#.

3. Proof by survival: typescript is empirically a huge success. They're doing something right when it comes to retrospectively bolting on static types in a dynamic language. Almost certainly there are more things than I can think of off the top of my head.

Even though I prefer C# to typescript or ruby _personally_ for most tasks, I don't think it's perfect, nor is it likely a good crib-sheet for historically dynamic languages looking to add a bit of static typing - at least, IMHO.

Bit of a tangent, but there was a talk by anders hejlsberg as to why they're porting the TS compiler to Go (and implicitly not C#) - https://www.youtube.com/watch?v=10qowKUW82U - I think it's worth recognizing the kind of stuff that goes into these choices that's inevitably not obvious at first glance. It's not about the "best" lanugage in a vacuum, it's a about the best tool for _your_ job and _your_ team.

emn13··on Cursor IDE support hallucinates lockout policy, causes user cancellations
Of course they had a choice: they could have stuck with google maps for longer, and they probably also could have invested more in data and UI beforehand. They could have launched a submarine non-apple-branded product to test the waters. They could likely have done other things we haven't thought of here, in this thread.

Quite plausibly they just didn't realize how rocky the start would be, or perhaps they valued that immediate strategic autonomy more in the short-term that we think, and willingly chose to take the hit to their reputation rather than wait.

Regardless, they had choices.

emn13··on Cursor IDE support hallucinates lockout policy, causes user cancellations
While some of what you say is an interesting thought experiment, I think the second half of this argument has, as you'd put it, a low symbolic coherence and low plausibility.

Recognizing the relevance of coherence and plausibility does not need to imply that other aspects are any less relevant. Redefining truth merely because coherence is important and sometimes misinterpreted is not at all reasonable.

Logically, a falsehood can validly be derived from assumptions when those assumptions are false. That simple reasoning step alone is sufficient to explain how a coherent-looking reasoning chain can result in incorrect conclusions. Also, there are other ways a coherent-looking reasoning chain can fail. What you're saying is just not a convincing argument that we need to redefine what truth is.

emn13··on AI agents: Less capability, more reliability, please
Perhaps the solutions(s) needs to be less focusing on output quality, and more on having a solid process for dealing with errors. Think undo, containers, git, CRDTs or whatever rather than zero tolerance for errors. That probably also means some kind of review for the irreversible bits of any process, and perhaps even process changes where possible to make common processes more reversible (which sounds like an extreme challenge in some cases).

I can't imagine we're anywhere even close to the kind of perfection required not to need something like this - if it's even possible. Humans use all kinds of review and audit processes precisely because perfection is rarely attainable, and that might be fundamental.

emn13··on The case of the critical section that let multiple threads enter a block of code
Is the advantage over an enum not kind of small? We're seeing bugs here because people tried to do the right thing but the tooling has absolutely no way of helping anybody to do that. Simply preventing accidental mistakes would prevent these. Adding complexity to make it harder (though never impossible) for consumers to misuse the API in a complex way seems like it's potentially going to far.

Then again, it's been years since I used this kind of C, so maybe my instincts are rusty here (no rust-pun intended!)

emn13··on It's not a crime if we do it with an app
While a humorous response by Milton and an interesting debating point, the argument he makes is pretty weak because it almost inevitably reduces to complete lawlessness, doesn't really define which government "granted" monopolies he's willing to give up, and ultimately relies on a fairly arbitrary definition of what government even is - and one that if you really let it go to the extreme not only obviously just doesn't work well for most people, it also does not avoid monopolies as is witnessed every day around the globe.

After all, the natural inclination of a powerful elite is to protect their interest. It's business 101 to want a moat, and tearing down one set of artifical legal protections that allows for a moat allows on the other hand for the far more extreme quite physically violent moat in the form of a putin-esque kleptocracy.

The argument merely sounds convincing because it's very selectively implying that certain monopolies are created by state power and might be weakened by free market principles without considering what a free market even is (generally a regulated one), nor addressing the fact that other monopolies will arise precisely because because the lack of regulation allows winner-take-all brute force strategies to work.

That doesn't mean Milton's ideas are without merit - but that there is a breaking point; dogmatically hoping for anarchy to avoid harmful centralization of power is problematic because of the dogma; not because it's never a valid approach.

But sure, if you're going to embrace Milton's (intentionally) vague proposition in the way it was likely intended - to provoke thought - then sure; there are state regulations that are in part to blame for some of today's near monopolies - the interaction between intellectual property, incorporation, and state-enforced contract law. As a matter of debate, sure, it'd be interesting to weaken all three and in particular their interactions. I just highly doubt that's very practical, nor would it be very easy to predict the outcome, especially once international power-plays start circumventing even the best of intentions.

emn13··on It's not a crime if we do it with an app
I'm curious what you base that on. For instance, we've never really allowed literally cutthoat competition, nor things like fraud and we've generally not allowed misrepresentation. Governments intervene heavily and always have to set those kind of boundary conditions - but there are really lots of them. Economies of scale seem to be very, very common ever since the industrial revolution; and even more so in today's information-economy platform era.

I'm sure there are plenty of cases where significant competition is a natural end state, but how common those are in comparison? I'm curious.

emn13··on It's not a crime if we do it with an app
One race-to-the-bottom phenomenon that (to me at least) appears to aggravate the impact of "corporate greed" is the social loop that goes as follows:

1. company decides to push the boundaries of the socially acceptable when it comes to cutting corners (e.g. screwing their customers, or employees, or environment, or debtors)

2. People don't like it, but rationalize this as being a natural consequence of incorporation and the profit motive. Hence while they grumble, there negative impacts to public perception don't actually cost the company as much as you might think

2b. Even if there's a boycott, there will be vocal minority that thinks it's all a bunch of whiny <target audience we're better than>. They'll actively harass or undermine said boycott or backlash, even if in a purely egotistical sense their interests are actually aligned with the boycotters

3. Social norm is reset; we all collectively expect even less from companies. That doesn't however mean the new norm is better or maintained, because as soon as there's some new major conflict between short-term profit and maintaining a decent reputation in public, we go back to step 1 from the new, lower baseline.

Stuff like increasing partisanship, and decreasing incentives for journalist (whether profession, citizen or influencer) to maintain their professional standing (as opposed to targeting clickbate) probably smears those gears nicely.

Many companies have historically clearly paid well over the odds to maintain their reputation, and done well doing so. It's just not true that nihilistic short term greed has always paid; obviously it didn't and still doesn't really. It profitable to do the little, but simultaneously also to do as many cheap things that materially affect public standing as possible.

By promoting the profit motive past a merely utilitarian means to an efficiency-optimizing end into a matter of national identity and point of distinction vs. in particular the USSR, we've shifted our culture beyond what's really rational. We (as a society) don't merely respect and understand the profit motive; we see it as a sign of merit - and significant enough merit that "winning" on that scale excuses a lot of other bad behavior.

emn13··on It's not a crime if we do it with an app
Outright monopolistic pricing is also "the market rate". Frankly, virtually any price somebody is willing to pay is almost by definition "the market rate". It's a meaningless defense for an artificially high price.
emn13··on The mistake of yearning for the 'friendly' online world of 20 years ago
Certainly helped!
Page 1 of 34Next →