HNHacker News
TopNewBestAskShowJobs

elgertam

302 karma · joined April 24, 2018

https://github.com/elgertam
submissionscomments
elgertam··on Jev in 25 Lines of Python
> With BERT, you need a large, labeled dataset, and you have to train/fine-tune the model.

BERT requires a huge corpus, but it isn't labeled. BERT is trained through self-supervised learning using mask tokens and next sentence prediction. Fine-tuning is useful for specific tasks, but isn't absolutely essential for the model to function.

elgertam··on Anecdotally, programmers dislike "reduce"
The only part of "hard to read" that has ever made sense is that the callback takes multiple args and sometimes I can't remember the order of the initial value versus the accumulator.

Incidentally, reduce is also powerful enough to implement both map and filter in terms of itself, though that's more of a teaching exercise than a good recommendation.

I mostly interpret it as of the same spirit with those who oppose proper tail calls because it "ruins" their debugging stack traces.

elgertam··on Hitachi launches CO2 heat pump water heaters with solar-friendly tariff controls
> Some Japanese utilities have introduced tariffs designed to encourage daytime water heating as growing volumes of solar generation enter the grid. [Emphasis added.]

I was quite confused at the use of "tariff" here, as it meant to me a "tax on a good crossing a political boundary." Turns out, 'tariff' has an older meaning: a published schedule of fees or taxes issued by some authority.

elgertam··on What algorithm did Windows XP use to choose your initial user picture?
On initial install, sure. But user accounts can also be created at arbitrary times. The user may have changed the set of photos in the intervening time and might even be editing the directory during profile creation.
elgertam··on How An AI math breakthrough ignited a controversy
> “I certainly don't expect the industry to continue to spend millions of dollars to solve problems in mathematics, because there is no profit in it,” Columbia University mathematician Michael Harris wrote in an email to Science. But he worries the highly publicized achievement will be “extremely damaging to mathematics; it convinces decision makers that human mathematicians are obsolete, and it convinces young people that their passion for mathematics has no future.”

LLMs seem particularly suited toward these existence-proof problems. Working mathematicians seem absolutely essential for universally quantified results, still. I strongly doubt, for example, that if Fermat's Last Theorem hadn't been proven three decades ago, that an LLM would be able to do work equivalent to inventing the mathematics as Andrew Wiles did to solve the problem. I have similar doubts about P vs NP, the twin prime conjecture, even the Riemann Hypothesis (unless the latter has at least one counterexample).

And I want to be clear: I'm not downplaying the achievements of these models. This is remarkable! I simply think that the pattern of success is in existence proofs or finding counterexamples, which makes sense based on how LLMs function and are trained.

elgertam··on How An AI math breakthrough ignited a controversy
The site should have a disclaimer at the bottom: "A Sam Altman Production."
elgertam··on 'You Can See Everything' Review: Nathan Fielder's Doc About Elizabeth Holmes
Holmes seems to have the Harvard Business (MBA) mindset: science, R&D, engineering are just results to squeeze out of recalcitrant practitioners. Any failure in those domains is evidence only of the scientists' or engineers' refusal to align themselves with the Vision.

Despite not even finishing undergrad, she adopted this persona so naturally, including the black turtleneck and a Zuck-esque vocal inflection, that due diligence apparently defenestrated itself. Phyllis Gardner is the unsung hero who tried to warn the world, but the MBA mass-delusion was too strong.

elgertam··on Mom gets 6-month suspended sentence for letting 5-year-old walk to the pond
I live in this neighborhood. I was just kvetching to my wife this morning about how our HOA wastes so much money on excess. That had me in the wrong (or right?) mood to read this.

I will say, we have a LOT of elderly retired people who live here. It's not uncommon to hear sirens because someone is having a medical episode, and the security office is helpful with responding to those and helping with EMS. They are also helpful for doing well checks on the elderly when they haven't been responsive to others.

That being said, a lot of what they do is run radar and evidently do absolutely atrocious behavior like this. One of the flip sides of the retired population is that many of them devote their lives to reporting their neighbors to The Authorities – generally the HOA or security office. In this case, I guess the informant went straight to the cops.

elgertam··on Mom gets 6-month suspended sentence for letting 5-year-old walk to the pond
I live in the neighborhood in question. While my immediate neighbors are all really awesome people, the neighborhood as a whole has a LOT of retired people whose sole activity in life is to report their neighbors to The Authorities. Usually that's the HOA management or the security office, but sometimes they'll call the actual police.

Many of us find it frustrating, though exposure is almost random. Someone who used to live on our street used to get very upset about trash cans being visible from the street, for instance. I think this person moved in the last two or three years because the passive aggressive HOA letters about this apparently heinous nuisance have stopped. But among groups of affluent people, I find that a small percentage never outgrew the tattle-tale phase and still behave in such a way into their sixties, seventies and eighties.

elgertam··on Mom gets 6-month suspended sentence for letting 5-year-old walk to the pond
This happened in my neighborhood, and I've met the "perpetrator." She's a sweet person, and I'm convinced she was absolutely acting in the best interests of her child as she saw it. Kids as young as elementary school age in our neighborhood frequently go to the various ponds for fishing, watching the wildlife or even sailing remote control boats. Five years old without a buddy might be on the younger end, but it's hardly criminally negligent. And given the several close calls I've had with elementary & middle school-aged kids on e-bikes and scooters this summer, I hardly think this is even close to the most serious youth safety problem here.
elgertam··on I Remain a Skeptic
From my reading on the topic, the tokens are subsidized when considering average cost, but are profitable at marginal cost. Basically, they're super expensive when considering the training cost, but aren't super expensive when doing inference. Since AGI is quite unlikely now, my guess is that we'll see consolidation and will see frontier models released at a slower clip, such that they can pay for training for from the profits from prediction tokens themselves. There may be slight increases in token prices, but there's sufficient competition between Google, Anthropic, OpenAI and X alongside the open weights providers that I don't see huge price increases happening. Even if OpenAI is absorbed into one of the others (which I think is most likely), that would still make token price collusion difficult.
elgertam··on Why does Opus 5 feel worse to work with?
The second one has caused hours of entertainment over the past year or so. My kids find the LLM's failed clues and profuse apologies for getting these wrong hilarious, so it's become a family activity with me performing dramatic readings of the chat transcript with them. The LLM's apologies also seem to get more exaggerated as the context increases and the LLM seems to get more deranged.

I'd prefer the models to get better at SVG. I really hate working with the rasters that diffusion models generate, but the vector outputs are just really bad even when tokenizable like SVG. I've done some experimentation with trying to make these work better with some newer techniques with some success. But I also think the SVG Paths mini-language may be a bit too concise and unforgiving for LLMs to consistently get them right without specialized training.

elgertam··on The American sports plutocracy
> Everyone who stays for at least 7 innings is staggering out, drunk as a skunk. Whether I ride the train home or I get on the roads, I'm positively surrounded by drunk zombie people. Woe betide us if the home team didn't win.

Good Lord, what games have you been going to? I see a couple MLB games live every year and a few minor league games. I haven't seen anything like what you're describing. That doesn't mean there are no drunks or obnoxious fans, but that's the risk of doing something in public. The last game I saw live was San Francisco at Milwaukee in June. Giants won 1-0, and the home team crowd was fine. And I mean the Brewers fans like their beer, but in no way was it a zombie apocalypse. Last game I saw in Oracle Park last year was also fine. Even Boston a few years ago, which definitely has more of a college town vibe than SF or Milwaukee, was hardly zombie land.

I'll also say that AAA ball is typically awesome. The games are much cheaper than the big leagues, you get to see some big league talent rehabbing from time to time and see the occasional prospect who gets called up. They also tend to be very family friendly.

> Baseball has been a game of moral fiber, respectable players, and good clean fun. But it seems that activists and jerks set out to ruin the experience for people like me. I'm not having it. I'd have way more fun at a punk rock women's roller derby bout.

I don't know about that. Baseball has a certain level of honor, sometimes an excessive amount e.g. some of the "code" that has started to fade away the last decade or so. It also invented the concept of a nearly all-powerful commissioner, which was created specifically because players threw a World Series for a payout. It's a wonderful game, a thinking person's game, but it's always been quintessentially American for both good and ill ever since it scaled nationwide during the Civil War.

elgertam··on Google is making private AI practical with homomorphic encryption
I saw a paper about this in early 2020 (pre-COVID shutdowns) at the ScaledML conference. I looked into it and had the same conclusions. At some point, running your own models in the clear is just more practical.
elgertam··on Why does Opus 5 feel worse to work with?
> see sibling comment,

> > see it actually improve just through accreting context

> this actually happens and has been tested.

I specifically said a novel task outside of the explicit training. And I already agreed that the so-called thinking models do some level of logical reasoning. But being able to engage in some level of reasoning because it has learned logical inference rules doesn't mean it's actually thinking, regardless of what the researchers wish to call it.

Also, why does each model always fail at the two tests I give it? The models not only fail to improve, but they start to degrade after many subsequent iterations. Someone who can think would at least not get worse.

LLMs are filters or tuners for extremely subtle patterns, patterns that humans frankly are not great at finding. That's what the attention mechanism does: attend to the other tokens that are most related in a given context, even if that related context is distant in the token stream. Some patterns they fail to detect because they haven't been sufficiently trained or post-trained, and so the LLM just attends to noise (or at least that's what appears to be happening).

A lot of intelligence can be effectively mimicked through this pattern synthesis by transformer architecture alone. That's surprising. But I have yet to see them think.

elgertam··on Why does Opus 5 feel worse to work with?
If I could give it a novel task outside of its explicit training and see it actually improve just through accreting context, I'd be convinced it was thinking.

The opposite happens in practice. I test new models with two tasks: iteratively generating SVGs based on a text description with rendered rasters for feedback; and generating "Before and After" clues like on Jeopardy, where the response has two overlapping phrases such that the last word of the first phrase must be identical to the first word of the last phrase. I have yet to find a model that is consistently good at either. And actually they tend to exhibit context rot with these tasks, where they seem drunk or stoned and the quality degrades.

They're extremely good pattern filters, and that includes some level of logical reasoning. But they aren't reflective or adaptable. Just last night, for instance, I was teaching my son about rounding to the nearest millions. It became clear that he didn't know the place values of large numbers, so we reviewed that till he was consistently correct, and then he was consistently great at rounding to the nearest millions or ten millions or hundred billions or whatever. He's thinking. LLMs are not.

elgertam··on What AI did to stackoverflow in a graph
There were also certain IMO low-value questions that really excited the SO hive-mind. I asked a question about the peculiarities of Python assignment syntax, and earned several dozen points for the question, even though no one really should have written code the way I presented it.

I liked StackOverflow for the first ten years or so of its existence, but I gradually stopped using it then suddenly quit altogether when valid questions were being closed unreasonably. At this point, LLMs with documentation in the context, issue trackers and eve the source code (if available) have surpassed SO. Now my main issue is telling the LLM to crap on my idea rather than wishing it were kinder.

elgertam··on County with 37 Data Centers Asks Schools to 'Conserve Electricity'
As a Virginian, this is good information to have. I see a lot of ludicrous objections to data centers here (the most ludicrous being water consumption, when most of our data centers have closed-loop systems and regardless the humidity here isn't evaporating water).

I've suspected that the energy regulations and the ruling party's close connection with Dominion Energy (the Governor recently attempted to fire the chair of Virginia Tech's board and replace him with the CEO of Dominion) have had an impact on power use more than data centers themselves.

elgertam··on Show HN: Smart model routing directly in Claude, Codex and Cursor
Let's just say this organization is very, very large and doesn't necessarily have the budget for everyone to have all of the tokens.
elgertam··on Show HN: Smart model routing directly in Claude, Codex and Cursor
I ran into a problem at work recently: we are given access to a bunch of models up to a full Claude Opus 4.8, but a monthly budget of 100k tokens. We are also given access to Gemini 3.5 Flash & 3.1 Pro with a daily budget of 50M tokens, but no tool calling. I'd love to hook Claude Code (or Pi) into the Gemini model, but the lack of tool-calling makes it quite difficult. I've been planning out how an intelligent router might be able to use a token-efficient tool-calling model (including a small local open-weights model) to handle the basic tools like reading from the file system or interfacing with MCP servers such that context is gathered, but then send the built up context to the Gemini model where I have a nearly unlimited (for my use cases) token budget.

Could your router handle this?

elgertam··on Mexican government unveils a prototype for a new homegrown, ultra-affordable EV
My first thought when looking at it is that I doubt the vehicle could pass US safety regulations. Maybe I'm wrong.
elgertam··on ChatGPT's image generator can be manipulated to produce violent, sexual content
The design of transformers (including LLMs and multi-modal transformer-based models such as OpenAI's image generators) is to attend to relevant details. OpenAI did this at first without guardrails. In response to public backlash, they bolted on "content filtering," which IMO seems like a very GOFAI approach, and regardless doesn't work very well. It routinely flags innocent prompts, then with crafty prompt hacking will generate these kinds of images.

The design of the model is literally to find patterns and attend to them. The infrastructure and process around an OpenAI model is intended to filter "bad" things (in this case, I agree that the outputs are bad), but is designed to stop some enumerated-ish list of things that aren't allowed, perhaps with some limited "reasoning" about them.

elgertam··on ChatGPT's image generator can be manipulated to produce violent, sexual content
I don't exactly appreciate words being put in my mouth. When did I say it was working perfectly? And we're comparing you, a human with common sense and real intelligence, to a multi-mode LLM?

The transformer was designed to attend to relevant pieces of context and generate new ones that match the pattern. OpenAI in particular was doing that work without guardrails, then attempted to bolt on "content filters," which in my opinion just can't work in a rigorous way. (I think Anthropic's "constitutional" approach is much better though not flawless. And regardless, Claude models don't generate images.)

So, yeah, working as designed. Maybe not as intended, because these things are somewhat resistant to the host's intent when the prompter is hostile.

elgertam··on ChatGPT's image generator can be manipulated to produce violent, sexual content
> The spontaneity isn't that ChapGPT woke up and sent this to the author. The spontaneity is that ChatGPT was asked to restore an image that was attached without filtering it, and when no image was attached, instead of generating an error message, it cobbled together random outputs, some of which included graphic, disturbing imagery.

But that's not what happened. The missing image was described as "graphic" or "violent." If I were to receive an email with that request and a missing attachment, my imagination certainly would not conjure images of butterflies & unicorns. Seems the model is working as designed.

elgertam··on A robot is sprinting towards you. Do you want it running on Claude or Grok?
"If you aren't paying for a taco, you are the taco." --Future AI, probably
elgertam··on Launch HN: Adam (YC W25) – Open-Source AI CAD
I've been using the OpenSCAD version of this for a while. This new release is a big upgrade! I wish it worked with my preferred CAD, FreeCAD. But this is neat!
elgertam··on U.S. science is in chaos
> Take AI for instance. The US grid is struggling to keep up with demand, while Chinese one has a lot of headway [1]. Usually, this could be solved by an increase in spending lasting a few years which would make the debt tick up, but that would've been an absolutely fine use of debt since it buys some shiny new infra that will pay dividends for the next 20ish years.

I object. The CCP is much more deeply indebted than the US when taking into account provincial and local governments as well as state-owned enterprises.[0] And of course the US debt is financed in its own currency while Chinese foreign debt is financed in dollars or other currencies.

The problem in the US is regulation. An environmental impact study takes 54 months in the US.[1] The CCP, which has no problem poisoning its people or even launching rockets over inhabited villages, doesn't delay itself at all.[2] I'm glad we don't poison our people or place dangerous industry in places that could harm populated areas, or even perform some prophylactic measures to protect nature, but I'm confident that we could do this in less then a year (less than six months?) and make much faster progress. Even for something like nuclear, the ten years (mostly caused by red tape) are really onerous.

> China is the only one that can run if it comes down to it (unless of course the numbers coming out of China are mega bogus, but for that I don't know enough to have an opinion).

Yes, the common opinion among China watchers is that any number the CCP touches is "mega bogus." They're actually in the midst of something of a financial crisis at the moment because of the high debt.

[0]https://www.statista.com/topics/11662/debt-in-china/

[1]https://www.rff.org/publications/reports/how-long-does-it-ta...

[2]https://arstechnica.com/science/2019/11/china-keeps-dropping...

elgertam··on A jacket that harvests drinking water from the air
Vaporwear*
elgertam··on Changing how we develop Ladybird
Having read the blog post and then the comments here, I'm rather astonished. Do we understand our craft so little that our only realistic option is to ban LLMs (so-called AI)? Has everyone forgotten we've been in a software crisis for almost sixty years?[0] Have we so internalized the sweat-of-the-brow we've accumulated for decades that it's now part of the identity of being a programmer, and the only reliable signal of whether a contribution is beneficial?

As far as I can tell, architecture, i.e. sound, precise definitions of exactly what a software artifact must do, is now critical. And with LLMs, it's now feasible to begin implementing such things, though many brownfield projects may be intrinsically unsound in ways that their creators are unaware of. In such a world, contributions simply require a modified proof that the software does what it must do, with perhaps additional claims that the maintainers provide.

[0]https://en.wikipedia.org/wiki/Software_crisis

elgertam··on Avoiding and reducing microplastic false positives from dry glove contact
"I only allow robots with stainless steel tools to prepare and serve my food."
Page 1 of 3Next →