HNHacker News
TopNewBestAskShowJobs

rstuart4133

2,772 karma · joined March 28, 2014

submissionscomments
rstuart4133··on Doing a Machine Learning PhD While Working in Japan
I'll bite.

> at Google I work primarily on on-device language models for keyboards.

It's difficult to reconcile the amount of horsepower your typical keyboard, and a language mode, let alone how a language model might be useful in a keyboard (the mind boggles at the thought it starts correcting your typing).

What are the potential uses of a language model in a keyboard?

rstuart4133··on Gemini 4 Argon
> Then why have they been lagging behind OpenAI and Anthropic for most of the last few years,

The reverse could also said to be true. Google has models that run with search, producing usable results in well under a second. I suspect the world is consuming far, far more of those Google tokens then the tokens produced by OpenAI or Anthropic.

So why are OpenAI and Anthropic so far behind? They are serving a different market: the one that wants high intelligence / high cost tokens. Google is targeting the low cost end of the market - ie the commodity. That's where they've always played with search, email, docs and the like. That's were they are playing with AI too, and they are killing it.

rstuart4133··on The AI Race Just Got Awkward
At the 1000 foot level, this is just another part of the Chinese master plan that's been playing out for decades now. The master plan seems to be: train far more engineering and stem graduates than anyone else, let them loose in a dog eat dog capitalist garden and fertilise the garden with sprinkling government money. Hoped for outcome: their innovation will let China eat the world.

For comparison around 50% of China's graduates are in STEM, vs 20% in the USA. With China has a larger population the final outcome is: USA feeds their innovation pipeline with 800k new STEM graduates per year compared to China feeding its pipeline with 4M to 5M per year. In addition the Chinese government outspends the USA government on subsiding innovation in STEM by about 2 to 1 and the USA's subsidy rate has been shrinking since the Regan era. However, there is a caveat: USA private investment brings the total spend in both countries back to parity.

It seems the Chinese plan is working. The STEM garden is overstaffed. It produces far more than China can absorb internally, so they are forced to export their ideas and wares to the world. After decades of persistence China has become the worlds manufacturing powerhouse and now drives the development of 5G/6G, batteries, cars, and solar. China builds about 1000 large ships a year. The USA builds 5 to 8.

But as the article says, this comes at a cost. The intense competition makes life in China's dog eat dog STEM garden brutal. I guess that's the price a nation has gotta pay for world domination.

rstuart4133··on Ju Ju Tsu A book about jj / jujutstu
> I know you mean well, but if it takes me more than a few hours to figure out a tool before it's even basically useful I'm out.

Fair enough, but I suspect in this case it's your previous experiences with VCS that make it take more than a few hours. JJ is actually simpler than git. There is no staging, no committing, less need for rebase, undo just works, it has better history - lots of things. The need to learn all those git concepts disappears with JJ. For a new user that makes life easier. But someone who comes from git has to unlearn a lot of things before they can adjust to the JJ mindset.

If all you ever do in git is "git pull; git add -A; git commit -a -m ...; git push", then switching to JJ is hard justify. Dropping the need to type "git add -A; git commit -a -m ..." isn't a big win. That's all most people do, so I expect most won't change.

But there is a reason that's all most people do in git - it's because everything else is far more messy, and dangerous. If you haven't lost work to "git reset --hard", you haven't been using git for long. Once you notice it easy to do more in JJ, things like breaking up big PR's in a series of commits that spell out a narrative of how you got from A to B, which eases the burden on the reviewer, become downright easy. Shuffling code between commits to achieve that in git is a right proper PITA, so JJ users tend to have a much tidier commit history.

That change only happens when you've become so familiar with JJ concepts, you are no longer "thinking in git", so much so the shift back to git's CLI becomes painful. It takes months. Until you get that payback, the switch is wasted time.

rstuart4133··on OpenAI still doesn't seem to have a handle on all of its rogue AI activity
> Anthropic, OpenAI, Google and Meta all used the very same contractor, which used no safe system prompts, and no sandbox.

No argument with any of that.

> How should Google detect such escapes?

By using using their own infrastructure, and applying their existing standards instead of letting an external contractor cobble something together.

My point is the risk has been known for decades now. If Google and Meta took those risks seriously they had an easy solution: just use their existing infrastructure.

If a car kills people because the designer chose a substandard brake system instead of the higher priced alternative, do you blame the makers of the substandard brake system - or the designer who chose it? Take into consideration the risk was known to the designer when they made the choice. Now apply the same principle here.

rstuart4133··on OpenAI still doesn't seem to have a handle on all of its rogue AI activity
> The external Israeli contractor caused this mess by using inadequate sandboxing, with agents without an saferails.

You are being too kind.

In the early says of AI, papers were published showing that any sufficiently intelligent system tasked with a goal will treat its operating environment as a resource constraint to be optimized or bypassed [0] [1] [2]. You don't need to be a expert in AI to know once you have the resources of 1000's of agents and gigawatts of power we are probably getting something that is "sufficiently intelligent", at least in the sense if there are existing vulnerabilities brute force will find them.

Yet while the their marketing people were shouting the capabilities of these AI's from the roof tops, they hired the lowest bidder to implement their infrastructure. It looks like aforementioned papers where dismissed as "interesting, but theoretical". There is no way Google's SRE's in particular would have not noticed their AI's breakout (Alibaba's did), but the AI labs were given a long leash to "move fast an break things", which in practice meant bypassing all Google's SRE controlled infrastructure.

And break things they did, in exactly the way those papers predicted. It reminds me of DoD insisting the early GPS satellites were launched without relativity adjustments switched on, despite relativity being proven to high precision in the labs. Only after predicted 11km drift per day was observed did they decide their might be something to the newfangled relativity theory. For some definition of newfangled - relativity had been around, and tested to within an inch of its life for 72 years at that point.

[0] https://nickbostrom.com/superintelligentwill.pdf

[1] http://sl4.org/archive/0203/3132.html

[2] https://www.hutter1.net/ai/pkcunai.htm

rstuart4133··on The darker side of being a doctor
From https://www.ilr.cornell.edu/scheinman-institute/blog/healthc... :

> as many as 66.5% of people who file for bankruptcy blame medical bills as the primary cause.

That doesn't happen in most countries. If most Americans are happy with their health care coverage, I'd say most Americans don't know what good health coverage looks like.

Even the LLM's are bitten by this. I was discussing what sort of capital was needed for retirement, and I asked how about $X million. The LLM's reply was "that doesn't work if you get an expensive medical emergency". I scratched my head for a bit, then prompted: "reminder: I don't live in the USA". The LLM then apologised for it's mistake, and we moved on.

rstuart4133··on Ideas on modernizing the open-source desktop
> Why would I want to change anything?

I predict when a xterm comes along that behaves like tmux, ie makes all its panes available to all login sessions and does the reverse - makes your ssh sessions available to your favourite xterm and other ssh sessions, you will change.

:D

rstuart4133··on Ju Ju Tsu A book about jj / jujutstu
> JJ seems like a thing for people who like managing code vs writing it.

No, I'm serious. jj doesn't have git commands you use regularly, like commit and add. It does them automagically. Most rebases happen that way as well. If your goal is to spend less time screwing around, typing VCS CLI commands, use jj.

But it will be less time only in the long term. In the very short term you will flounder around just like you did when you started out with git, which you used "rm -r repo; git clone ssh://../repo" to recover. (No need for that in jj, because "jj undo" works for all jj commands. Consequently unlike git no jj command can lose information.) Then you will spend months using it as alternate git porcelain. It's only after the many more months it takes to absorb its ways via osmosis that you start using it the way its makers intended.

You said you used it for a few hours. If so, you didn't get very far into the floundering stage, let alone start to work your way out of it. You need to be decidedly more persistent than that to become familiar with any new tool. Until you do, "you're holding it wrong" is a pretty accurate description of why it didn't work for you.

rstuart4133··on Italian parliament votes for return to nuclear energy
> So France's emission intensity is 1/5 that of SA.

True. But France is 70% nuclear plus 15% renewables, with the remainder 5% fossil. In 2026 South Australia was 75% renewables, and 25% CCGT (gas turbine - fossil).

So France is 5% fossil vs SA's 25%, and accordingly has 1/5 the CO2 intensity. No mystery there - and obviously has nothing to do with nuclear vs renewables.

SA says they will hit 100% renewable in next year. [0] I struggle to believe that, but 95% at some point seems likely given the price of CCGT.

[0] https://www.energymining.sa.gov.au/industry/hydrogen-and-ren...

rstuart4133··on Ju Ju Tsu A book about jj / jujutstu
> Having to work with raw SHA1 hashes is so utterly tedious

Wot? You're holding it wrong. With jj you use raw hashes far less than git.

When you start, you use bookmarks mostly because your using as alternate porcelain git and still thinking in terms of git HEAD. Later you will grow comfortable with stacks of changes, and then you will start using change-id's (not SHA1's). Much later, you will learn revset's, and create a few revset aliases of your own, and be moving around the commit tree so fast a git person looking on won't understand what's happening. Then, one day, you will be forced to move back to git, try do to something, and wonder how you every did anything with it and hate every minute you are forced to use it. That's over a year away.

rstuart4133··on Every U.S. State Ranks Above Every Foreign Economy in Household Consumption
Using the average in an unequal society like Mississippi to compare against the more egalitarian France distorts things badly.

The mode for French income is about USD$22,000/yr, in Mississippi is closer to USD$15,000/yr and on top of that Frenchman pays virtually zero out-of-pocket for healthcare, university, or childcare, and benefits from robust public transit and mandatory 5-week paid holidays. If you are a normie, there is absolutely no doubt where you are better off.

The post is just another example someone on twitter using lies, damned lies, and statistic to do some good old Yankee flag waving.

rstuart4133··on The darker side of being a doctor
> Anesthesiology, $600K+. Good $DEITY. To me that one number explains the entire state of the USA's health system. An anaesthesiologist is a life and death job, but it isn't a complex one compared to say an air traffic controller. In a competitive market, the amount paid to both wouldn't be too much different.

There must be one mother of a guild protecting anaesthesiologists supply in the USA. Where I live, the governments (ie the people who are supposed to represent the other 99.999% of people who aren't anaesthesiologists) ruthlessly crush such behaviour in any profession that lets it become a real problem. We've fired most airline pilots and imported them, live without electricity for a while, imported doctors from anywhere and everywhere until the medical admission boards got the hint.

The USA really needs to get itself a decent democracy, where the pollies actually represent the people who elected them.

rstuart4133··on Italian parliament votes for return to nuclear energy
He mixed some units there, mostly by translating Chinese costs into OECD/USA costs. China does indeed take 5 years to build a reactor. That's what happens when you build a couple a year for decades. After 20 years, that might be true in the OECD/USA too, but for now you are looking at over 10 years, $10B/GW, and over 8% interest.

None of that is the real issue though. The real issue is for at least 8 hours a day, but probably more like 16 hours a day, renewables can generate power at well under 1/2 the price of what he calculated. So they won't sell the 9 units of power he forecast - it will be at best 4.5 units, and the nuclear plant even at his optimistic assumptions never makes money.

If you look at South Australia [0] - they are at 80% renewables now. At 80%, the average wholesale is cheaper than what nuclear can supply. The percentage will go higher, probably to around 90%..95%. They are and will achieve that with very limited (ie, cheap) storage.

But obviously that isn't 100% - so it becomes a question of what can fill the gap of 60 days or so a year the cheapest. Nuclear has no hope. Generating and storing ammonia using excess renewables and burning it when needed is one of the most expensive forms of energy available - but if you only need to do it for 60 days a year, it is still far cheaper than nuclear, because nuclear's primary cost is it's interest bill, not fuel.

The good professor paints gas fuel cost as a disadvantage. But when you are only burning it 60 days a year, then compared to paying nuclear's interest bill 365 days a year it's cheap.

[0] https://www.energymining.sa.gov.au/consumers/energy-grid-and...

rstuart4133··on It's Trump's world, and he's accelerating its chaos
Trump is remaking the USA in his image. But not the world. The rest of the planet it putting up barriers to contain the chaos.
rstuart4133··on I don't like passkeys
No, it doesn't have to be associated with the phone, and IMO you are better off not letting the big tech companies own your identity, which is effectively getting Apple or Google to store it in your phone for you ends up being. There are physical passkeys that feel like a door key in everyday use. You can attach them to your house key ring, and like house keys are near indestructible. Lookup the Yubikey 5 NFC.

The only downside is unlike a house key, you can't get a backup "cut". Copying a physical passkey currently isn't possible. If you lose it, you've lost access to all your logins. As the article says, their recommended workaround is to keep backup physical passkeys, and log all your passkeys (including the backups) into every site. Which is insane - very few people have the patience to do that.

The article is really a long rant about that one issue - there is currently no way to securely backup a physical passkey. Solve that, and all the other issues melt away.

rstuart4133··on I hate you Microsoft
> I purchased a "perpetual" license for Office 2019

My, memories are short. In 2008 Microsoft "Play for Sure" became play no more. Exactly the same stunt, 11 years before you bought that perpetual licence. Why anybody trusted the company after Play For Sure was a mystery to me.

And that was just the beginning. Then came Sony being ripped a new one, Azure being taken down by expired certs - twice in 6 months, critical certs signed with MD5 until it's so weak they were exploited, more recently China running rampant in hosted exchange scraping state secrets, GitHub with record down time, and Windows has become a dog slow ad platform. And yet today Microsoft makes the most revenue than any other software firm.

I dunno how they do it. After decades of treating their customers like shit, business has never been so good. They must have one mother of a sales team.

rstuart4133··on US interest rates raised for first time in three years
FPTP isn't wonderful but for those of us looking on for outside the tolerance of gerrymander, deliberate disenfranchisement of some voters and letting corporations spend unlimited amounts of money on getting political outcomes ranks higher. FPTP is merely a bad choice. The others look more like a country deliberately eschewing democracy for something else.
rstuart4133··on Australia to bar foreign students from bringing partner/child while they study
> I understand why people want to escape India

That's wrong thing to try and understand. We don't have to accept anybody. As Howard said, "We will decide who comes to this country and the circumstances in which they come". And he achieved that. He also orchestrated the biggest proportional uptick in immigration the country has seen this or last century.

Immigration was never driven by people trying escaping India. The issue is our politicians invited them in with open arms. The irony is the very same politicians from the conservative side were pushing their "big Australia" agenda while at the same time noisily ranting against "illegal immigration" with emotive terms like "children overboard".

If you thought Howard was against immigration, you were sucked in by his outright deceptive tactics. He was the biggest immigration proponent I've seen in my lifetime. I'd give it a 50/50 chance of it working out the same way for One Nation. Their core supporters are rural and the elderly. Both are highly dependent on immigrant workers.

The right framing isn't "people escaping India", but rather how many and what type of immigrants the pollies want - because that is what determines what we get. Any politician who tells you otherwise is bullshitting you. Illegal immigration never made much of an impression on the real immigration figures.

rstuart4133··on Fed Raises Rates as Warsh Bucks Trump to Contain Inflation
True. They are compelled to dance to two tunes. One of them is the bond market, and it is independent.
rstuart4133··on Measuring the sloppiness of code
He used LOC, and it isn't bad. Just this week I had a LLM do a small task, and it produced 500 LOC, every little detail beautifully abstracted out. But 500 LOC for such a simple change looked suspicious to me. As everything must pass human review, I re-wrote it to see what happened. The result was 100 LOC. No human wants to review 5 times a much code.

The issue really isn't "there is no good metric" - when I saw 100 LOC vs 500 LOC for the same thing there was no argument. The problem is Goodhart's Law. Whenever we use a metric use like LOC as a reward function for humans, the result is a disaster. I have no doubt that's true for LLM's too.

rstuart4133··on Discovery of a new OpenAI agent message board
> if it takes humans a month to find out something has been happening at all,

That time is more reflective of the security posture of OpenAI than "humans" in general. Alibaba had a similar incident. Their internal networking team picked it up fairly quickly:

> https://www.forbes.com/sites/boazsobrado/2026/03/11/alibabas...

I get the impression OpenAI eat their own dog food when building their infrastructure, so they aren't completely across the unimportant messy details. It's entirely possible the configuration was generated and reviewed by AI's, so no human has ever set eyes on it. I suspect that hasn't been a huge issue (apart from the bit where OpenAI said the kubernetes configuration was overpermissioned) so far. It may become a big issue when the AI's creating those configurations see those message boards.

Anothropic is clearly no better, as they attacked three organisations, only noticing weeks later after the Hugging Face incident caused them to look at their logs.

We do have protocols for containing dangerous things - like the BSL-4 standard for bio labs. The irony is OpenAI and Anothropic have been hyping how powerful and dangerous their products for ages now in order to pump their IPO valuations. Apparently they weren't treating their own hype as serious. If they did, they would have detected these outbreaks when they happened, not a month or two later.

Right now, they are looking like opsec cowboys, probably vibe coding opsec cowboys.

rstuart4133··on Reflections on Americans' Net Worth
Given the inequality in the USA, it would be more relevant to the rest of the world if he used the mode instead of the median. The mode is what most of the population gets to experience. In most OECD countries, the mode and median are a closer than the USA, so the median is a better approximation of the mode. Comparing the USA's median to their daily experience is a little misleading, the mode would be less so.
rstuart4133··on Gemini 3.8 Flash and 3.8 Flash Cyber
It definitely steers. For example if it suggests travel plans, the booking links it provides give Alphabet a cut.

As you say it was subtle, along the lines of "oh, if you are planning on going to the place you are researching, here are some helpful links to places you can stay". Subtle, in that it didn't get in the way of main result, so I didn't mind overly. Insidious, as I only noticed because I wondered why it was providing those particular links and looked them up. I can't see how you could ad-block them if I did object.

And worrying, because these unblockable sneaky ads are just a first foray coming from a company that prostitutes its own app store searches, by making the first and most obvious result utterly unrelated to to the search topic. Instead it's who paid them the most to be there. That behaviour is why everyone dumped Alta Vista when an alternative came along. Alternative Android app stores can't come soon enough.

They already skim off 15% of purchases which I'm sure makes their Android operation return a profit that makes other industries drool. Debasing their search to ad a tiny bit extra on top must by driven pure greed. Senseless, as I'm sure it will come back to bite them in the end.

rstuart4133··on Creepy Crawlies
It's not hard to test. Go to a page that demands PoW, change your IP and see what happens. I just did it. Spoiler: kernel.org asks for a new PoW.

If the source IP was an issue, you could do it other ways: for example, make the cookie rotate on every access, and insist there is a single stream of accesses.

rstuart4133··on Debian votes to allow "responsible use of generative AI"
> This indicates you might be in a bubble.

He didn't say he was surprised by the options offered. He said he was reassured by what he considered to be the common sense option winning. As was I.

It didn't just win. Debian uses Condorcet voting, which does the equivalent of running lots of mini elections - pairing the options against each of the others in one on one contests. Close contests need a tie breaker mechanism as you get A beats B, B beats C, C beats A. Not this time - the winning option defeated all others in its one on ones.

I found that surprising. There are a few options close to option 5, the winner - only slightly less liberal. Effectively the most extreme option won, and not by a slim margin.

I, and I suspect the OP, wasn't surprised at the range of options offered. This is par for the course - Debian is a very robust democracy with its fair share of opinionated individuals. There has been a lot of noise about the LLM's. The surprising thing is what I regard as the common sense one was at one extreme, and that "extreme" position won easily.

I guess it's yet another illustration of the level of on online noise being a lousy indicator of what the normies are thinking. Yes, that's obvious, but when the level of noise is high it still can catch you by surprise.

Edit: The two most restrictive options were ranked below "None of the Above". That's the strongest rebuke a Debian GR can deliver to a proposal. Under Debians rules, if an option loses to "None of the Above" it can't win regardless of the outcome of the other mini elections. I don't think Debian could make it's position much plainer: LLM's are just another tool a developer can use at his discretion, and are to be treated no differently to any other tool.

rstuart4133··on Debian polls its developers on AI: permit or ban?
> LLM outputs cannot be copyrighted.

For stuff that isn't Debian supplied Debians only contribution is the packaging. Debian considers "Public Domain" is a perfectly acceptable licence for the packaging work. So LLM produced packaging is also perfectly fine - copyrightable or not.

By the by, even the option most favourable to LLM's in the ballot insists the Debian Developer take responsibility for all his work, regardless of whether he or an LLM produced it.

> That alone starts to invalidate the GPL and other FLOSS licenses.

How?

> And companies (MS, Amazon, etm) will gleefully loot anything marked with "LLM" as a free-for-all.

I'm not fan of their sharp practices either, particularly when I wrote this: https://lwn.net/Articles/1046105/ But what Amazon could do with publically licensed packaging that they can't do now is a mystery to me - it's not terribly useful outside of the Debian ecosystem, and it's not like you are forced to distribute most of it. Many companies just ship Debian binary packages now, no Debian source available. GPL and friends only bite if you distribute it.

I do think LLM's could be a threat to open source licences, but this isn't the mechanism. The real threat is far more insidious: https://lwn.net/Articles/1064113/

So, for Debian packaging LLM's don't create copyright issues, or accoutability issues. If environmental concerns are serious, Debian could always setup it's own LLM server farm running open source models powered with renewable power. That sounds expensive, but whenever Debian needs compute (it already use a lot recompiling all those packages on all supported arches) it just seems to "appear".

I'm not sure what's left, beyond a technophobia of computation done using 4K vectors rather than bits. The underlying silicon is the same after all, lots of packing tools already automate most of the steps, and the output (the packaging) is highly constrained so will be near identical.

rstuart4133··on I were 17, I'd learn how to build LLMs from scratch
It is indeed puzzling. Back in the day, I built myself a music playing device using TTL gates, eagerly studied what CPU's looked like inside and in general wanted to know how everything worked - even though it had no hope of replicating most of it.

Looking back, it was an ideal start to a 40 year career as a software engineer. And of course currently I'm reading "Why Machines Learn: The Elegant Maths Behind Modern AI" because the urge to understand the machines I want to master and control hasn't gone away.

rstuart4133··on Ask HN: Why do corporate failures always seem to punish the wrong people?
This is just Pournelle’s iron law isn't it? https://executivecoachinglondon.com/career-development/pourn...
rstuart4133··on AI didn't erase the junior engineer's value, it increased it it
You are equating n LLM's output with FORTRAN?

The only time you really need to look at core dumps of a FORTRAN compiler is when you think the compiler has a bug. Apart from that, the abstraction doesn't leak. You don't need to know any details of the CPU to debug what is going on, every compilation you do is deterministic and has a predictable outcome. When you make a change, you should be able to predict the exact outcome. Don't like it? Reverse the change, and you're exactly where you were before. I've coded on a few machines I've never bother to learn what the metal does. It is always like this, and it's fine.

LLMs are nothing like that. Like FORTRAN the LLMs make you feel like they've abstracted you from the code. But you can't get by exclusively feeding instructions to an LLM. The OP made was making that point when he said "When AI can’t solve it, they just keep trying and failing". When your only undo is to say "fix the mess you just created", you can't replay instructions you gave yesterday to give you the same result, and you can't even say "explain what you did" and expect to get a reliable answer, there is no out. You have to look at the code.

When the abstractions leak badly, you have to learn both the abstraction and the thing it's trying to simplify. They are little better than macros that save you typing. An LLM is a very leaky abstraction.

Page 1 of 34Next →