HNHacker News
TopNewBestAskShowJobs

martinald

7,261 karma · joined May 7, 2013

Feel free to reach out: martinalderson AT gmail DOT com

meet.hn/city/gb-Cardiff

submissionscomments
martinald··on Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers
If you set aside political menace, this is a huge problem with Anthropic's strategy.

You _cannot_ say that Mythos is super dangerous and can only be rolled out to certain people, but then release Fable with anything other than bulletproof cyber denials.

Clearly with LLMs, bulletproof denials are ~impossible due to the way LLMs work.

So you've ended up in a situation where Anthropic are simultaneously claiming it's a incredibly dangerous model _and_ there are (minor, potentially) problems with the security "protections".

As technical people we understand that nothing can be perfect, esp in LLM world. But all my non technical friends were really confused how they had managed to make the model "safe" so quickly when it was released and the general sentiment was it shouldn't have been released - and now to an outsider I think it looks like it was never safe at all to release, so I can totally see how the current US administration have got themselves very upset with it.

_Even if_ there was no political bad will, it's a bit of a silly scenario to end up in, and really quite easily foreseen.

martinald··on The EU Open Source Strategy
Well you can do that right now with Chrom(ium)OS.
martinald··on Google to pay SpaceX $920M a month for compute capacity at xAI data centers
Keep in mind Google also rents GPUs via GCP, so they could be just reselling these to GCP customers?

Thing is though, Anthropic was really against the wall with lack of compute pre xAI deal. And tbh, Gemini reliability has been abysmal which probably points to real compute shortages.

And nearly _every_ major DC project is really up against it with massive delays, etc. Stargate UAE has been badly affected by the Iran conflict.

So maybe long term this isn't a great business, but _right now_ I'm not convinced it's all financial engineering. There is a enormous shortage of compute and xAI has a load of it _available now_.

martinald··on Gov.uk has replaced Stripe with Dutch provider Adyen
Why do you think that? People definitely pay taxes by card.

But regardless this contract is _not_ for HMRC payments, its for gov.uk pay which is basically a centralised service that other services can use.

martinald··on Anthropic confidentially submits draft S-1 to the SEC
Don't think so - the 3x is a separate cap. It actually reduces it down from market cap.

Eg say spaceX has $50bn of float at $1.5T valuation. If there wasn't _any_ cap at all, the full $1.5T would be used as the market cap. With the (new) 3x cap, it means only $150bn of the $1.5T valuation is taken into account in the index weighting.

Before this change, SpaceX wouldn't clear the 10% requirement to be listed in QQQ at all. So the 3x basically allows them to be included but _does not_ increase their market cap from $1.5T to $4.5T.

Btw, for clarity, I'm not saying there isn't questionable behaviour going on here. My main point is that even if SpaceX, openai and anthropic all went to 0 (unlikely IMO), it's not going to have a material impact on people's retirements which is what OP was proposing.

martinald··on Anthropic confidentially submits draft S-1 to the SEC
But the US has never had $1T+ IPOs before. And also a huge amount of enormous private companies that don't want to go public for various reasons.

Also, the rules have changed before. It's not the first time these rules have changed.

I see both sides of the argument (it's definitely _not_ good for 401k investors if Anthropic/OpenAI/SpaceX make huge leaps in technology that allow for far higher earnings that they aren't able to access, for example).

But my main point is that these investors regardless would "only" have 5% exposure to these. That surely cannot be considered a systemic risk that the OP is inferring.

martinald··on Anthropic confidentially submits draft S-1 to the SEC
Is it? I thought the idea was diversity of risk, not "mitigating risk". You clearly don't want 100% of your 401k in OpenAI or Anthropic. But you probably do want 1 or 2% of it in, to give you the long term growth potential?

Regardless SPY is actually a pretty "risky" index fund on some measures - it pays a (very) low dividend compared to many other intl/ETF funds and is weighted very heavily towards tech stocks (atm).

If you genuinely wanted to mitigate risk you would probably not choose SPY.

martinald··on Anthropic confidentially submits draft S-1 to the SEC
Let's get it in perspective though. The S&P500 market cap is currently $70T.

Assume that Anthropic, OpenAI and SpaceX all IPO and get included in SPY with the new fast listing rules. They are likely to be worth $3-4T combined, which means 'retail' investors are going to have perhaps 5% of their portfolio in it.

_Arugably_ that's a pretty fair allocation for retail investors to have to these "moonshot" style companies.

Also - if any one of these IPOs don't go well; I suspect the other(s) will have to postpone, further reducing exposure.

martinald··on Rotary GPU: Exploring Local Execution for Large MoE Models Under Limited VRAM
Why is this a paper? It's just using the n-cpu-moe option on llama.cpp? What am I missing here?
martinald··on The mysterious Hy3 LLM is topping OpenRouter Model Rankings by a large margin
Hi! Big fan of OpenRouter and the data you provide. It'd be awesome if you would consider providing volume of tokens per hour, mostly for my own curiosity as to quite how peaky demand is.

Thanks!

martinald··on Ruby vs. Java vs. TypeScript: my experience on building a Cowork DOCX plugin
How is it surprising to people that zip and XML are in stdlibs for a programming language?

Btw, you should have looked at dotnet for this as well. There is a very good library ( DocumentFormat.OpenXml) that can handle all docx/xlsx/pptx files. And dotnet can ship standalone binaries (though AOT probably won't work).

martinald··on Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%
They can (not entirely sure how 'grey' market this is) either have subsidiaries outside of china (eg: singapore) that provide the inference and/or just rent it off the public gpu clouds.
martinald··on GitHub Actions down again today
Totally agree. There's days (or even afternoons) where I trigger more actions than I would have done in a month.
martinald··on I love Linux, but I can't quit Windows
I use sudo -A with some openssh ui for sudo. I tell the agent to use sudo -A for anything that it needs and then it pops up with a sudo password prompt.
martinald··on I love Linux, but I can't quit Windows
Two thoughts (I was in the same situation, constantly trying desktop Linux then pinging back to Windows after hitting issues).

1) Fedora is really worth a try, it's extremely polished. The best thing is the packages in the repo are generally much more up to date that debian based distros, which maeans less random PPAs to work around it, which cause issues.

2) The biggest change is having Claude Code/Codex able to diagnose and tweak things extremely quickly. If something goes wrong, I ask claude code (in a specific folder with various docs about workarounds) and it goes and fixes it 99% of the time very quickly.

Coding agents being able to fix Linux actually makes it _more_ stable than Windows for me. In my experience Windows is less buggy _in general_ than desktop Linux.[1] However, once you hit random issues you are basically screwed if basic attempts don't work. With Linux you can have a coding agent go thru all the reams of logs to find the issue and even clone the underlying source code to find issues.

[1] For example, there is some ridiculous problem with wayland and notifications on GNOME at least, see this: https://gitlab.gnome.org/GNOME/gnome-shell/-/work_items/358?... which has to be disabled with an extension unless you want to go insane

martinald··on New Claude Code programmatic usage restrictions
Currently, if you use claude -p (non interactive mode) in for example CI/CD, you can use your included subscription tokens.

They are now changing it to be:

You get $20/$100/$200 of "credit" that can be used for claude -p. Problem is, once you are out of that it is the normal API rates (outrageously expensive).

martinald··on Claude Platform on AWS
I get this, but isn't this a complete compliance failure?

What's the point of having all those loops to onboard vendors if you can just buy from AWS marketplace (which AFIAK is not a particularly high bar to achieve for SaaS options)?

Like imagine $POOR_QUALITY_VENDOR. If they go through the normal channels they might get shot down. If they get procured on AWS Marketplace, then it feels to me in many organisations 'its fine', though AWS does minimal checking?

martinald··on AWS North Virginia data center outage – resolved
I wrote this recently which maybe people will enjoy in the same vein :) https://martinalderson.com/posts/august-29-2026-a-scenario/
martinald··on Wi is Fi: Understanding Wi-Fi 4/5/6/6E/7/8 (802.11 n/AC/ax/be/bn)
Powerline is in my experience vastly worse than WiFi in nearly all cases. It's slow, suffers from bad jitter/interference (often worse than WiFi) and the chips run so hot (especially the last gen ones, AV2000 iirc - I believe they don't sell them any more because they overheat and fail, or at least 2/2 of the ones I had did this).

Even with many walls I was getting 300-400mbit/sec on WiFi vs 100mbit/sec on powerline.

martinald··on Wi is Fi: Understanding Wi-Fi 4/5/6/6E/7/8 (802.11 n/AC/ax/be/bn)
Not sure if that's true really - its just 6GHz is a lot less congested in general (for now).
martinald··on Incident with Issues and Webhooks – Resolved
Opus has 1M context now. In my experience it starts getting increasingly dumb after about 700k, but below that it is very usable. I don't think I've ever ran out of context window since they brought that out.
martinald··on GitHub Is Down
Many things at once I suspect:

1. Models have got way better, which means you are far more likely to get something working. I know I used to have little 'tool'/'weekend projects' all the time that wouldn't get off the starting blocks before, now it takes a few minutes often to build them, and once I've built them I tend to want to have them saved on github. Quite how useful they turn out to be is another question though...

2. Related, because the models are a lot better I can generate far more code per unit time. On Sonnet last year I'd have to babysit the model and constantly 'steer' it, which meant a lot of the CC time was actually me reviewing it. Now with Opus4.7 it can often just churn away for 10-30minutes and get something reasonable.

3. Most importantly, just the volume of new users to coding agents - loads of new developers shipping far more far frequently.

4. Many users who were not on github, now signing up and pushing code to it. "Vibe coders" basically who don't have SWE experience and their agent tells them git would be a good idea.

Each of these would be a big increase in scale, but combined it is vvv high

martinald··on Belgium stops decommissioning nuclear power plants
Most of the "danger" from nuclear waste passes in a few years as the most radioactive isotopes decay quickly (which is obvious when you think about it).

Interestingly the US/UK/USSR dumped loads of nuclear waste in the ocean in the 1950s-70s and I recently read that there was basically no trace detectable of any of it.

martinald··on Alphabet Announces First Quarter 2026 Results
Gemini is not reliable whatsoever: https://openrouter.ai/google/gemini-3.1-pro-preview/uptime (the orange chart is either AI studio or Vertex, I suspect AI studio, but it's not good either way).

The reason you don't hear people complaining (esp on HN) is because noone is using Gemini with coding agents. Claude Code, Codex (and IMO OpenCode et al with open weights models) are miles ahead of Gemini CLI/Jules/Antigravity/whatever other coding products Google have.

martinald··on Intel Arc Pro B70 Review
Don't think that's true. The drivers are bad (not sure terrible is fair, they have improved a lot) esp for older directx etc games. But Vulkan support is pretty good and that's all you need for LLMs really.
martinald··on Microsoft to stop sharing revenue with main AI partner OpenAI
Really interesting. Why would Microsoft have done this deal? I'm a bit lost. Sure they get to not pay a revenue share _to_ OpenAI but surely that's limited to just OpenAI products which is probably a rounding error? Losing exclusivity seems like a big issue for them?
martinald··on Show HN: OSS Agent I built topped the TerminalBench on Gemini-3-flash-preview
Yes I understand, but do you not have issues that it drifts out of date and confuses the agents (especially on longer running tasks)?

Like even "full" Visual Studio and Resharper have issues with this. Eg, you start editing file x, 'intellisense' runs, says there are loads of errors... because you haven't finished editing yet.

martinald··on Show HN: OSS Agent I built topped the TerminalBench on Gemini-3-flash-preview
Very interesting! I've often thought static analysis could really help agents (I wrote this last summer: https://martinalderson.com/posts/claude-code-static-analysis...), but despite being hyped for LSPs in Claude Code it turned out to be very underwhelming (for many of the reasons that they can be annoying in a "real" IDE, ie static analysis starts firing mid edit and complaining and cached analysis getting stuck).

Curious to know if this has been an issue with your AST approach on larger projects?

The hash line based numbering is very interesting too (though I see on Opus 4.5+ far far fewer editing errors).

I've often thought that even if model progress stopped today, we'd still have _years_ of improvements thru harness iteration.

martinald··on Our newsroom AI policy
const isAiContent = (str) => str.includes('—');?

:)

martinald··on Prefill-as-a-Service:KVCache of Next-Generation Models Could Go Cross-Datacenter
Maybe I'm missing something in this paper, but this seems to me to be just pretty "standard" caching stuff, albeit:

a) very time sensitive b) huge files c) scoped per user

Sort of reminds me of video streaming on CDNs for live video (but per user)?

I still think the big win is going to come based on time of use/live capacity. In a pure economics sense you want to charge a lot for inference when it's oversubscribed and far less when it's off peak (see electricity markets).

We have seen this with anthropics peak times, but it's very blunt currently. We also saw this with batch processing back in the day, but that breaks down because agents are 'chatty' and need to send new responses ASAP. You can't wait ages for each response - it would take weeks to do a simple agentic task if you had to wait hours between turn.

So I think what we'll see is async agents queued up, that you can then decide when to run them - either 'immediately' for time sensitive stuff (for more $$$) or 'best effort' where they can be scheduled to run whenever the provider wants to (3am say). If you also have diagnostics that usually agent task xyz takes y tokens total you can do far more efficient scheduling of these. This also reduces the amount of KVcache gymnastics significantly, as you can dedicate that agent task to a certain rack and schedule it all efficiently.

tl;dr I think the issues with inference efficiency need to be solved at a higher abstraction level of per agent "task" not purely on a per chat message basis. If you can schedule a load of agentic use cases off peak you don't need to preempt them because there is spare capacity by nature.

← PreviousPage 3 of 34Next →