HNHacker News
TopNewBestAskShowJobs

martinald

7,261 karma · joined May 7, 2013

Feel free to reach out: martinalderson AT gmail DOT com

meet.hn/city/gb-Cardiff

submissionscomments
martinald··on Running Gemma 4 locally with LM Studio's new headless CLI and Claude Code
Just FYI, MoE doesn't really save (V)RAM. You still need all weights loaded in memory, it just means you consult less per forward pass. So it improves tok/s but not vram usage.
martinald··on Tell HN: Anthropic no longer allowing Claude Code subscriptions to use OpenClaw
if (process.argv.includes('-p')) and then setting a different http header?
martinald··on How the AI Bubble Bursts
Yes I wrote a detailed article about this Forbes claim. https://martinalderson.com/posts/no-it-doesnt-cost-anthropic...

Key points - if you compare it to openrouter costs for ~similar sized models it is ~90% gross margin.

And this claim came from Cursor - not Anthropic!

martinald··on Walmart: ChatGPT checkout converted 3x worse than website
Would be interesting to know for other retailers though and how much of this is down to what Walmart sells?

I'm confused by the comment that it failed because it forced single item purchases. Most of my "ecommerce" use is researching and buying one item at a time.

martinald··on GitHub appears to be struggling with measly three nines availability
I wonder how much of this is down to the massive amount of new repos and commits (of good or bad quality!) from the coding agents. I believe that the App Store is struggling to keep up with (mostly manual tbf) app reviews now, with sharp increases in review times.

I find it hard to believe that an Azure migration would be that detrimental to performance, especially with no doubt "unlimited credit" to play with?

You can provision Linux machines easily on Azure and... that's all you need? Or is the thinking that without bare metal NVMe mySQL it can't cope (which is a bit of a different problem tbf).

martinald··on Mayor of Paris removed parking spaces, reduced the number of cars
It depends though. At least in London a lot of cycleways were made by removing bus lanes and replacing them with high quality segregated cycle lanes.

This has led to a big increase in %age terms of cyclists in London, but a fairly significant decline in bus passengers.

I think roughly 300m/yr cycle journeys were added, but bus has lost 500m pax/yr (mainly because of increased congestion making them less and less attractive). Note this isn't all down to bus lane removal, but it's a significant part of it.

martinald··on Wayland set the Linux Desktop back by 10 years?
No I do get that, it's definitely been a slow and painful migration. But just having a very insecure X11 "forever" with no fractional font scaling wasn't a long term plan either imo.
martinald··on Wayland set the Linux Desktop back by 10 years?
FWIW I recently switched full time to Linux and have had absolutely 0 problems with GNOME, Wayland and Fedora, though I am using an AMD GPU.

wl-copy works fine, askpass works, copy and paste works, screen sharing with Google Meet works, drag and drop works. Using an iphone as a webcam works as does recording my screen.

Most importantly using multiple monitors with fractional scaling works perfectly. AFIAK this is not possible to do well (at all?) on X11, which is a complete show stopper for me.

If anyone's reading this and sitting on the fence, I would really give Fedora a go. I've found it so much more polished than Ubuntu, and loads of things which didn't work on it work out of the box on Fedora (at least compared to 24.04 LTS).

martinald··on Claude March 2026 usage promotion
Very interesting. As I wrote in this article https://martinalderson.com/posts/is-the-ai-compute-crunch-he... a couple of weeks ago:

"One thing I really suspect we'll see a lot more of is much more generous rate limits at 'off peak' times - likely to be early morning UTC - as there is no doubt a lot of "idle" compute sitting there"

I strongly suspect this will end up in the opposite happening - where peak tokens are far more "expensive" (whether that be thru usage limits of API costs) than off-peak.

PS: Anthropic have managed to improve reliability but are absolutely shredding opus tok/s at peak times. It absolutely crawls on the web (maybe 2-3 tok/s?) and I believe that on non-max plans it's also incredibly slow on claude code.

martinald··on No, it doesn't cost Anthropic $5k per Claude Code user
Hi, OP here! Even if Qwen wants to run at a loss, why would Together, DeepInfra, SiliconFlow, etc _all_ also want to run at a similar loss?
martinald··on WSL Manager
It's interesting because I'm the same in so much that I use windows basically as a WSL2 host and not much else. I use a MacOS a lot.

_However_, still find the Linux desktops that I've tried are too buggy. While the hardware support is incredible (compared to Windows out of the box), I constantly hit bugs with fractional scaling on multiple monitors. I'm hopeful that Ubuntu 26.04 may finally iron out the last problems with this. The latest version of Fedora I installed did fix all this but I'm far too used to Debian based OSes.

martinald··on Better JIT for Postgres
Yeah, the other problem is I've really struggled to have postgres use multiple threads/cores on one query. Often maxes out one CPU thread while dozens go unused. I constantly have to fight loads of defaults to get this to change and even then I never feel like I can get it working quite right (probably operator error to some extent).

This compares to clickhouse where it constantly uses the whole hardware. Obviously it's easier to do that on a columnar database but it seems that postgres is actively designed to _not_ saturate multiple cores, which may be a good assumption in the past but definitely isn't a good one now IMO.

martinald··on Better JIT for Postgres
Anything jsonb in my experience is quickly CPU bound...
martinald··on You Want to Visit the UK? You Better Have a Google Play or App Store Account
Because for many people with poor eyesight, poor English or computer literacy tapping a passport is far easier than typing the data in with no risk of transcription errors.
martinald··on You Want to Visit the UK? You Better Have a Google Play or App Store Account
You are literally sharing biometric passport information with the government for an ETA in this app. Information sharing is the whole point.
martinald··on You Want to Visit the UK? You Better Have a Google Play or App Store Account
But there's really good reason for this. On the app it can use NFC to read your passport data exactly. Until WebNFC supports reading passports, it is a much more efficient way.

It's not like they are getting some long term benefit of having the app on your phone. It's just because WebNFC can't read passports.

martinald··on Making MCP cheaper via CLI
But this is just the nature of LLMs (so far). Every "conversation" involves sending the entire conversation history back.

The article misses imo the main benefit of CLIs vs _current_ MCP implementations [1], the fact that they can be chained together with some sort of scripting by the agent.

Imagine you want to sum the total of say 150 order IDs (and the API behind the scenes only allows one ID per API calls).

With MCP the agent would have to do 150 tool calls and explode your context.

With CLIs the agent can write a for loop in whatever scripting language it needs, parse out the order value and sum, _in one tool call_. This would be maybe 500 tokens total, probably 1% of trying to do it with MCP.

[1] There is actually no reason that MCP couldn't be composed like this, the AI harnesses could provide a code execution environment with the MCPs exposed somehow. But noone does it ATM AFIAK. Sort of a MCP to "method" shim in a sandbox.

martinald··on Making MCP cheaper via CLI
MCP defines a consistent authentication protocol. This is the real issue with CLIs, each CLI can (and will) have a different way of handling authentication (env variables, config set, JSON, yml, etc).

But tbh there's no reason agents can't abstract this out. As long as a CLI has a --help or similar (which 99% do) with a description of how to login, then it can figure it out for you. This does take context and tool calls though so not hugely efficient.

martinald··on Show HN: Emdash – Open-source agentic development environment
Thanks. Btw, doesn't work at all for me. I installed, tried to connect to my WSL2 instance on localhost via SSH, which worked. Selected a folder and got Claude Code is not installed (it is very much installed :)).

Then tried running the Linux version on WSL2 (not ideal because the wayland server on WSL2 is slow) - doesn't work. This 404s: https://github.com/generalaction/emdash/releases/download/v0...

Grabbed the version before and got "PTY unavailable: ... was compiled against a different Node.js version using NODE_MODULE_VERSION 127, this version requires NODE_MODULE_VERSION 123".

Hope you can fix the bugs. I love Conductor on my Mac, but I need something for my WSL2 machine. Ideally Windows which can SSH into WSL2 (for UI speed) or runs on Linux itself. This is very close to what I need if you fix the bugs :).

martinald··on Show HN: Emdash – Open-source agentic development environment
Please codesign your Windows installer exes :)
martinald··on Google restricting Google AI Pro/Ultra subscribers for using OpenClaw
This seems unlikely while we have open weights models available that are ~as decent as the frontier ones.

Given the API prices for open weights models of similar size are 5-10x less than the frontier models the APIs are very profitable on a pure unit economics approach. I strongly suspect they make money off their monthly plans as well.

martinald··on Google restricting Google AI Pro/Ultra subscribers for using OpenClaw
It's a fair point, but I think people are thinking too much about 'cost' and 'subsidies' and just the fact that everyone is so compute stretched.

While it's sort of the same thing, I think it's much more a symptom of not enough compute vs some 'dump cheap tokens' on the market strategy.

One related thought I had was that given OpenAI is the only one _not_ doing this of the big3, it probably indicates they have a lot more spare compute.

It doesn't make sense to me that given the absolutely brutal competition any of these companies would block use of 3rd party apps unless they had to. They clearly have enough cash, so I don't think it's about money - I think it's that an indicator that Google and Anthropic are really struggling with keeping up with demand. Given Anthropics reliability issues last week this does not surprise me.

martinald··on A16z partner says that the theory that we’ll vibe code everything is wrong
I'm not saying _the end user_ clones it. I mean someone else does (more efficiently with agents) and runs it as a _new_ SaaS company. They would provide support just like the existing one would, but arguably at a cheaper price point.

And regarding agents being non deterministic, if they write a bunch of SQL queries to a file for you, they are deterministic. They can just write "disposable" tools and scripts - not always doing it thru their context.

martinald··on A16z partner says that the theory that we’ll vibe code everything is wrong
Thanks,yes exactly what I think.

Or an industry specific Workday, with all of workdays features but aimed at a niche vertical.

I wrote about this (including an approach on how to clone apps with HAR files and agents) if you are interested. https://martinalderson.com/posts/attack-of-the-clones/

martinald··on A16z partner says that the theory that we’ll vibe code everything is wrong
I sort of agree with this, but what a lot of people are missing is it's unbelievably easy to clone a lot of SaaS products.

So I think big SaaS products are under attack from three angles now:

1) People replacing certain systems with 'vibe coded' ones, for either cost/feature/unhappiness with vendor reasons. I actually think this is a bigger threat than people think - there are so many BAD SaaS products out there which cost businesses a fortune in poor features/bugs/performance/uptime, and if the models/agents keep improving the way they have in the last couple of years it's going to be very interesting if some sort of '1000x' engineer in an agent can do crazy impressive stuff.

2) Agents 'replacing' the software. As people have pointed out, just have the agent use APIs to do whatever workflow you want - ping a database and output a report.

3) "Cheap" clones of existing products. A tiny team can now clone a "big" SaaS product very quickly. These guys can provide support/infra/migration assistance and make money at a much lower price point. Even if there is lock in, it makes it harder for SaaS companies to keep price pressure up.

martinald··on Cloudflare outage on February 20, 2026
Just crazy. Why does a staging environment matter? They should be running some integration tests against eg an in memory database for these kinds of tasks surely?
martinald··on Claude Code's compaction discards data that's still on disk
AFIAK claude code includes _all_ messages you sent to the LLM in compactation (or it used to). So it should catch those bits of nuance. There is so much nuance in language that it picks up on that is lost when writing it to a plan.

Anyway, that's just my experience.

martinald··on Claude Code's compaction discards data that's still on disk
Yes I think the same here tbh, hard to keep up with.
martinald··on Claude Code's compaction discards data that's still on disk
Well a few things.

Firstly, it's very useful to have your (or at least some) previous messages in. There's often a lot of nuance it can pick up. This is probably the main benefit - there's often tiny tidbits in your prompts that don't get written to plans.

Secondly, it can keep eg long running background bash commands "going" and know what they are. This is very useful when diagnosing problems with a lot of tedious log prepping/debugging (no real reason these couldn't be moved to a new session tho).

I think with better models they are much better at joining the dots after compactation. I'd agree with you a few months ago that compactation is nearly always useless but lately I've actually found it pretty good (I'm sure harness changes have helped as well).

Obviously if you have a total fresh task to do then start a new session. But I do find it helpful to use on a task that is just about finished but ran out of space, OR it's preferable to a new task if you've got some hellish bug to find and it requires a bunch of detective work.

martinald··on Claude Code's compaction discards data that's still on disk
I'm fairly sure that Claude adds a note where it can find the original transcript after compactation (somewhat recently).

Fwiw I built a little CLI that could help with this, https://github.com/martinalderson/claude-log-cli. It allows Claude to search its own logs very efficiently. So I'm sure you could add something like "if the session is continued from a previous one, use claude-log cli to find users original prompt with claude-log" which would pull it out very efficiently. I built it to enable self improving claude.md files (link to the blog in the GitHub) but it's so useful for many tasks.

← PreviousPage 5 of 34Next →