HNHacker News
TopNewBestAskShowJobs

martinald

7,257 karma · joined May 7, 2013

Feel free to reach out: martinalderson AT gmail DOT com

meet.hn/city/gb-Cardiff

submissionscomments
martinald··on Nvidia dramatically reduces amount of OpenAI infra financing it may guarantee
Nothing would really change IMO? 99% of users don't have anything like a RTX5070 (mobile especially).

Even if it did, it still doesn't make much economic sense running a model locally vs on a datacentre.

For example, I managed to just about squeeze a Q2 quant of Qwen 3.7 27b on my 9070XT. I get around 60tps decode (slightly faster prefill). _but_ it uses 300W of power to do so. At UK electricity rates of 30c/kWh this works out at something like 42c/MTok. I can get far far better models on openrouter cheaper than that, plus I'm not horrendously constrained on context length.

martinald··on Qwen 3.8 27B
You can run these on CPUs at a somewhat reasonable speed.
martinald··on Why does Opus 5 feel worse to work with?
Keep in mind all this kind of stuff can make the model less capable. If it has to think in "plain" English, it may well be squashing quality of code etc output.

I'm not sure how true this is, but when using "forced" json output it def had a big drop off in quality - https://arxiv.org/html/2408.02442v3.

I think you're better not fighting it with hacks like this and find a different model.

martinald··on DeepSeek peak/off-peak pricing update
Why would it save the stock market? Cheaper models if anything transfers more value to hardware companies and datacentre companies. The two companies that would be most affected are OpenAI and Anthropic, which aren't public.
martinald··on Grok 4.6
Yeah I've been sorting of amazed how polished Grok build is. It's also super fast (written in Rust).
martinald··on Grok 4.6
There's also a bit of selection bias going on here because we forget about labs that don't have a jump and just focus on the ones that do. Notably Google is definitely not having that capability jump.
martinald··on Muse Code and Muse Spark 1.2
Also looks incredibly fast. 150tps on openrouter (nearly all deepseek providers are around the 50tps mark).
martinald··on Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers
It depends. If you are doing blog content with the idea of upselling users to your product, probably not as much (because the LLM can just give the user the answer).

However, if you are looking for the best product/service/whatever, then yes it really does matter. I've bought _so_ many products because of LLM recommendations. For example, I wanted a new webcam, I asked the LLM to find me the best ones with a large sensor and Linux compatibility. It gave me a shortlist then I chose one and then I bought it.

This experience is far better than trailing through dozens of pages of (even pre LLM) SEO slop.I just tried the same on Google search and all the links recommended a camera with ~10% the sensor size that I bought.

martinald··on Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers
That's not _entirely_ true. Search Console now has a Generative AI page where you can see impressions per page now. Bing has something similar.

Also, I would assume that really "GEO" is just like "SEO". If you rank on the '1st page' of results for whatever common searches, you are _very_ likely to rank the same way for LLM questions, because all the LLM is doing (nearly all of the time for 'best roofers in Houston') is doing a web search and summarising the first x results. So if you are on page 1 for that term, it's very likely IME that you will get mentioned on LLM answers for that.

martinald··on IMAX vs. IMAX 70mm: The difference between these two cinema formats
It's interesting to me when I go to the cinema how poor the quality is compared to at home on an OLED TV. The blacks are so grey and dark stuff is hard to make out.The resolution is far less noticeable than that (but obviously worse).

Seems to me the real improvement in cinema would be OLED-style contrast ratio panels. Is this a thing being worked on?

martinald··on Claude Is Down
now failing for me, esp auto mode classification
martinald··on Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
I don't think so. According to some very basic research there are around 8bn searches a day, or 250bn a month.

Let's assume Google serves AI overviews on every SERP (they don't) and don't cache them (they do, afiak).

And let's assume that each AI overview is 2000 tokens (blended input/output), that's 500T tokens a month.

It's rumoured that anthropic is serving somewhere close to 10Q tokens a month.

Now it may be that AI overviews uses vastly more tokens than that per search, but I doubt it based on speed to render the overview.

My very rough napkin math on this is that maybe AI overviews is consuming 100T tokens/month max (after adjusting for caching and SERPs that don't have them), which would be 1% of Anthropic token volume.

martinald··on OpenAI and Hugging Face address security incident during model evaluation
If it is marketing it's the most silly marketing of all time. They are under extreme pressure from the US Govt to prove safety and saying "our model escaped" is not ideal.

Perhaps there is some 4D chess going on to get open weight models banned, which may be possible but this is an odd way to go about it imo (it hardly proves the point, unless the point they are trying to prove is that without safeguards the models are too dangerous, therefore open weights are de facto dangerous?).

Having said that the AI companies are not generally very good at PR, so perhaps it is just marketing after all...

martinald··on Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Yes agreed - I wrote this up a while back https://martinalderson.com/posts/whats-going-on-with-gemini/

My view then was they are optimising the models for inference ability on their own hardware AND use cases, which is often speed and time to first token.

They've somehow seemed to end up with terrible compute shortages, which again is surprising given how good Google is at infra deployments AND have their own hardware. From rumors out there they are turning down enterprise deals for Gemini because they don't have the compute.

The problem is they're falling further and further behind on frontier class on coding especially, and since I wrote that article it's got even worse with open weights models undercutting them on price AND intelligence.

martinald··on Kimi K3 is now live
Will be interesting to see how it stacks up pricing wise on the various inference providers.
martinald··on Meta reuses old RAM in new servers with custom bridge chip
GPUs are even more extreme. A 5060 is something like 15,000x faster than a 3dfx Voodoo card from ~2000 by my limited research.
martinald··on Launch HN: Manufact (YC S25) – MCP Cloud
MCP makes a lot, lot more sense when you think of it as as a auth standard and not a comparison with CLIs. It obviously does more than just auth, but having standardised auth (which CLIs definitely do not) is the real 'killer' feature.
martinald··on I Am Behind on C# 14 Features, and I Can't Prove It but Does It Matter?
I don't think that's inevitable with RL.

Imagine in C# you are training the model with RL loops in a harness. One uses C#12 and one uses C#15 (when released), with union types (and importantly - includes the release notes in the harness). Union types if used properly will reduce the amount of bugs/issues in theory from "forgetting" about certain conditions, because the compiler will enforce that better.

In theory, the one with union types will "win" (less errors/fewer edits required) in certain conditions, which makes it more likely to be used going forward.

Basically I think it looks less about 'ingest lots of slop' but 'how do we give our RL harnesses the best possible tools and documentation to make the best* code'. I think this is exactly what good engineering teams do.

For example, if I put 'use C#15 union types' in my CLAUDE.md/AGENTS.md on a .net11 preview project, it is very good at using them when required. It doesn't take much instruction for an agent to use new language features.

_However_ what it does do is change the language feature adoption from 'many developers' to 'eval writers and people that put features into CLAUDE.md'. This obviously changes things massively - though I sort of suspect very few developers _actually_ adopt new language features quickly.

Final thought is that I think we may see a lot of different features being adopted. Instead of what makes code readable to humans, what makes code better on evals. I sort of suspect we'll end up with some Frankenstein language in the future that is difficult for humans to write but agents can write extremely well, with esoteric language features that no (sane) human would think to use.

martinald··on Apple raises prices of MacBooks, iPads
Micron said that they tried to tell 2 of their largest customers (one almost certainly Apple) that the prices they were demanding would result in the cancellation of a lot new construction in 2023, which wasn't in the industries best interests.

It is sort of Apple's fault. They are probably the biggest single buyer of DRAM and NAND globally and they pride themselves on their supply chain management under Cook.

It seems they over optimised this too far.

martinald··on Om Malik has died
Really sad. I grew up reading his writing. I emailed him some thoughts on one of his blog and he immediately replied in a lovely way very recently. What a shock and a loss.
martinald··on Wikipedia cofounder Larry Sanger blocked from editing Wikipedia
Why not? It's the first time many developing countries have had access to high quality internet at an often relatively affordable price?
martinald··on Trains halted across Germany because of communication system problem
Well, the EU insists that track & train operations are separate. (ironically the UK _is_ combining passenger operations and track somewhat back together, which is only possible because of brexit).

The bigger issue tbh is the enormous cost inflation in civil engineering in general. This seems to be a problem everywhere. There's no doubt some of this is caused by material cost increases, labour shortages etc, but I'd say the huge amounts of regulation added over the years is really a core driver of this.

martinald··on AI's Affordability Crisis
Wrote this a while back. https://martinalderson.com/posts/no-it-doesnt-cost-anthropic...

OpenRouter is the best guide to real costs.

martinald··on Wikipedia cofounder Larry Sanger blocked from editing Wikipedia
Yes agreed, for example, there was an interesting table on the starlink page I used to check every so often showing which countries had access to starlink as it was rolled out. Was interesting to see the expansion.

Of course, some editor decided it was 'marketing' for starlink so it got deleted despite loads of people protesting. It was the only source I could find easily for showing which country got starlink when.

A huge list of prose is still on the page (not marketing?) showing the updates in a very hard to read and not comprehensive way. Something is really quite wrong over there.

martinald··on Inference cost at scale with napkin math
The point is that tok/s/GPU stays ~roughly stable. So you need say 4 GB200s minimum to fit the modules, but this provides 4x the tok/s as 1 GPU.
martinald··on Inference cost at scale with napkin math
Yes 32B dense is a weird one to choose.

But in reality, 32B dense is very similar* to 32B activated on MoE in terms of inference costs. And I highly suspect eg Opus is around that level of active params.

A 284ba13b model at scale, is almost certainly cheaper to serve than a 32b dense model.

*as you can shard the model across multiple GPUs at scale. but in reality you have some loss of efficiency from GPU coordination and expert routing

martinald··on Inference cost at scale with napkin math
In general, less for fuel cost alone. But you obviously need to buy the turbines.
martinald··on Ask HN: Will programmers write more efficient code during the memory shortage?
I was thinking about this recently. If you discount web/electron bloat, actually the memory bloat of software isn't hugely terrible.

I still often notice that servers on Linux use <1GB of RAM even with relatively high use. I don't think that's really changed massively in 20 years.

martinald··on How Japan's railways stayed one while splitting apart
Problem with this approach though is it makes the system very vulnerable to political changes. How much of a problem this is in Germany I'm not sure.
martinald··on Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers
I get that, but anyone else releasing a model of similar capabilities has the advantage that they haven't spent the last few months hyping the danger up to fever pitch.
← PreviousPage 2 of 34Next →