HNHacker News
TopNewBestAskShowJobs

joefourier

1,700 karma · joined September 29, 2016

submissionscomments
joefourier··on Study: Young users (9 to 18Y) ditch Google for AI, with unknown consequences
I use LLMs every day as a search engine replacement, but I’m concerned about this because there’s a vast gulf of capabilities between different models, or even the same models with different reasoning depths.

Back in the day everyone had essentially the same Google, and there was no concept of paying to get better search results. But now the average family is going to rely on the ad-ridden ChatGPT Free tier which is more likely to be wrong, less in depth, with additional commercial incentives. It’s only a minority of adults now that pays for frontier models, and how many are letting their children have access?

You can see this divide right now with medical advice for instance. If you have access to ChatGPT 6 Pro (only available for the $100/month tier I think), it will research Pubmed etc and give you a careful measured answer that likely reflects the latest medical consensus, while a casual user relying on Google AI might be told to eat toilet paper to relieve constipation: https://www.reddit.com/r/medicine/comments/1wmbn87/a_pt_show...

joefourier··on The LLMentalist Effect (2023)
Ah, are you saying that because most people don’t interact with agents, they aren’t aware that LLMs can have initiative and pursue goals?

I think the line is blurring though, mainstream chat interfaces are adding more and more “agentic” features.

ChatGPT will happily execute code in a sandbox, search the web and design downloadable PDFs purely through the standard OpenAI chat interface. They can also send you emails or do tasks on a repeated schedule.

joefourier··on The LLMentalist Effect (2023)
We are way past that point, any harness can trivially make LLMs start conversations or pursue goals. An encyclopaedia wouldn't have hacked Huggingface on its own.
joefourier··on I built non-autoregressive decision models with RL a year ago
There's no need for black and white thinking. Javascript and the internet browser are the most common interface sure, but there's still room for specialised desktop software, especially those that require serious performance like anything to do with 3d graphics or real-time audio.

But also, frontier LLMs are enormously expensive and slow. Using Astra for things like simple text classification is not going to scale, and you're likely to end up in the same boat as those people who saw their Vercel bill shoot up to $96k/week when their site got traction, if not worse.

joefourier··on Show HN: Friday – Self-hosted persistent memory for AI coding agents (MCP)
While the idea of improving codebase search and documentation has some merit, any modern coding agent harness does not have the problem of starting each session blankly and giving generic advice. The actual issue that I have experienced is that it will waste tokens re-reading the codebase over and over again to understand your query, although you can mitigate that by proper use of AGENTS.md and documentation.

But more importantly, is it too much effort to write the readme yourself? Or at least to ask the AI to rewrite it so that it doesn't have all the tell-tale signs of AI slop?

joefourier··on Why I'm still bearish on LLMs after Navier-Stokes
> current frontier models

> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1

The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.

joefourier··on Shopify is moving from React Native back to Swift and Kotlin
The app not existing would unironically be better in so many instances, though. The web version doesn’t take up 1GB of disk space, install persistent services, send you push notifications by default, and it’s trivial to block ads in comparison.

I’d be very happy if companies didn’t artificially degrade their web version to force installation of an “app” that’s effectively a web browser in disguise.

joefourier··on Small Models Have Arrived
> It's especially palpable just looking at the last few years of LLM's: a frontier model with all the world knowledge you can stuff in it and every tool at its disposal has always performed the best at all tasks. Suggesting otherwise has become an extraordinary claim requiring extraordinary evidence.

Absolutely false. At least when it comes to multimodal inputs, even a simple classifier will outperform the largest LLMs who still hallucinate details or don’t describe audio and images accurately.

And there’s also the issue of cost/inference speed. Running a trillion parameter model for all tasks will be incredibly costly, require a cloud API, while a tiny CNN can be run locally or at a cost multiple orders of magnitude lower.

joefourier··on Position: LLMs Can't Jump
What about multi-token prediction and speculative diffusion? That’s a different mechanism of prediction, even if it serves only to accelerate decoding.
joefourier··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
You cannot just "try all possible optimisations". It takes time, effort, and money that could otherwise be spent elsewhere (especially for training, where each training run is especially costly, and optimisations might be promising early on, but cause the final performance of the model to be worse). You need smart people interested in unglamorous work, and if you're swimming in VC money, it's far more straightforward to just throw more GPUs at the problem.
joefourier··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Anthropic's API has two nines availability and Claude Code is a TUI made with React that can regularly consume more than 1GB of RAM, and the codebase is utter slop. They couldn't fix the flickering bug for over a year!

And yet, Fable and Opus are among the best coding models out there (matched only by GPT5.6 Sol).

It's not about the people there being smart or not, it's about their and the company's priorities, resources and what they choose to focus on.

joefourier··on Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
> Like isn't it weird that the 1 million parameter model with the same architecture can't solve basic puzzles but suddenly the 1 trillion parameter can conjure up counter-examples for the Jacobian conjecture?

I'm not sure what you mean? You can see the intelligence of LLMs progress predictably and stably according to scaling laws. LLMs have to encode language in addition to intelligence so there's a minimum bound for them to output sensible text (you can train specialised tiny models to solve basic puzzles without language). Start at around 127M and compare models of increasing parameters and you'll see a clear progression in intelligence.

> It's unintuitive since, to the best of my knowledge, one of the basic tenants of algorithm development was that you can't just brute-force your way towards a solution for some complex problems, e.g. naive sorting algorithms suddenly won't beat quicksort if you put more processing to them

How is that a basic tenet? Simple, easier to parallelise algorithms that have lower memory requirements, or can take better advantage of hardware, or don't hit a plateau the more compute you throw at them, can absolutely beat cleverer algorithms. E.g. brute forcing rendering with Monte Carlo path tracing will give you more physically accurate results than ray tracing or rasterisation algorithms that rely on a bundle of hacks to approximate global illumination, transparency smooth shading, etc.

joefourier··on SpaceX to buy Cursor for $60B
I stopped using Cursor because of how terribly optimised it is (worse than VSCode despite being a fork). It would routinely take up 50% of the CPU resources on my MacBook M4 and gigabytes of RAM for absolutely no reason.

I switched to Zed, and I'm never going back to Electron/non-native IDEs.

joefourier··on How to earn a billion dollars
Why is it unethical? I'm both a freelance engineer and a business owner that sells software, and I've both sold my labour for equity/revenue share, and for a flat hourly rate.

If I charge a client $50k for some software and they made $1 million profit from it, good for them? As long as they pay our mutually agreed upon rate on time and there was no hostile negotiation, why should I feel suddenly entitled to more money if that wasn't in our contract? How do I know how much of the value is from my work and not their marketing or idea?

What you're saying seems as crazy as me saying that someone who bought my software for $99 and used it on a multi-million dollar project is being unethical unless they give me more money. How on Earth does that make sense? Should I be forced to switch to a royalty model? What if I make more selling copies at a flat rate, what if I don't want to have to investigate the finances of thousands of customers and have to deal with that whole trouble?

For me it's the same thing regardless of whether I'm selling my labour or a product. I can choose whether to accept a flat hourly rate, equity, or a mix of both, and usually the better deal is the hourly rate.

If I find a way to hire a software engineer for market rates (say, $200k/year in the US) and get $2M revenue from their work, good? They can ask for a raise or a bonus, we can renegotiate, they can leave if they're unhappy, but I'm not obligated to give them more money than was in our agreement anymore than they're obligated to give me their salary back in the project fails.

joefourier··on How to earn a billion dollars
> My response is “great! Let businesses take a lesson here: give all your employees a chunk of the company. Let’s all share in the success!”

Don't >95% of tech companies offer stock options or equity, from startups to FAANG?

joefourier··on xAI is looking more like a datacentre REIT than a frontier lab
Yeah most of the performance increases have mostly been from architectural improvements like reduced precision tensor cores. AFAIK FP4 is basically the limit for floating point matmuls, after which you need to switch to integer addition if you want to reduce bits, and I don’t think we’ve figured out 1-bit LLMs just yet.
joefourier··on xAI is looking more like a datacentre REIT than a frontier lab
Demand is so high and supply so low customers will go to anyone that has any gear, period. Anthropic is paying xAI for GPUs from 2022, not the latest Nvidia release.
joefourier··on xAI is looking more like a datacentre REIT than a frontier lab
And yet Anthropic is paying xAI over a billion dollars a month for those out of date GPUs in their first datacentre (H100s being nearly 4 years old at this point).

Even A100s are still barely available on the major clouds despite being 6 years old.

joefourier··on Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
Do you think the work will still apply to speculative/alternative decoding methods like MTP and block diffusion, which are making batch=1 decoding less memory bound? Kernel launch overhead and memory transfer become less and less significant as a % of time when computing multiple tokens at once.
joefourier··on The AI bubble isn't like the internet bubble
You'd be surprised, people are somehow buying Tesla P40s and M40s on eBay for almost $300 and $180 respectively (M40 being the same gen as GTX 950). Google Colab still offers T4s and it's taken them years to add modern GPUs. Hope they're powering them with renewables at least.

And people in general are holding on to their old machines for very long periods of time now, especially CPUs. I've had to support first gen Intel i7s at work! That's pre AVX.

joefourier··on The AI bubble isn't like the internet bubble
Outside of training the biggest LLMs at big labs, GPU lifespan isn't as short as the OP made it out to sound. A100s are 6 years old and still a reliable work-horse, and the 80GB version hasn't depreciated that much on the used market. On the consumer side, 3090s are actually still selling for very close to 2020 MSRP.

Even the ancient V100 (soon to be 10 years old!) had somewhat of resurgence on the second-hand market, with a healthy market for interconnects in China.

If I had a datacenter and power consumption was not a concern, I'd be holding on to my A100s for years at least for inference.

joefourier··on Perceptual Image Codec: What Matters in Practical Learned Image Compression
> And why is V100 even used? V100 is four generations old and not even supported anymore.

It wouldn’t surprise me that due to bureaucratic processes, it’s still somehow the most readily available GPU for Apple researchers despite being almost 10 years old now. I recall even last year seeing V100s used by Microsoft researchers who weren’t working on LLMs.

joefourier··on Was my $48K GPU server worth it?
Not a single new 64GB GPU, but multiple used GPUs.

They’ve significantly increased in price (so much for hardware depreciation…) but you can still get a modded 22GB 2080 ti for $320, or a Mi50 32GB for ~$450 each (used to be $150 a few months ago, alas), or a Mi50 16GB or <$200 but you’d need to stack 4 of them.

There’s also some more exotic configurations but those are probably the simplest options. You won’t get the performance of an RTX Pro 6000 Blackwell of course, and the power consumption will be pretty high so it’s only worth it if you have cheap electricity. But it is possible.

joefourier··on Was my $48K GPU server worth it?
What quant? You should have no problem running it at Q4 with 256K context, Q5 or Q6 even although maybe not at full context. I can run Q4 on a 4090 with just 24GB VRAM.
joefourier··on Was my $48K GPU server worth it?
Who is going to buy a $4299 M5 Max MBP with 64GB of RAM just to run Gemma 4 31b? Firstly you don't need 64GB for that model. Secondly if you want a machine that sits in the corner and does nothing but LLM inference, you don't buy a MacBook Pro, you buy some GPUs which are going to cost you a fraction of that (~$1k for ~64GB of VRAM is possible). The people buying Apple Silicon for inference general aim for the Mac Studios with enormous amounts of RAM (128-512GB), to run very large models.

The idea is obviously to be running the LLM on your work laptop. As a developer I'd need a laptop with 24GB of RAM for work anyway, and 48GB, which is enough for a very good quant of Gemini, is just $400 extra.

joefourier··on Was my $48K GPU server worth it?
Why didn't you take into account batching, input tokens, different costs of electricity, and the fact that a laptop can still hold a decent % of its resale value, and is useful for many other tasks than running an LLM?
joefourier··on Access to frontier AI will soon be limited by economic and security constraints
> I'd compare it to OpenAI 5 years ago except I think even then OpenAI had way more!

Say what? 5 years ago OpenAI had received around $139 million in funding, and they’d just come out with GPT3 with 175B parameters, a 2048 context window, trained on 300B tokens on a 10,000 V100 cluster which would have cost maybe $4-13 million at the time for their training run.

Meanwhile Deepseek V3’s famously frugal training was $5M, and Chinese AI companies are raising billions in funding. Sure American AI companies are raising tens (and maybe hundreds in the case of OpenAI, if you count their circular funding rounds) of billions but they’re grossly inefficient, and we’ve already hit the limits of the scaling laws where there’s little point in increasing the number of parameters of a model.

joefourier··on Princeton mandates proctoring for in-person exams, upending 133 year precedent
It's incredibly common all over Europe, not just Switzerland. Not only the metros but the trams and even buses often rely on this system where there's no turnstile or barrier, you just walk in.

Not sure it's about being a high trust society or not, there's frequent inspections where they block the doors, and you get a hefty fine if you're caught without a valid ticket. I certainly wouldn't call Prague or Rome or Dublin high trust societies on par with a Swiss city.

joefourier··on I have seen the dystopian future of elderly care
Personally I feel like it would be less undignified and infantilising to have a machine take care of my basic bodily functions than a human being. There's no feeling of judgement or being shamed in front of someone else, and the machine could even restore a feeling of autonomy since it would feel like you're using a tool instead of being helplessly reliant on another person's help.
joefourier··on I returned to AWS and was reminded why I left
> I also used dedicated servers in the late ’90s (and they still offer great value today). But before AWS, provisioning new hardware typically took days, not minutes.

VPSes and non-custom configs for dedicated servers were pretty instant as far as I know, I think the advantage of AWS was more that you could scale up and down much more easily since you weren’t locked down in a monthly contract, and that you could automate server provisioning through an API.

Page 1 of 12Next →