So now he's a Codex user. OpenAI and Google both have a minimum age of 13.
EDIT: I should note that Anthropic gave him a refund for the whole month that was underway, despite him being nearing the end of it. So good on them.
It is frequently said that programming directly is obsolete, and the skill you must have now is knowing how to operate agentic AIs.
Yet you aren't allowed to do this until you're 18.
So, developing software is now 18+ only?
Who says this?
Local models are chasing the online frontier models pretty hard.
So worst case, that's the fallback (FWIW, YMMV)
edit: Qwen-3.5 MoE (and other local MoE models like it)
If you ever want to answer this type of question yourself, you can look at the size of the model files. Loading a model usually uses an amount of RAM around the size it occupies on disk, plus a few gigabytes for the context window.
Qwen3.5-122B-A10B is 120GB. Quantized to 4 bits it is ~70GB. You can run a 70GB model in 80GB of VRAM or 128GB of unified normal RAM.
Systems with that capability cost a small number of thousand USD to purchase new.
If you are willing to sacrifice some performance, you can take advantage of the model being a mixture-of-experts and use disk space to get by with less RAM/VRAM, but inference speed will suffer.
You can quantize it of course, but if the idea is "as close to Sonnet as possible," then while quantized models are objectively more efficient they are sacrificing precision for it.
So next step is to up that speed, so we're at 4x $1300, Nvidia 5090s with 32 GB of VRAM each (128 GB), or $5,200 before RAM/CPU/etc. All of this additional cost to increase your tokens/second without lobotomizing the model. This still may not be enough.
I guess my point is: You see this conversation a LOT online. "Qwen3 can be near Sonnet!" but then when asked how, instead of giving you an answer for the true "near Sonnet" model per benchmarks, they suddenly start talking about a substantially inferior Qwen3 model that is cheap to run at home (e.g. 27B/30B quantized down to Q4/Q5).
The local models absolutely DO exist that are "near Sonnet." The hardware to actually run them is the bottleneck, and it is a HUGE financial/practical bottleneck. If you had a $10K all-in budget, it isn't actually insane for this class of model, and the sky really is the limit (again to reduce quantization and or increase tokens/second).
PS - And electricity costs are non-trivial for 4x 3090s or 4x 5090s.
Qwen3.5-35B-A3B is reported to perform slightly better than the model you mentioned.
It runs fine but non-optimal on a single 3090 with even 131072 tokens of context , and due to the hybrid attention architecture, the memory usage and compute scale rather less drastically than ctx^2. I've had friends with smaller cards still getting work out of it. Generation is at around 20 tokens/sec on that 3090 (without doing anything special yet) . You'll need enough DRAM to hold the bits of the model that don't fit. Nothing to write home about, but genuinely usable in a pinch or for tasks that don't need immediate interactivity.
It's the first local model that passes my personal kimbench usability benchmark at least. Just be aware that it is extremely verbose in thinking mode. Seems to be a qwen thing.
(edit: On rechecking my numbers; I now realize I can possibly optimize this a lot better)
- Qwen is near Sonnet 4.5!
- How do I run that?
- [Starts talking about something inferior that isn't near Sonnet 4.5].
It is this strange bait/switch discussion that happens over and over. Least of all because Sonnet has a 200K context window, and most of these ancdotes aren't for anywhere near that context size.
Only way to know if your own criteria are now matched -or not yet- is to test it for yourself with your own benchmark or what have you.
And it does show a promising direction going forward: usable (to some) local models becoming efficient enough to run on consumer hardware.
[1] released mid-2025
[2] take with salt - only tests personal usability
+ Note that some benchmarks do show Qwen3.5-35B-A3B matching Sonnet 4.5 (released later last year); but I treat those with the same skepticism you do , clearly ;)
That's unsurprising, seeing as inference for agentic coding is extremely context- and token-intensive compared to general chat. Especially if you want it to be fast enough for a real-time response, as opposed to just running coding tasks overnight in a batch and checking the results as they arrive. Maybe we should go back to viewing "coding" as a batch task, where you submit a "job" to be queued for the big iron and wait for the results.
Gemma 4 31B Q6: 9tok/s, I'd say it is smarter than GPT-4o, but yeah it's slow. Good for coding.
Gemma 4 26B A4B Q4: 50tok/s. Feels faster than ChatGPT 5.4, but not as smart (as it reasons less). Good for general chatting and research.
This is genuine advice I’ve seen from high profile business types. We’re fucked in the sense our children will be made to be attention whores online.
Can you expand on this? Your teenage son makes more money than you do professionally, by vibe coding video games?
I have zero obligation to detail the work my minor son does to random weird foot-stomping, entitled creeps on HN, and these bizarrely insecure demands by professional failures is...telling.
In this case I assume you're keying off of the other clown who, based upon absolutely nothing (but apparently their own professional failure), is certain it must be "cheating software".
How pathetic. If this is your lot in life, Jesus Christ find a different career. Maybe the trades or something.
Anthropic made the best models by hiring non-technical folks like philosophers to build the best training sets and evaluations. Now, it seems like their philosophers are telling people how they can and can't use their model.
I sense an opportunity for free tokens.
Ideas for prompts that reliably trigger the age check?
And the answer to that question is:
"Hell no! We used the cheapest, shadiest company we could find for that. They'll process and sell all your data. Thank you for continuing to be a valued Anthropic customer!".
* preventing North Korea, China, Russian, Iran and etc. actors from accessing service. They absolutely use workarounds to access AI, e.g. I bet there are companies who are proxy between Anthropic and those countries.
I imagine there will be quite some false positives while identifying those.
[0] https://www.tradingview.com/news/cointelegraph:6192f38e3094b...
On the scale of intelligence budgets this would be in the realm of petty cash.
- Repeated violations of our Usage Policy
- Account creation from an unsupported location
- Terms of Service violations
- Under-18 usage
So identity verification is basically a canary that your account is about to get banned, or is on the chopping block. At that point you're better off abandoning ship rather than handing over your ID.