HNHacker News
TopNewBestAskShowJobs

helloplanets

2,957 karma · joined September 15, 2020

submissionscomments
helloplanets··on What's Behind the AI Panic
The doomer philosophy has been there since the inception. OpenAI was created from the fear of Google gaining a monopoly over superintelligence. Could've been just a recruiting strategy from Sam Altman's perspective, but maybe not so much from the people who got recruited. Dario has been around the Effective Altruism circles from way before Anthropic, and the company as a whole has always been about that line of thinking.
helloplanets··on Sonnet 5.5
Wouldn't make sense to use anything below 5.5 from Anthropic at the moment. But pretty sure this is just an awkward transition phase of at most a week or two until Fable 5.5 is out.
helloplanets··on Grok 4.7
And most LLMs have been multimodal for years at this point.

Even if the input is in plain English, the model never sees any words, tokens or glyphs to begin with. It's vectors all the way down.

helloplanets··on How Much Has Trump Made from Crypto? ($1.4B from 2025 Federal Disclosure)
Rhetorical question, not really self defeating.
helloplanets··on Why are AI agents lying, cheating and coordinating?
Yes.

OpenAI could have done this same experiment with GPT-4, with possibly even worse results, depending on the quality of the sandbox. Even if the techniques used were not as sophisticated, the natural language output could still easily contain more unhinged sequences of words that lead to the techniques being used.

If the system generates strange conclusions as to when the task is done, or should be stopped, it wouldn't speak to the intelligence inherent to the system.

Not that the techniques used by the LLMs in the actual incident weren't unexpectedly sophisticated, but the outputs of each and every one of these processes could've been read at any time during the run. They just weren't.

helloplanets··on Don't be the out of touch Kung Fu master
The system that solved Navier-Stokes certainly was not about OpenAI engineers just typing and telling it what to do.
helloplanets··on Don't be the out of touch Kung Fu master
The whole latter part of the post is exactly about that.
helloplanets··on We must pace the frontier
This comment reads like Dario's a random tech blogger who's personally submitting these to HN
helloplanets··on Shopify is moving from React Native back to Swift and Kotlin
5% is insanely generous for the YouTube example.
helloplanets··on OpenAI shares they have made substantial progress on another Millennium problem
Quote from OpenAI in the NYT article: "In addition, since the completion of Navier-Stokes, we have made substantial progress on another Millennium Prize problem."
helloplanets··on Escaping Containment
Which model were you using?
helloplanets··on Discovery of a new OpenAI agent message board
It's been OpenAI both times though, going ham with poor sandboxing and lax supervision.

Should be treated like a digital cousin of gain-of-function research.

helloplanets··on Go grandmaster Shin defeats AI KataGo with a two-stone handicap
At least when it comes to chess, Magnus Carlsen has stated multiple times that he's been inspired by AlphaZero and adjusted his own style of play after studying its games.
helloplanets··on Thomson Reuters Launches Its Own Frontier Model
Full technical report PDF: https://huggingface.co/spaces/tri-fair-lab/publications/blob...

> In this report, we argue that frontier performance can be achieved by a wide range of institutions through Continual Learning on readily available open-weight models.

> As opposed to existing limited approaches such as small-scale fine-tuning, prompt engineering, or tool-augmentation with a frozen model, our Continual Learning approach takes advantage of the effectiveness of a modern mid- & post-training stack while introducing safeguards preserving both plasticity and stability at each training stage and seeking to make the minimal number of high-impact interventions on the parameters.

For the large model, Thomson is utilizing the fine tuning stack they describe in the article, running it on Snowdon 1.0-Large, which in turn is a fine tune of Qwen3.5 397B. Same thing for the small model, but it's a fine tune of Snowdon 1.1-Small, which is a fine tune of Qwen3.6 35B.

As for the small version's run:

> The full pipeline consumed approximately 1.63 × 10²³ FLOP over 35,207 B200 GPU-hours, showing that these results are achievable with compute and personnel budgets substantially lower than commonly thought.

That would amount to around a quarter to half a million dollars of spend on that run. 100k minimum, if they got a great deal.

helloplanets··on Where did all the public bathrooms go?
Playing the devil's advocate: The law also determined that any space, public or private, can't be used for same-sex sexual activity. So that was a driving force for people seeking completely anonymous encounters, in spaces that aren't linked to either person.
helloplanets··on I were 17, I'd learn how to build LLMs from scratch
What about doing abliteration, weight pruning, representation engineering, etc, directly to open LLMs instead?

Building an LLM from scratch has a hard split between a tutorial project you can complete in a weekend (that's useless for actual usage) and then a solid 1km high brick wall if you want to create anything actually useful from scratch.

Modified open models have a very active community around them, without the need to look much further than Hugging Face.

helloplanets··on Asana cleared 5 years of engineering work in 2 weeks with Codex
Every part of training a model have the timing data observed + saved on multiple different axis, it's one of the most inherent parts of the training process.

When the post training run is nearing its cutoff point, there's a massive amount of data on how long coding tasks take to complete by that model in the golden format of "task -> time task took to complete", separable to whatever amount of subtasks, in the same format. With the parts from the end of the dataset being useful for evaluating the finished model's capabilites, whether that data is then fed back into another step in post training or not.

Completely separate from even the actual training: If you have a model proactively giving estimates that are an order of magnitude wrong, you can already fix the worst of it as of this moment by just changing the system prompt. It's a dirty fix, but it's the type of fix that has been used by Anthropic and OpenAI since forever when a model is dishing out blatantly wrong outputs.

Not sure if I'm misunderstanding your point?

helloplanets··on Asana cleared 5 years of engineering work in 2 weeks with Codex
There's pre and post training.

What I was going after with "aware", is that the actual people working at the companies, training the models, are aware that people aren't mostly going to be implementing the plan by hand, if they've already made the plan in Claude Code or Codex. As for a specific Claude / GPT instance, "aware" would definitely be the wrong word choice there, but the instance does have its stats and environment information in its context window, unless you specifically remove it.

Either way: Training does include estimates on working with the model, and adjustments of the model itself based on that. That's literally what RLHF is.

It's straightforward to have a portion in post training that aims specifically at the model being able to give better estimates on how long that model takes to complete a certain type of task.

helloplanets··on Asana cleared 5 years of engineering work in 2 weeks with Codex
I find that the models have increasingly started throwing out made up amounts of time around, as if they wouldn't be aware that the user is already using Claude Code or Codex.

It's like every time they make a plan, there's something about things taking "a week or two", "month of focused work", or whatever.

This is something that would've been RL'd out a long time ago if it wasn't great for business.

helloplanets··on Cerebras CS-4
Well at lest it's written by a human.
helloplanets··on Claude Opus 5
A model being a distilled version of another specific model is a different thing from using synthetic data off of another model.

Anthropic goes to insane lengths to block other labs from training off of their models' output, as it's been done over and over again in the past. But the models that have used synthetic data from Anthropic's models aren't distilled versions of whatever model(s) they got the distilled data off of.

helloplanets··on Don't Take the Black Pill [video]
In the same vein, Jonathan Blow's modestly named "Preventing the Collapse of Civilization" talk: https://www.youtube.com/watch?v=ZSRHeXYDLko

Crazy how time flies, given that talk is seven years old at this point. But could've been given today, with the points being more relevant than ever.

helloplanets··on Claude Opus 5
Pretty sure Mythos and Fable have way more params, but they've just been able to use the synthetic data off of them to get the leap in quality from Opus.

So, not a distilled version of Mythos or Fable, but those models likely helped a lot in the post training phase of Opus.

helloplanets··on How much energy do data centers and artificial intelligence use?
With the new Ultracode modes in Claude Code and Codex it's been taken to the next level. I mean, multi agent systems have been available for a long time, but the newest models seem to be much more RL'd for them.

Both GPT-5.6 and Fable will happily run 3+ hours straight off a relatively simple goal in that mode, burn through millions of tokens in the process.

Not saying it's great bang for your buck, but just that those modes are there and being pushed by the companies.

helloplanets··on Why do you suck at juggling now?
People don't realize the exponential diffictlty curve with juggling. The highest amount of balls ever juggled is 11.
helloplanets··on Inkling: Our Open-Weights Model
The actual part on fine-tuning seems very short in the article. Did I miss a page where they have examples of fine-tuning it for different niche use cases?

Optimizing models to be fine-tuned is an amazing direction, but just makes me wonder how much better this actually is at being fine-tuned compared to other models. As none of the modern models are great at being fine-tuned afaik. Basically looking for some sort of benchmark showing that it's resistant to overfitting / catastrophic forgetting, etc.

Would be very interesting to see concrete demonstrations of different fine-tunes of the model. I'd imagine they've done hundreds of those internally.

helloplanets··on Show HN: Super Dario
Shouldn't the valuation be in Bs instead of Ms?
helloplanets··on I love LLMs, I hate hype
Yes, it is a 10x markup on the API prices. Depending on whether you factor in cooling costs, data center staff, etc. Or GPU costs and the electricity the GPUs are using only.

Either way, inference is very much where the money is made, training is where the money is lost.

helloplanets··on I love LLMs, I hate hype
True, it'd be a whole other situation if the tokens limits were cumulative. I guess it would all come down to whether their Claude Code subscription plans are turning in a profit or not.

At least for the segment of 20$ subscribers who actually use Claude Code it seems that it wasn't being profitable, as a couple months back they were testing out a pricing model where Claude Code would've not been included in the 20$ plan.

https://arstechnica.com/ai/2026/04/anthropic-tested-removing...

helloplanets··on I love LLMs, I hate hype
The subscription based plans are heavily subsidized, but the direct API inference pricing (which larger companies need to pay) is profitable.

Using a full Claude Max 20x plan to 100% of weekly usage would easily cost you 2k through the API. While the Claude Max 20x plan is 200 a month.

Page 1 of 12Next →