Why care for current iterations of nature, when we will all get to experience infinite varities of it as immortal digital consciousnesses?
313 karma · joined November 18, 2023
Why care for current iterations of nature, when we will all get to experience infinite varities of it as immortal digital consciousnesses?
Otherwise we agree that benchmarking is hard, the benchmarks contain hard problems, and that there are many hard working people trying to accurately gauge what is going on. It is getting harder to watch though as all that is on the line taints the overall endeavor.
Either way, I agree that HN is quickly becoming more manipulated and low SNR, like the rest of the entire internet.
Maybe back when this was a scientific endeavor; not now when enormous, enormous amounts of capital are on the line. Along with an entire cult's chosen eschatology.
Honestly I'm pretty tired of Anthropic's press releases too, but this one is pretty benign. If I was a hater, I'd save up my new-account-energy for their next "paper" that insinuates Claude might be actively introspecting.
HN is way too central for shared sentiment in the tech world for these companies not to do some amount of astroturfing. AI companies have shown at every single turn that they act out of self-interest and greed, not of moral principles. So it isn't surprising, even if it is still sad, to see those who are commanding the most capital in human history act with such callousness.
I think the appropriate course of response is to stop adding to public spaces on the internet. No doubt painful for those of us who have so benefitted from the freely shared thoughts of others. But if well-funded bullies are going come in, steal everything, ruin the commons, and then say "this is the new normal, deal with it", there isn't much the rest of us can do other than stop feeding them.
They pay people for expert training data they do not share because it gives them an edge over other AI companies. And, as always, deep learning is enormously data-hungry, and we've gotten to the point where publicly available data has been exhausted.
AI companies absolutely retrain models regularly to keep up with the cutting edge. There is a reason why this announcement references an internal, unreleased model, rather than "we just put a lot of new math papers within the GPT5.5 context window and found this."
- https://www.theverge.com/cs/features/831818/ai-mercor-handshake-scale-surge-staffing-companies
- https://outlier.ai/math/en-us
- https://www.opentrain.ai/
- https://www.pin.com/blog/ai-labs-hiring-train-models/
Much of this is data annotation, reasoning trace evaluation, and problem set curation. But there is no way they haven't atleast paid some mathematicians to work on research grade problems in tandem with their models, and then used that for training data.Does this expert data likely contain this proof within it? No. Would it temper the impressiveness to know they have a large amount of novel mathematical training data, an internal Lean harness for evaluation of open conjectures, and spent hundreds of millions in compute to calculate this? Yes.
I'll gladly admit I think what these companies are doing is unethical, and I'm sure that biases my thinking toward skepticism.
That said, there remains way too much that is hidden to be able to effectively evaluate what is going on. You have the perfect storm:
- AI companies do not share their custom internal harnesses.
- AI companies do not share their custom internal training data.
- AI companies do not share how much compute they allocate to trying to solve problems of this nature.
- AI companies are primarily marketing their models to investors as human-replacing rather than human-augmenting.
- AI companies are under enormous financial pressure to make their business work.
The last two points incentivize them to find these types of "first proof" successes as aggressively as they can, and I'm sure they've thrown the whole book at it.Is it likely that they literally had a mathematician discover this, put it into the training data, and then prompted it out? Of course not.
But it would make a world of difference--in evaluating the impressiveness of this discovery and LLM capabilities in general--if we were to know the extent to which the training data crosses over this problem, the harness with which this was ran, and how much compute was spent.
Until they bring more transparency to the whole process--something which some of the mathematicians commenting on this even asked for--I will personally take discoveries of this nature with a good dose of salt.
In all seriousness though: My suggestion is that those shepherding the frontier of AI start acting with more transparency, and stop acting in ways that encourage conspiratorial thinking. Especially if the technology is as powerful as they market it as.
Without knowing all this model has been trained on though, it is pretty hard to ascertain the extent to which it arrived to this "on its own". The entire AI industry has been (not so secretly) paying a lot of experts in many fields to generate large amounts of novel training data. Novel training data that isn't found anywhere else--they hoard it--and which could actually contain original ideas.
It isn't likely that someone solved this and then just put it in the training data, although I honestly wouldn't put that past OpenAI. More interesting though is the extent to which they've generated training data that may have touched on most or all of the "original" tenets found in this proof.
We can't know, of course. But until these things are built in a non-clandestine manner, this question will always remain.
From TrendForce's analysis:
"The laptop market's 2026 shipments have been revised down from the previously expected annual growth of 1.7% to -2.8%, and further adjusted to -5.4%. Brands with highly integrated supply chains and more flexible pricing, such as Apple and Lenovo, have more flexibility to handle rising memory prices. However, low-end and consumer laptop brands face difficulty passing on costs and are constrained by processor and operating system requirements, making further spec reductions difficult."
Google can obviously just make this machine more expensive, but to launch a completely new brand of consumer laptops in a year where production is already very constrained is only going to exacerbate the core issue.
Checking now: The way they describe it in their FAQ is that if the price changes, then they will bill you the new price. But I read that as regarding if the primary model provider changes their headline token cost; not in the case of pricing differences for models that have many different backends that host them.
Regardless, I would be more concerned about the streaming costs if the service continues to blow up and they scale aggressively through VC investments. If their 5.5% skim accounted for what they needed, you'd think they could effectively grow organically..
They also show headline prices for the cheapest provider of whatever model, but then need to hit different backends some of which may be more expensive. For now they absorb those costs, but the VCs always come knocking.
Just my opinion though. Totally agreed that they have one of the best positions amongst all AI providers from a financial standpoint.
But, if not: It is different because Drew DeVault is scathingly anti-AI, and has a history of sticking to strong opinions (for better or worse). Seems like the best bet for off-premise source control if you are concerned about AI scraping and downtime.
There is definitely enough empirical validation that shows image models retain lots of original copies in their weights, despite how much AI boosters think otherwise. That said, it is often images that end up in the training set many times, and I would think it strange for this image to do that.
Regardless, great find.
But they also desperately need users (and the data those users bring) to build their products, and the people who do have the power to manipulate this site are on their team. And it does get tiring to see a new Claude feature with like 1 comment and 25 points right at the top, multiple times in the last two week. Keeping their needs in mind, it has begun to look like manipulation, even if the above effect could explain it.
I'm glad the technology foments it excitement for you. The idea that we can share intellectual processes broadly and implement them without the previously requisite skills will obviously change the world. That it could change the world for the better, excites me too.
But many of us have our excitement tampered by the messaging, the questionable ethics behind how it has been done, and the fact that a real % of the space is basically driven by eschatological thinking. And it especially annoys me that Anthropic is the company whose messaging simultaneously encourages that eschatological thinking, and preys upon the emotional reactions it creates.
I think it is increasingly clear--if you look at recent public sentiment and feel what is in the air--that they are a villain in this aspect. I don't think we want the people who believe they are building the future to be doing so both out of fear--of China--and gaining power through others' fear of what they are doing.
But villains can ultimately do good in the world, despite their villainy. Let's hope that is how it plays out.
But, I'll gladly admit that I am bias: I'm tired of seeing blatant astroturfing by a company whose main marketing tactic is to play on societal fear, while simultaneously employing safety theatre to look like the "good guys".
So take my opinion with a grain of salt :)
Until we have embodied AI's with eyes and hands that provide good enough approximations, the aspect of design bottlenecked on human experience will stay bottlenecked.
Combine that with the obvious hackernews manipulation that somehow gets each and every haphazard release instantly to the top, and you can see they're starting to feel some real heat.
Think about how valuable HN is for a company whose primary market is professional devs.
Can you link any? All I've seen is stuff like Anthropic claiming 90% of internal code is written by Claude--I think we'd agree that we need an unbiased source and better metrics than "code written". My concern is that whenever AI usage in professional developers is studied empirically, as far as I have seen, the results never corroborate your claim: "Any developer (who was a developer before March 2023) that is actively using these tools and understands the nuances of how to search the vector space (prompt) is being sped up substantially."
I'm open to it being possible, but as someone who was a developer before March 2023 and is surrounded by many professionals who were also so, our results are more lukewarm than what I see boosters claim. It speeds up certain types of work, but not everything in a manner that adds up to all work "sped up substantially".
I need to see data, and all the data I've seen goes the other way. Did you see the recent Substack looking at public Github data showing no increase in the trend of PRs all the way up to August 2025? All the hard data I've seen is much, much more middling than what people who have something to sell AI-wise are claiming.
https://mikelovesrobots.substack.com/p/wheres-the-shovelware...