The $200 cut to 10x was announced only formally on Twitter
However none of their product pages mention concrete #s, so I guess value will only degrade more over time.
357 karma · joined June 23, 2026
The $200 cut to 10x was announced only formally on Twitter
However none of their product pages mention concrete #s, so I guess value will only degrade more over time.
Previous $200 tier: 20x pro value
New $200 tier: 10x
New $500 tier: 25x, plus ultrafast mode but it drains 6x the cost (!!)
It eventually stopped making sense because of inference costs. Running something internal with 30% GPU utilization is just too cost inefficient compared to using an API. Idk how Reuters will manage to solve this fundamental problem.
It gets really agentic at the higher reasoning levels. It does the basics like goal tracking and tool calling well, but more than that, it gets obsessed with solving problems and will do insane/unusual things to get to the solution. It actually reminds me of GPT-5.6-Sol-max which is similarly obsessive.
It doesn't surprise me at all that it outscores Opus 4.6. Opus had way better world knowledge but was more "human" with agent stuff - sort of lazy and uncreative, basically giving up once the obvious solutions failed. These newer models work more like magic, they are so creative and persistent at finding ways to get to the solution.
It's pretty thoughtful about diagnosing the problems of the current system, but I don't know about the solutions. Disbursing money to researchers directly (or via industry) seems captive to the same incentives as disbursing money via universities.
The paper compares women in STEM to women outside STEM (as the baseline). However, the paper tracks a cohort of high school/college students from 1979, meaning they would've reached 35 by the 1990s.
Add to that, stock gains in Korea are often used to finance housing purchases (or real estate investment) so many retail investors are scared of being "locked out" of housing (which is a requisite status symbol for dating or marriage) if they're not making the same capital gains others are.
Currently Korean social media is full of stories of leveraged day traders who've gotten rich the past year, HBM employees who've made bonuses worth decades of salary (e.g. memes of Samsung employees in luxury cars), etc. Lots of comments along the lines of "everyone is getting rich except me". It's all reminiscent of the crypto frenzy in the US a few years ago but way more intense and concentrated.
For example, I'm often working with Codex in a WSL terminal. GPT-5.6 often does things autonomously that I thought would need my intervention (e.g. for Windows admin rights). It figures out complex workarounds or makes wild assumptions about what I'd be OK with, rather than just asking me for help or clarification. I've had to restrict its tool permissions compared to older models as a result.
I imagine this due to RLVR training, but it's clearly very dangerous. How is it that these same labs calling for open-weight safety restrictions are training such obvious "paperclip maximizers" without introspection?
There's a much more reasonable explanation here.. tech jobs had the most wage growth during COVID and this was a pullback. The time frame they used (2022-2024) was also pre-coding agents, where GPT-3.5/GPT-4 were frontier models.
Similarly, many people prefer Waymos to human-driven taxis already!
Some companies (most notably Deepseek) also manage to host their own LLMs so efficiently they undercut all third-party hosting services.
We used to pay OpenAI >1m$/month for fraud classification, NER, etc. Sadly the US companies no longer care about non-coding-agent uses.
I imagine uptake will continue to increase as the corporate infra improves. Right now it's still bad - for example, AWS Bedrock is awful, models are months late and implemented with basic errors. Google Vertex is even worse. Finding a decent provider is the hardest part.
- Grand claims backed by no evidence
- Core designs that make zero sense
- Pointless graphs that show nothing of interest
- Yet endless robustness checks on minor methodological assumptions (especially confidence intervals and t-tests)
For example, on the winning entry, not only is the graph completely wrong (as mentioned by the OP), but the interpretation would be nonsensical even if it was (implying bigger models "get more RL"?). And their own results even show the core dataset is worthless, because all their metrics are near-perfectly correlated. There's no way a serious human reader trying to evaluate "is this benchmark useful" would ever miss this.
I don't mean to pick on them - all the winning entries seem like there was no human effort put into them. And again, there's no way a human who actually attempted to read and understand these would ever think these are good by even the most minimal of standards.
Kaggle is legitimately a really awesome website, as someone who's competed before and won a few contests pre-LLMs. Stuff like this winning devalues the entire product and makes it look like a joke. If almost all entries look like this now, it'd be better to allow for the possibility of no winner.
- Companies can still make money from commodities
- Chinese labs only have 5-10% the valuation of OpenAI/Anthropic, so massive monopoly profits aren't necessary. Profit expectations for tech companies in China are really low in general, complete opposite of the US.
- Open weighting is a great way to get talent/attention/reputation
https://artificialanalysis.ai/models/comparisons/kimi-k3-vs-...
https://artificialanalysis.ai/models/comparisons/kimi-k3-vs-...