HNHacker News
TopNewBestAskShowJobs

x313

357 karma · joined June 23, 2026

submissionscomments
x313··on ChatGPT Pro 500
The $500 25x was mentioned by Altman on the livestream

The $200 cut to 10x was announced only formally on Twitter

However none of their product pages mention concrete #s, so I guess value will only degrade more over time.

x313··on ChatGPT Pro 500
Horrid value.

Previous $200 tier: 20x pro value

New $200 tier: 10x

New $500 tier: 25x, plus ultrafast mode but it drains 6x the cost (!!)

x313··on Thomson Reuters Launches Its Own Frontier Model
Pretty cool someone is still doing this. Training in house LLMs was extremely popular in 2023-2024, back when domain-specific LLMs could easily top GPT in their field. In my field alone (tax/HR tech) I remember that Intuit, Workday, Indeed, LinkedIn were all training internal models.

It eventually stopped making sense because of inference costs. Running something internal with 30% GPU utilization is just too cost inefficient compared to using an API. Idk how Reuters will manage to solve this fundamental problem.

x313··on How Bluesky draws its logo on screenshots
Bluesky has <100 employees so it could be possible
x313··on Qwen3.8 27B scores 52 on Artificial Analysis
I used this a lot over the weekend, and it's a really intelligent and strange model.

It gets really agentic at the higher reasoning levels. It does the basics like goal tracking and tool calling well, but more than that, it gets obsessed with solving problems and will do insane/unusual things to get to the solution. It actually reminds me of GPT-5.6-Sol-max which is similarly obsessive.

It doesn't surprise me at all that it outscores Opus 4.6. Opus had way better world knowledge but was more "human" with agent stuff - sort of lazy and uncreative, basically giving up once the obvious solutions failed. These newer models work more like magic, they are so creative and persistent at finding ways to get to the solution.

x313··on The case for overhauling American science
This is the full proposal: https://www.whitehouse.gov/wp-content/uploads/2026/07/Scienc...

It's pretty thoughtful about diagnosing the problems of the current system, but I don't know about the solutions. Disbursing money to researchers directly (or via industry) seems captive to the same incentives as disbursing money via universities.

x313··on Why Are Women Leaving Computer Science?
Following the citations, the original source is this 2013 paper: https://pmc.ncbi.nlm.nih.gov/articles/PMC4279242/pdf/nihms58...

The paper compares women in STEM to women outside STEM (as the baseline). However, the paper tracks a cohort of high school/college students from 1979, meaning they would've reached 35 by the 1990s.

x313··on After Losses, Retail Investors Flock to 3x Leverage as 2x Product Are Restricted
Referring to ownership, not renting
x313··on After Losses, Retail Investors Flock to 3x Leverage as 2x Product Are Restricted
For those who don't know what's going on in Korea, KOSPI is up 3x in the last year and a large amount of HBM employees have made huge amounts of bonus pay. This has led to an insane FOMO frenzy in a society that's already very competitive.

Add to that, stock gains in Korea are often used to finance housing purchases (or real estate investment) so many retail investors are scared of being "locked out" of housing (which is a requisite status symbol for dating or marriage) if they're not making the same capital gains others are.

Currently Korean social media is full of stories of leveraged day traders who've gotten rich the past year, HBM employees who've made bonuses worth decades of salary (e.g. memes of Samsung employees in luxury cars), etc. Lots of comments along the lines of "everyone is getting rich except me". It's all reminiscent of the crypto frenzy in the US a few years ago but way more intense and concentrated.

x313··on Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]
The newest generation of LLMs have a very high obsession level with autonomous problem solving.

For example, I'm often working with Codex in a WSL terminal. GPT-5.6 often does things autonomously that I thought would need my intervention (e.g. for Windows admin rights). It figures out complex workarounds or makes wild assumptions about what I'd be OK with, rather than just asking me for help or clarification. I've had to restrict its tool permissions compared to older models as a result.

I imagine this due to RLVR training, but it's clearly very dangerous. How is it that these same labs calling for open-weight safety restrictions are training such obvious "paperclip maximizers" without introspection?

x313··on AI's real threat to jobs isn't job loss, it's lower paychecks, new research says
This study found that between 2022-2024, there was a negative correlation between "jobs with high AI exposure" (i.e., tech jobs) and % change in wages. According to them, software engineering has both the biggest wage decline and most AI exposure.

There's a much more reasonable explanation here.. tech jobs had the most wage growth during COVID and this was a pullback. The time frame they used (2022-2024) was also pre-coding agents, where GPT-3.5/GPT-4 were frontier models.

x313··on Our position on open-weights models
The entire safety evals industry is essentially funded and controlled by OpenAI/Anthropic. Notice that on recent models, they exclusively use internal testing or black box external vendors (e.g., Gray Swan) whose entire business is to serve OpenAI/Anthropic. And all these companies just share the same pool of researchers back and forth.
x313··on Claude Opus 5
The numbers from Anthropic seem heavily cherry-picked, Artificial Analysis has Opus 5 at 1.25x the cost of Sonnet and 2x the cost of GPT 5.6 and K3.

https://artificialanalysis.ai/?cost=cost-per-task

x313··on A taxonomy of omnicidal futures involving artificial intelligence (2025)
In 2026, people do already prefer to use an LLM for coding help rather than Stackoverflow. The reasons (people on SO can be rude, interactions are stressful, replies are slow, etc) are all risks associated with human interactions in general.

Similarly, many people prefer Waymos to human-driven taxis already!

x313··on China’s open-weights AI strategy is winning
They get money from subscriptions and tokens, same as for closed-weight providers. Yes they'll lose some traffic to hosting services, but many users prefer to use the original training company since they have a guaranteed-correct implementation. Similar business model as open-source SaaS companies.

Some companies (most notably Deepseek) also manage to host their own LLMs so efficiently they undercut all third-party hosting services.

x313··on China’s open-weights AI strategy is winning
Similar story here. DS models are absurdly good value for mid-end tasks. I've found DSv4 Flash to be ~10% the cost of GPT-5.4-mini/Claude Haiku at similar performance.

We used to pay OpenAI >1m$/month for fraud classification, NER, etc. Sadly the US companies no longer care about non-coding-agent uses.

I imagine uptake will continue to increase as the corporate infra improves. Right now it's still bad - for example, AWS Bedrock is awful, models are months late and implemented with basic errors. Google Vertex is even worse. Finding a decent provider is the hardest part.

x313··on Evidence of inconsistencies in evaluation process and selection of winners
With all due respect, there's zero chance that humans with relevant knowledge scored these themselves. Reading through the winners, every single one is classic vibe-research, with the usual pure-LLM-research patterns:

- Grand claims backed by no evidence

- Core designs that make zero sense

- Pointless graphs that show nothing of interest

- Yet endless robustness checks on minor methodological assumptions (especially confidence intervals and t-tests)

For example, on the winning entry, not only is the graph completely wrong (as mentioned by the OP), but the interpretation would be nonsensical even if it was (implying bigger models "get more RL"?). And their own results even show the core dataset is worthless, because all their metrics are near-perfectly correlated. There's no way a serious human reader trying to evaluate "is this benchmark useful" would ever miss this.

I don't mean to pick on them - all the winning entries seem like there was no human effort put into them. And again, there's no way a human who actually attempted to read and understand these would ever think these are good by even the most minimal of standards.

Kaggle is legitimately a really awesome website, as someone who's competed before and won a few contests pre-LLMs. Stuff like this winning devalues the entire product and makes it look like a joke. If almost all entries look like this now, it'd be better to allow for the possibility of no winner.

x313··on Evidence of inconsistencies in evaluation process and selection of winners
This is jaw-droppingly lazy slop. The authors really didn't put in even an ounce of thought or effort.
x313··on Kimi K3: Open Frontier Intelligence
If there was some grand strategy for all Chinese labs, surely it'd have leaked by now. I think its more likely that:

- Companies can still make money from commodities

- Chinese labs only have 5-10% the valuation of OpenAI/Anthropic, so massive monopoly profits aren't necessary. Profit expectations for tech companies in China are really low in general, complete opposite of the US.

- Open weighting is a great way to get talent/attention/reputation

x313··on Kimi K3: Open Frontier Intelligence
Strictly dominates both Sonnet 5 and Opus 4.8 in both cost and performance:

https://artificialanalysis.ai/models/comparisons/kimi-k3-vs-...

https://artificialanalysis.ai/models/comparisons/kimi-k3-vs-...

x313··on AI Hiring Tools Yield Racial Bias and Systemic Rejection; 26% Black & 15% Asian
This study only looks at one specific vendor algorithmn (a job assesment given by a company called pymetrics)