In my experience if you're using OpenAI/Claude models and paying API costs, almost every other harness beats Claude Code/Codex in cost.
181 karma · joined May 7, 2026
In my experience if you're using OpenAI/Claude models and paying API costs, almost every other harness beats Claude Code/Codex in cost.
Also, advocating for slowing LLM progress does not benefit Anthropic or OpenAI.
This was a fringe belief until recently, but the progress of AI in research is impossible to ignore. Epecially in math, where not only has AI outstripped humans in generative ability, but is able to create scientific knowledge which is beyond the capacity of human comprehension.
There's clearly no intelligence task that AIs can't do due to some magic fundamental constraint. And it's hard to imagine a world where current limitations like poor sample efficiency or lack of continual learning won't eventually be solved.
Total AI compute is estimated to grow somewhere in the 1-10 million-fold range in the next decade. Please don't underestimate the phase change that's still coming.
Sure, maybe there's some plateau due to RL being fundamentally limited in some surprising way, but this is nothing but a hope.
Completely discredits the index if it just gets modified to match social media vibes.
I wonder if you can use these for attacks, like this previous paper showing that if you know how a model reasons, you can "fake its thinking" to control it? https://news.ycombinator.com/item?id=48631888
But this has only been shown on simple tasks, so I think this paper is still quite neat. The interesting thing is that they show "future horizon length" varies across models.
They're probably rapidly opening + closing new jobs to increase visibility, as matching models on job boards tend to prioritize new posts.
They also target a cost-insensitive market (corporate/coding users) compared to Google/OpenAI which support massive amounts of free users.
GLM-5.2 actually has really good intent understanding though, on par with GPT-5.5 and Opus from my experience.
I don't believe this would work on two LLMs that have different pretraining. Even if it did you would need two LLMs that have exact same internal activation shapes, dimensions, expert counts, token vocabulary, realistically it would never happen outside of finetunes or academic experiments.
I find this rather disturbing. Anthropic has quite a habit of overclaiming on questionable research results when they definitely know better. For example, their linked circuits blogpost ("The Biology of LLMs") was released after these methods were known to have major credibility issues in the field (e.g., see this from Deepmind - https://www.lesswrong.com/posts/4uXCAJNuPKtKBsi28/negative-r...). Similarly this new blog is heavily based on another academic paper (LatentQA) and the correlation/causation issue is already known.
Shoddy methodology is whatever, but it feels like this is always been done intentionally with the goal of trying to humanize LLMs or overhype their similarities to biological entities. What is the agenda here?