Yeah I set thinking as my default and never looked back. It’s my daily driver and extended thinking is usually not too slow. The way that the “instant” model trades quality for speed is not worth it and I don’t need the instant gratification. (But I also don’t do entertainment chatting so ymmv.)
Or rather, it’s hard to ask everyone to side-by-side compare both products on their use cases. So the choice really comes down to word-of-mouth even though their use cases may be better served by Codex.
Codex also gives you a lot more usage for $20/mon than Claude, so there’s not also that fear that high or xhigh reasoning will eat up all your quota. It really comes down to whether you want to try to save some time or not. (I default to xhigh because it’s still fast enough for me.)
people have been complaining about this since GPT-4 and have never been able to provide any evidence (even though they have all their old conversations in their chat history). I think it’s simply new model shininess turning into raised expectations after some amount of time.
Word on the street is that Opus is much much larger of a model than GPT-5.4 and that’s why the rate limits on Codex are so much more generous. But I guess you could also just switch to Sonnet or Haiku in Claude Code?
Claude Code is not a platform and you’re not meant to be building on it. Netflix is also not a platform and you shouldn’t be running code (open source or not) to mass download Netflix movies either.
They have so much mindshare right now that they can’t lose, and the number of users that use opencode and would be affected is miniscule—-on the level of complaining about your online bank not supporting Konqueror.
Mainly OpenSCAD is not a BRep modeling tool! It is not on the same level of power as CAD tools with a BRep kernel and this especially shows when you want to do a fillet over an arbitrary edge. Unfortunately these kernels are hard to make and integrate and I only know of two open-source BRep kernels out there: OpenCASCADE (used by FreeCAD and build123) and truck (not sure what the status of it is).
This is 100% crank “science” that has wrapped up a banal finding in big words and LaTeX. The claim is roughly as exciting as, “some computer programs print nothing to stdout.”
The output shown is not “null” or “void”. It is the empty string, which these LLMs are perfectly capable of outputting. Technically, it outputs the stop token, analogous to \0 at the end of a C string.
Simultaneously, if you hire human translators, you are likely to get machine translations. Maybe not often or overtly, but the translation industry has not been healthy for a while.
I think Google has already shown that in the long run, people accept ads and prefer them to paying a subscription fee. If that weren’t true, then YouTube Premium would have double-digit % of youtube users and Kagi Search would be huge.
That may be true, but you can’t compare average GenAI with the best humans because there are many reasons the human output is low quality: budget, timelines, oversights, not having the best artists, etc. Very few games use the best human artists for everything.
Same with programming. The best humans write better code than Codex, but the awful government portals and enterprise apps you’re using today were also written by humans.
Model capability improvements are very uneven. Changes between one model and the next tend to benefit certain areas substantially without moving the needle on others. You see this across all frontier labs’ model releases. Also the version numbering is BS (remember GPT-4.5 followed by GPT-4.1?).
Note that this is not relevant for reasoning models, since they will think about the problem in whatever order it wants to before outputting the answer. Since it can “refer” back to its thinking when outputting the final answer, the output order is less relevant to the correctness. The relative robustness is likely why openai is trying to force reasoning onto everyone.
This article spent a lot of words to say very little. Specifically, it doesn’t really say why working towards AGI doesn’t bring advancements to “practical” applications and why the gazillion AI startups out there won’t either. Instead, we need Trump to step up?
More and more I feel like these policy articles about AI are an endless stream of slop written by people who aren’t familiar with and have never worked on current AI.
That’s an interesting point. It’s not hard to imagine that LLMs are much more intelligent in areas where humans hit architectural limitations. Processing tokens seems to be a struggle for humans (look at how few animals do it overall, too), but since so much of the human brain is dedicated to movement planning, it makes sense that we still have an edge there.
The past few years I’ve been hearing crazy stories of workarounds and scripts to deal with all these new features in Windows. Isn’t that what was preventing people from using Linux? Replacing utilman.exe with cmd.exe is not something a normal user would ever do.
> Does "career development" just mean "more money"?
Big companies means more opportunities to lead bugger project. At a big company, it’s not uncommon to in-house what would’ve been an entire startup’s product. And depending on the environment, you may work on several of those project over the course of a few years. Or if you want to try your hand at leading bigger teams, that’s usually easier to find in a big company.
> Is it still satisfying if that software is bad, or harms many of those people?
There’s nothing inherently good about startups and small companies. The good or bad is case-by-case.
GPT-4 is very different from the latest GPT-4o in tone. Users are not asking for the direct no-fluff GPT-4. They want the GPT-4o that praises you for being brilliant, then claims it will be “brutally honest” before stating some mundane take.
It feels crazy to keep arguing about LLMs being able to do this or that, but not mention the specific model? The post author only mentions the IMO gold-medal model. And your post could be about anything. Am I to believe that the two of you are talking about the same thing? This discussion is not useful if that’s not the case.
The 5 seconds delay is probably due to reasoning. Maybe try setting it to minimal? If your use case isn’t complex maybe reasoning is overkill and gpt-4.1 would suffice.