Kurtzerkag has a good video about it: https://youtu.be/TYhNHX372ek?si=5ESz3ykXykyUcmYg
15,328 karma · joined August 29, 2017
Unless stated otherwise, opinions here are my own and personal.
Kurtzerkag has a good video about it: https://youtu.be/TYhNHX372ek?si=5ESz3ykXykyUcmYg
And, is it really even an oligopoly anymore? Open weight models are incredibly competitive in every way; whether you want to use US providers, Chinese official providers, self host, etc.
As he confirmed, for at least a period of time, and for some users, “Sol xhigh” was actually “Sol high”, etc.
I’m sure it’s lowkey intentional, probably encourages users to spin up new chats; hence less context.
If you’ve tried setting either of these up, you’ll know how various tricky settings can impact throughout and model output quality; and those are much simpler stacks.
Even homelabbers are getting into disaggregated compute; e.g. one GPU for prefill, another for decode.
Obviously Anthropic and co are using a mixture of GPUs and hardware and clusters; not everything is just GB300 or whatever; so you then get into hardware quirks, kernel optimisations that may deliver huge speedups at the cost of a tiny bit of KL divergence, etc.
And I believe they’ve publicly said they use TPUs for inference too, but I doubt exclusively; and I’m sure that’s well optimised too.
Finally, Google has publicly stated they intentionally and silently degrade/poison models in response to distillation attacks; who knows what the other companies do.
At their scale, you’d have to be setting money on fire if you’re not doing dynamic inference optimisations based on load.
API and consumer subscriptions are treated differently; all trackers measuring via API won’t notice this.
But even so, I often realise I self-justify not doing something and attribute it to my ADHD, excessively. It becomes a clutch, and I gravitate away from the willpower I do have.
That reminds me, it’s time to reach out to some friends and make weekend plans, even if I’d prefer having a quiet one. I’ve had too many recently :)
I don’t enjoy going to the gym, but I love swimming and hiking, and love that. I haven’t worked out in a gym for years, I don’t want to, and I’m quite healthy and reasonably active.
Some people are just introverts, and I think the ‘optimal’ amount of IRL time is different for every individual. Much like my gym vs swimming example; quality social interaction can take many forms; some of my closest and most meaningful friendships over decades have been based on games (voice calls; discord chats; hanging out in person sometimes and going on trips together sometimes), and that works for me; and I don’t think I’m missing anything.
Prioritise genuine human and social interaction, and for introverts, get out of the comfort zone once in a while (being a hermit is not good). But whatever form that takes is secondary to the quality and how it fits your personality and quirks.
After claude or codex finishes a commit, I switch console tabs and ask the other to review it; with the commit ID. I often find it helpful to inject a bit of human knowledge, and callout any areas of attention I see from a quick skim. (But that could just be me wanting to not abstract myself away from software engineering that much :)
Sometimes, I do ask Codex to read my ~/.claude/; and vice-versa. But, generally, I try to keep as much knowledge (e.g. investigations, reports, deep dives) inside the git tree as possible; so that is not necessary.
I don't use skills, but I do have ~/.codex/AGENTS.md and ~/.claude/CLAUDE.md. These are high-level instructions for what (1) I consider readable, maintainable code and patterns, and (2) workarounds for empirically observed model behavior IOdon't like, such as Astra being a bit of a "over-correct over-validation nit-picker". I keep these human-authored, and update regularly based on what I find annoying.
I do all of this before I submit a PR; but of course, for trivial stuff (e.g. CSS changes, copy/string changes, etc), I don't bother.
My practices and workflows do change over time. Back in the ~Opus 4.5 days I'd often define a rubric/criteria in a markdown file, iterate with AI to improve it, and that's the "spec". I've stopped doing that since GPT 5.6; partly because models have gotten a lot better at understanding high level intent from the context; and partly because nearly all models these days feel 'gradermaxxed' when working like that.
Finally, consumer $100/$200mo subs get you _so_ far, I get a lot of value from both. I used to have multiple Claude subs for a while, but trying to 'get full value' made me work on projects just for the sake of it; so 2x$200/mo is my cap :)
Other than trivial PRs; everything I do with Opus/Fable gets reviewed by Astra; and everything I do with Astra gets reviewed by Opus/Fable.
Using only models from a single vendor, is like testing your website/webapp only on Chrome.
There’s also the fact that LLMs aren’t perfect, and sometimes even the best models act really stupid sometimes.
I fully believe GP that there was zero intent to gate this behind collecting telemetry. Sounds like a little tech debt and a little oversight, and the simplest explanation is that it is.
Fable and Opus, since 5.1 and 5, will happily hill climb on my CUDA kernels for transformers.
The other explanation is just as part of ‘token efficiency’
It’s more of something to scare companies with legal teams. If you’re an individual or hobbyist doing a side project the risk is essentially zero.