686 karma · joined July 14, 2024
But even when Opus is running healthy, it still doesn't address the underlying issue that these models can only do so much. I have had Opus build out a bunch of apps but I'm still finding my time absorbed as soon as it comes to anything genuinely exceeding "CRUD level difficulty". Ask it to fix a subtle visual alignment issue, make a small change to a completely novel algorithm, or just fix a tiny bug without having to watch for "Oh, this means I should rewrite module <X>" is something that simply isn't possible while still being able to stand over the work.
It's not to say I don't get a massive benefit from these tools, I just think it's possible to be asking too much of them, and that's maybe the real problem to solve.
Would really love some path forward where the AI parts only poke out as single fields in traditional user interfaces and we can forget this whole episode
Dealing with Claude going into stupid mode 15 times a day, constant HTTP errors, etc. just isn't really worth it for all it does. I can't see myself justifying $200/mo. on any replacement tool either, the output just doesn't warrant it.
I think we all jumped on the AI mothership with our eyes closed and it's time to dial some nuance back into things. Most of the time I'm just using Opus as a bulk code autocomplete that really doesn't take much smarts comparatively speaking. But when I do lean on it for actual fiddly bug fixing or ideation, I'm regularly left disappointed and working by hand anyway. I'd prefer to set my expectations (and willingness to pay) a little lower just to get a consistent slightly dumb agent rather than an overpriced one that continually lets me down. I don't think that's a problem fixed by trying to swap in another heavily marketed cure-all like Gemini or Codex, it's solved by adjusting expectations.
In terms of pricing, $200 buys an absolute ton of GLM or Minimax, so much that I'd doubt my own usage is going to get anywhere close to $200 going by ccusage output. Minimax generating a single output stream at its max throughput 24/7 only comes to about $90/mo.
Planet announced last week there will be a 14 day delay on all commercial satellite imagery from the middle east. It shocks me how transparent we are about information war and voluntarily lying to ourselves at particular moments
Even if utilisation weren't a metric, "efficient" can be interpreted in so many ways as to be pointless to try and apply in the general case. I consider any model I can foist into a Lambda function "efficient" because of secondary concerns you simply cannot meaningfully address with GPU hardware at present (elasticity and manageability for example). That it burns more energy per unit output is almost meaningless to consider for any kind of workload where Lambda would be applicable.
It's the same for any edge-deployed software where "does it run on CPU?" translates to "does the general purpose user have a snowball's chance in hell of running it?", having to depend on 4GB of CUDA libraries to run a utility fundamentally changes the nature and applicability of any piece of software
A few years ago we had smaller cuts of Whisper running at something like 0.5x realtime on CPU, people struggled along anyway. Now we have Nvidia's speech model family comfortably exceeding 2x real time on older processors with far improved word error rate. Which would you prefer to deploy to an edge device? Which improves the total number of addressable users? Turns out we never needed GPUs for this problem in in the first place, the model architecture mattered all along, as did the question, "does it run on CPU?".
It's not even clear cut when discussing raw achievable performance. With a CPU-friendly speech model living in a Lambda, no GPU configuration will come close to the achievable peak throughput for the same level of investment. Got a year-long audio recording to process once a year? Slice it up and Lambda will happily chew through it at 500 or 1000x real time
edit: just checked the version that ships with Steam on Linux, yep, works great in a VM
The massive DC overbuild does not match demand, prices tank in 3-5 years.
Third possibility: some approach like Taalas renders the current storyline meaningless. Would put 3 in 10 odds of this happening but I'd looove to see it.
Fourth: entire planet gets profoundly sick of emdashes, we all move back into caves and live in eternal gratitude of the moment humanity woke up to how little all of this really matters.
Anthropic sells two products: a consumer subscription with a UI, and an API with metered pricing. You want the API product at the subscription price. That's not a principled stance about interface freedom, it's just wanting something for less than it costs.
The nvim analogy doesn't land either. Nobody's stopping you from writing your own client. You just have to pay API rates for it, because that's the product that matches what you're describing. The subscription subsidises the cost per token by constraining how you use it. Remove the constraint, the economics break. This isn't complicated.
"I don't give a shit about Anthropic's credit liability," right, but they do, because it's their business. You're not entitled to a flat-rate all-you-can-eat API just because you find metered pricing aesthetically displeasing.
It almost makes me feel sorry for Dario despite fundamentally disliking him as a person.
trying to hacksmash Claude into outputting something it simply can't just produces endless mess. or getting into a fight pointing out issues with what it's doing and it just piles on extra layer upon layer of gunk. but meanwhile if you ask it to boilerplate an entire SaaS around the hard part, it's done in about 15 seconds.
of course this says nothing about the costs of long term maintainability, and I think everyone by now recognises what that's going to look like
> “We are also working on providing a new licensing arrangement which will allow third parties to apply to use our data. We will provide more information on this in the coming weeks.