And Anthropic has still generally been eating their lunch at the top end.
1,256 karma · joined September 3, 2015
And Anthropic has still generally been eating their lunch at the top end.
Same teamates also post 'Sol deleted my git repo!' or 'Sorry, ignore those 300 PR comments i was just looking!' ~once a month.
I'm sure there are bad MCPs and great CLI tools that parse poorly/well via harness, but I'd be curious on an better research study.
75 y/o MIL is still fighting my wife to take her middle/highschool 'stuff' back to this day. She won't throw it out though.
I mean, you're probably thinking 30 years ago; not any time recent. Indy devs are still a thing though.
I mean, I used Mac on ~every machine i touched for about a decade, got annoyed and haven't touched one in ~5 years now. It's not unheard of.
> Beautiful, Fun & Agentic Linux
'agentic' linux
Great hardware. I don't miss OSX though.
If you're doing the training yourself, you at least have a verifiable supply chain and an audit trail. Today, we have no idea if a black-box model handed to us is coded to recognize specific domains or patterns and back door an application in a sneakily targeted way.
Black box is a black box. Many open-weight models clearly haven't been trained or created in the manner their creators claim---which raises the obvious question: if they lied about the recipe, what else did they lie about? Putting those models in a production capacity scares the living crap out of me.
That said, building from scratch isn't about achieving mathematical perfection or manually auditing 15 trillion tokens---that's impossible. It's about eliminating third-party supply chain risk and having actual governance over the pipeline.
Of course, that doesn't mean we can magically guarantee gradient descent won't produce weird emergent behaviors, or that we can blindly trust OpenAI not to backdoor things. But at least with the latter, you're making a calculated operational decision rather than blindly trusting an opaque black box of entirely unknown provenance.
I'd update this to
'I don’t LLMs for anything that needs to be private'
What's to prevent the LLM from sliding a heavily obfuscated binary blob into the application that does nefarious things? If you aren't creating the LLM itself from scratch, I don't feel it can be trusted.
i'm... not sure? This assumes ~stagnation in task-possibility. We've had ~exponential progress for like 3+ years now; I'd have never dreamed the tooling I hammer daily would exist in my lifetime just.. 3? years ago. And it's improving daily.
Maybe Open will win, maybe Closed will keep pushing the envelope. The world here is raw enough i don't think anyone can make any significant claim other than 'holy shit this is useful and moving Fast'.
This is the mental mental leaps I'm struggling with here. Did you not live through that era where they were explicitly and repeatedly called out as 'attacks'? They were generally tolerated/hardenee around as they provided value-in-discoverability.
Because they aren't giving you a cheaper service that fits your use case.
Best Case scenario, it's a trillion-dollar behemoth stealing from a billion-dollar behemoth so they can add their own explicit restrictions/weights on top to influence the masses.
There is no 'robin hood' here, any perceived value you get is clearly and explicitly tainted. "I don't care if it doesn't show me non-party-line results - It makes me a cheap UI !". Ethics/morals be damned.
You don't trust the multi-billion dollar behemoth, but you trust the militarized multi-trillion dollar behemoth to play 'robin hood'?
i can't get my brain around the mental loops here.
From the article -
> 28.8 million exchanges with Claude through almost 25,000 fraudulent accounts
wouldn't that be considered an attack? Not sure what I'm missing here.
the statement isn't "GLM 5.2 has large token usage", it's "GLM 5.2 has large token usage vs modern Opus".
I haven't used it, but this wouldn't surprise me. I see ~30% lower token usage for better results with Opus 4.8 vs 4.6 (and i had great results with 4.6)