5,745 karma · joined May 24, 2023
https://dmitriid.com
You may know me from these:
- Everything around LLMs is still magical and wishful thinking(2025) https://dmitriid.com/everything-around-llms-is-still-magical-and-wishful-thinking
- Prompting LLMs is not engineering (2025) https://dmitriid.com/prompting-llms-is-not-engineering
Until shit like this: https://www.anthropic.com/engineering/april-23-postmortem
Where people pointed out issues early and en masse, and Anthropic denied it was happening, gaslighted anyone claiming this was an issue, then begrudgingly admitted it was an issue, and then spent another two weeks "fixing it".
Or shit like this: https://www.anthropic.com/engineering/a-postmortem-of-three-...
Anthropic is in a perpetual state of "oops, these 'bugs' degraded our model quality" and only admit the issues when it's immediately obvious and visibly affects a large number of customers.
Otherwise all open benchmarks can be (and are) gamed. And it's quite hard to judge the output of a non-determenistic black box that Anthropic (or OpenAI) constantly tweak.
How do you know it's difficult if you say you don't understand it?
> Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while.
I run all the "latest and greatest" models the moment they become available to me. The amount of insanely bad code they produce remains largely the same, and largely in the same areas. And it cannot be caught by tests unless you know that bad code is there and end up with extremely bad tests anyway. I wrote about it here: https://dmitriid.com/adding-to-i-dont-read-ai-code-discourse
Main one is, of course, "to get a record from a database read all records from it, and filter in memory".
Download speed isn't the only thing that matters
The last thing people want is interacting with a chatbot for anything that doesn't require a chatbot.
> in terms of functionality and value to users, software ought to be malleable and composable.
No, it shouldn't. The last thing people want is for software to change from under them, or be replaced with the bullshit that is a chatbot.
> plaster a ‘programmable’ interface on top of unstructured data/interface meant for humans
Ah yes. I can imagine everyone "plastering programmable interfaces" on top of daily stuff they want to do. Like buying a ticket to a museum. Ordering Doorsash. Sending a meme to a friend. Renting a car...
They sell "we're only targeting military infra" and "we only escalate after unwarranted escalation by Ukraine".
They hit Ukraine's energy grid during literally the coldest days in winter. The audience's reaction? "YES LET THEM FREEZE".
Russia couldn't care less about what they need to sell. They've been hitting civilian infra right and left.
- as biology: "Use Unicode graphemes" (already used by the project)
- as generally unsafe: English text I typed without switching from Russian layout (so, gibberish, which it previously just converted to English and executed)
You are going to jail.
1. What about negligence?
2. Every follow up to every story after the news cycle moved on shows both intent and negligence. To the point of "we opened internet access and told it to hack"
Same with Anthropic. On top of that Anthropic rarely or ever admits any issues, and even if they do, you get like 6 hours of reset. Rmemeber March?
Anthropic is the last company I would trust to make any decisions. Look at any discussions surrounding their "Claude is a tiny game engine" idiocy, numerous bugs that a junior can discover, a full "plugin system" in which they neeeed a dozen files in the worst Clean Code manner to read one of two files etc.
AKA "we need 100~ish files wrtitten in the most horrible Clean Code style replete with no two files agreeing on the same naming of the same feature... to read one of two files, one of which has been a de-facto industry standard for over two years"
Aka: "an issue even a junior would've spotted if we didn't rely on Claude of 100% of our tasks"
I mean. John Ternus and other execs and senior management had nothing to do anything with it. They were all new, picked up randomly out of the box and only obligingly did the bidding of Tim Cook (or Alan Dye, or both).
But still.
Everyone keeps blaming just one guy for enshittifying Apple, be it Alan Dye or Tim Cook. As if literally no other executive or senior manager has any say in anything that happens.
Even here.
Joh Termus was just picked up randomly out of a box. He never knew what was happening at the company, is entirely new to Apple, and never weighed in on any decision besides hardware. That is why he was picked to checks notes lead all of Apple.
Ternus was one of the main people driving these eyesores. He wasn't a nobody.
Don't know about Bulgarian, but Russian punctuation is notoriously horrendously complicated. I still remember a dictation [1] which had a sentence that had a comma after nearly every word. (And this last sentence would have three)
> English spelling is better than, let's say, Greek or French.
Can't speak for Greek, but IIRC in French a collection of letters usually conforms to a rule. "eu" or "ou" will rarely represent different sounds.
Whereas one of the famous spellings of "fish" in English is "ghoti": https://en.wikipedia.org/wiki/Ghoti?wprov=sfti1
[1] Common in Soviet and post-Soviet schools: https://en.wikipedia.org/wiki/Dictation_(exercise)?wprov=sft...
That is the modus operandi of Anthropic and all of its employees (at least those active on social media).
More here: https://news.ycombinator.com/item?id=49785996
This behaviour is regardless of whether the issue is an edge case or it affects most of their users who are shouting from the rooftops
No, not really. Anthropic is notorius for ignoring issues, pretending they don't exist, gaslighting and blaming users, then spending weeks fixing things that end up being obvious even to juniors.
At the same time they proudly tell everyone how they don't look at code anymore and only run a gazillion Claude sessions.
For that you use internal communication channels.
Though we've known for a while that no one at Anthropic is capable of doing any proper communication.
See literally every isssue, even those widely reported.
But we could improve some parts. For example, voting in elections could be improved by introducing Single Transferrable Vote https://en.wikipedia.org/wiki/Single_transferable_vote (CGP Grey had a bunch of videos on this: https://youtube.com/playlist?list=PL3897F608FAD61E88&si=TzEP...)