2,395 karma · joined October 23, 2010
It is getting questions like "David has 18 apples and Ivan has 7 apples. How many apples do they have together?" wrong half the time, while Gemma3 12B could very consistently answer that. Other smoke tests (like Chinese translation, and the infamous "Rs in Strawberry" test) also show poor results.
I don't know if it is a quantization/release issue, if the parameters needed for accurate responses have changed (i.e. it needs "thinking" tokens to handle its base error rate), or if the model has been so focused on audio/video that the text processing is bad.
1. A bunch of people with new Claude Code codebases in December now are working with a larger codebase, causing more context. Claude reads a lot of code files, and doesn't effectively prune from the context as far as I can tell. I find myself having to hint Claude regularly about what files to read (and not read) to avoid having 75k of unrelated files in the context window.
2. Claude Code tries to do more now, for the benefit of people who don't know exactly what they want. The trade-off is that it's worse at doing exactly what people want, when they do know. The "small fix" becomes a large endeavor for Claude.
The project was also already dead; the English Wikinews has had 10 "articles" posted in the last 3 weeks, two of which were trivial sports stories (a second-division Queensland football match, and the retirement of a pitcher whose last substantial year in MLB was 2019). The most recent story is that an amateur jazz group recently played at a library.
It will no longer be an attractive nuisance to the few who stumble across it. Rest in peace.
GPT-5.4-nano is less impressive. I would stick to gpt-5.4-mini where precise data is a requirement. But it is fast, and probably cheaper and better quality than an 8-20B parameter local model would be.
( https://encyclopedia.foundation/benchmarks/dashboard/ for details - the data is moderately blurry - some outlier (15s) calls are included, a few benchmark questions are ambiguous, and some prices shown are very rough estimates ).
For many "simple" LLM tasks, GPT-5-mini was sufficient 99% of the time. Hopefully these models will do even more and closer to 100% accuracy.
The prices are up 2-4x compared to GPT-5-mini and nano. Were those models just loss leaders, or are these substantially larger/better?
Two months ago, Claude was great for "here is a specific task I want you to do to this file". Today, they seem to be pivoting towards "I don't know how to code but want this feature" usage. Which might be a good product decision, but makes it worse as a substitute for writing the code myself.
The demo leans heavily on "choose the words for the sentence", which avoids spelling/keyboard issues, and maybe generalizes around the problems of N->N language maps better. The "decoy selection" for multiple choice answers also isn't great - I am getting sentences mixed with numbers for the translation of "three".
It also has the Duolingo-esque audio "reward" sounds. I personally hate them, but a lot of people feel otherwise.
The "outliers" are posts that were deleted and re-submitted, or possibly boosted by the moderators. The "visible for 44 days" post was only submitted 14 days ago; either it was deleted or it is a glitch in this guy's script.
Please.
It is priced at about twice what people will expect, has very limited apps, and very limited "immersive media". Some of that will get better over the next year ... but a promise of "better next year" won't move units today.
And she's not even complaining about not going to Stuyvesant or LaGuardia?
This is NYPost trash, not a real concern.
My prediction is that they will raise a nine-figure sum over the next decade, and never release a product that comes close to the performance of an NVIDIA card today.
I am still digging through the 54-page paper to try to find the data set for this "death penalty" test to tell if there is anything there beyond "people who use more violent language tend to be viewed as more violent".
They do comment on the dialect issue: << Appalachian English evokes them to a certain extent (m = 0.015, s = 0.030, t(89) = 4.8, p < .001), but much less strongly than AAE (m = 0.029, s = 0.053, t(89) = 5.3, p < .001), a trend that holds for all language models individually (Figure S11, Table S14). The difference between AAE and Appalachian English is found to be statistically significant by a twosided t-test, t(178) = 2.3, p < .05. The fact that Appalachian English is associated with the Katz and Braly (1933) stereotypes to a certain extent is not surprising since the two dialects share many linguistic features (e.g., usage of ain’t), and the stereotypes about Appalachians bear similarities with the stereotypes about African Americans (e.g., lack of intelligence; Luhman, 1990) >>
As far as the underlying research paper: the researchers seem to be conflating "low-status English dialects" with "African American English". In particular, I have never considered the use of the word "ain't" to be associated with a certain race.
If the researchers assume "African Americans are low status" and conclude "African Americans are associated with low-status jobs", the conclusion is entirely about the researchers, not the LLMs.
The research paper's Git repo at https://github.com/valentinhofmann/dialect-prejudice does nothing to ameliorate these concerns.
1) The panels being de-commissioned now aren't modern panels, they are 20+ years old. They were substantially less efficient (per square-meter) than today's panels when they were new, and they (probably) are aging worse.
2) A lot of these "the solar panels have to be deconstructed for economic reasons" arguments are the same arguments as "Coyote v. ACME has to be destroyed for economic reasons". The reasons are completely fake; designed purely to feed Moloch.