7,288 karma · joined August 18, 2016
Luna is obviously very competitively priced, and I'm expecting Haiku 5.5 will be strong based on this release and get back to more competitive pricing since the model family shrank this generation (though Anthropic has proven my expectations wrong on the latter before)
Where GLM, Kimi and co shine for me is when you need to offer near-frontier capabilities in your product and straight up can't afford frontier models: if you're offering Opus in a product with API pricing, $20 a month Claude Pro is offering about $500 of comparable usage in a harness that flexes to a lot of tasks.
Offering GLM/Kimi increases the max complexity of problems you can solve successfully compared to stepping down to Luna/Haiku, while letting you offer a reasonable amount of usage. Once $20 a month Pro is comparable to "just" $100 of usage in your product, it's much easier to close the gap with UX, a better constrained harness, etc.
Aren't you still getting paid more money than god to write React if you work at Anthropic? I wasted 5 minutes digging into random stupid nooks and crannies in the desktop app to find where I could update: only to find on Linux you need to use apt.
How hard would it be to put a notice where the normal Check For Updates goes that says "This install is managed by [package manager], use [command] to update"
AGI is going to be so awful for product quality on the more basic things. It feels like these are small papercuts that humans would implicitly smooth over, that RL'd models are actually getting worse at dealing with because of their single-mindedness about completing the given task.
Also I honestly thought we were two arrogant people conversing in good faith: you don't find unarrogant people saying things like "If you don't care so much about adding something significant to the world..."
Very few things can't be reduced to insignificance. MacOS was just cribbing PARC, Facebook is just a glorified PHP forum, Dropbox is just SFTP, etc. etc.
Even in deep-tech, Zipline is just wrapping from deeper-tech (batteries motors etc), GLP-1s were a VA throwaway that dusted off, etc. etc.
LLM wrapper doesn't mean anything, it's an implementation detail. Besides being an LLM wrapper what is a given thing?
The beer app wasn't just a beer app, it was the intersection of the first time accelerometers were doing something that the average consumer could interact with in their pocket, the first time there was something to spend money on for your phone besides a wallpaper, a ton of things.
If anything now when something seems trivial or like a toy, but it has a lot of traction, I want to spend energy figuring out why is it more meaningful than it appears.
Sorry who else did everything I mentioned? I think the guy behind Laya tried after noticing Jev's traction... but the site's auth went down and has stayed down for a day now.
"substantially more straightforward than almost any other product in its category"
More straightforward than the spite projects based on constrained decoding? Or Laya with it's couple of days post-training ModernBERT?
-
I have no doubt other teams can build models like this and I've love for a frontier lab to give us an even smarter model with these ergonomics... but in the rush to show Jev what's up, we're mostly getting slop.
PS: I don't know anyone who's done anything of note who uses trivial like that. The commentariat do, and the "I could have done that" crowd do, but I don't pay much attention to them until they actually do the thing.
https://benchmarkheaven.com/jev-models
You linked to some weird subtable that labeled: " Not the default — not the JevBench Score", that can only be reached after you see what I just linked... lmao are you really this hard up about things?
Also every single question (even in the hard set) is single dimensional?: https://github.com/fstandhartinger/jevbench/blob/main/datase...
Jeeze, this is getting sad. I guess after all the mass-psychoses where people thought pointless things are going to change the world, we were due for a mass-psychosis where something interesting just has to be pointless?
(also most signs point to this being LLaDA 2.0-adjacent so throw in solving some substantial mid-training)
I think it's 100% a hot take to call what they built trivial. Or at least it used to be.
There was a time when that kind of stuff was something between sour grapes and cluelessness about the gap between an idea and an actual commercial product deployed at scale, but now that's just weirdly normalized.
In fact, if anything I'm the weirdo for repeatedly taking issue with the way people are trivializing it ¯\_(ツ)_/¯
That only makes sense if you try to rope in data previously used to establish the model's priors, but that wouldn't make sense in this context. That same additional data is what enables things like...
> use generalized models to generate ad hoc specialized classifiers.
I feel like good engineering doesn't just ignore those things, or at least it didn't before recently. Now I guess social media has added a pressure to reduce everything to a hot take.
You really need to try that to find out?
And again have you actually tried Jev? It has a ton of world knowledge: it's able to infer user personas based on TV show watch histories using fairly recent titles... where the hell do you think that capability is emerging in 395M params?
The irony is if you really want to die on this hill, there are much better angles by focusing on LLMs that've had diffusion heads attached for fast inference with as much of a constrained decoding intelligence penalty: at least that'd put you in the ballpark.
I was being charitable that you know the field and are clueless about Jev, mea culpa for giving you the space to think I'm the one that's missing something.
You're going to post-train 100s of instances of BERT? Traditional ML had world knowledge more than a fart?
The closest/fairest comparison is still an LLM, but no one has actually chucked enough compute at post-training to make a better Jev yet.
I'm sure in more time that'll happen, and so my excitement is expanded to Jev-like things... but so far most Jev like things are this weirdly reactionary attempts to steal thunder: is it so bad if we have some team actually invest in a quality post-training receipe to compete?
And even if BERT wasn't woefully underintelligent for the task... have 100+ instances of BERT running locally faster than Jev API response times? Sweet rig you must have...
LLMs would not be fast enough without constrained decoding tricks that people fundamentally don't seem to understand make the models much dumber, and sure wouldn't be cheaper or faster.
Again I feel this deep discomfort because presumably you're somewhat intelligent but your opening salvo made it hard not to scream DO YOU EVEN HAVE A SINGLE CLUE WHAT IT DOES instead of giving you my actual answer... yet you're speaking from the chest! If I didn't try it for myself I would have been 100% sucked into you and this ocean of clueless negativity.
-
I apologize if that sounds harsh but it angers me because why should I have to deal with this kind of noise in an already insanely noisy environment? What do you gain from being cluelessly pessimistic?
And dwelling a but more I think it breaks one of my most used filters which was assuming people who know the "old world" of AI/ML are better at judging the "new world" full of hype and noise. Maybe my frustration is also just fear that things moved so quickly that the "old world" is becoming increasingly irrelevant. That'd be really disappointing.
Like even 5 minutes of tinkering captures why this isn't anymore like BERT or any past classification model than ChatGPT is like those old Markov Chain generators, yet folks cannot shut up about how this is nothing new.
Absolutely scary and makes me wonder how much of the field is just people super confidently discrediting otherwise promising/interesting directions for development for a cheap dunk!
The comparison between ELIZA and LLMs is valid you boil it down to "humans evolved for 6-7 million years, had spoken language for 500k years, but have only had something non-human that could generate convincingly novel language well enough to hold a conversation for a few decades".
There's no inherent reason it can't turn out having a non-human generate convincing enough language for conversation isn't a complete evolutionary blindspot the same way the short form feed has pretty much one-shotted society...
Even the artifacts are getting picked up.
If I have enough money, I can buy a lot of businesses (or be in a controlling position) and run them poorly (read: efficently) and as wealth concentrates the capital for a new entrant to compete becomes scarcer and regulation on new competition mysteriously materializes, and without any new competition I can raise prices and no one can really do anything.
And now I have more money to do this more times in more places, and wealth is more concentrated, and legislation just seems to go my way more and so on.
That vague playbook is being applied to just about everything today.
Spirit Airlines went through two bankruptcies, and the people who lead the first one were liquidating it while it was still running to pay themselves back for the first one, and turning debt they had forfeited for equity back into debt to increase how much they'd take out of the remains.
I mean you can't even escape that playbook in childrens sports! That playbook is being run into making children sports league a profit squeezing cartel as we speak.
It's just such a simple playbook that if we really go all in on the idea that anything that's technically legal to make money will be allowed, then having more money and basic "high agency" mindset will lead to making money in increasingly perverse ways.
To me the end game seems to be a race between authoritarianism (will these high-agency wealthy types just be able to keep the poorer class under lock and key with TikTok feeds to placate) and collapse (maybe they don't get the poor under deep enough control before the poor realize they've been given a deal that justifies trying to raze it all and start again)(edit: or maybe better framing, they realize they collectively have no meaningful part in the wealth and raze it)
And reading the release it feels very obvious this is also a ton of aligning their data mix with coding and science: we don't know that this model doesn't have terrible world knowledge or is ruined for anything related to subjective preference
They also repeatedly mention knowledge almost as if they saw that skepticism coming, but then limit knowledge to topics where more understanding of how code/scientific writing looks would produce the same graph as having actual world knowledge maintained.
That's not nefarious (they literally build coding models), but it also means the resulting model isn't necessarily competitive with a frontier model in a broader way.
This feels like the inverse approach to what Thinking Machines did with Inkling (trying to train as "un-spikey" a base model as possible)
The unlock isn't AGI smart enough to invent quantum mechanics, it's suddenly being able scale human intelligence using grains of sand instead of decades of food and energy and nuturing.
Cancer should be cured, and we should be a post-quantum interstellar fusion-powered civilization.
I wish the AGI crowd would finally shut up now that it's clear no one is even trying for AGI (OpenAI revised that to "$100B in profit")
What we're getting is incredible, where we're headed is incredible, but some people have such a fetish for futuretelling they can't just shut up and enjoy the ride.