HNHacker News
TopNewBestAskShowJobs

BoorishBears

7,288 karma · joined August 18, 2016

submissionscomments
BoorishBears··on Decisions API is in public beta
Decisions voice isn't a product for anyone else who was confused: it's a canned guide for hooking up a voice model to the decision model browser use thing
BoorishBears··on Decisions API is in public beta
I'd imagine it's chasing the lowest possible latency
BoorishBears··on Sonnet 5.5
Wow with the Dev Day announcement on "partner apps", it sure seems like OpenAI just kneecapped a huge area of growth for Chinese frontier models...
BoorishBears··on Sonnet 5.5
I've seen the opposite: if you're doing work that doesn't need the frontier, it's really hard to beat lower-tier frontier models.

Luna is obviously very competitively priced, and I'm expecting Haiku 5.5 will be strong based on this release and get back to more competitive pricing since the model family shrank this generation (though Anthropic has proven my expectations wrong on the latter before)

Where GLM, Kimi and co shine for me is when you need to offer near-frontier capabilities in your product and straight up can't afford frontier models: if you're offering Opus in a product with API pricing, $20 a month Claude Pro is offering about $500 of comparable usage in a harness that flexes to a lot of tasks.

Offering GLM/Kimi increases the max complexity of problems you can solve successfully compared to stepping down to Luna/Haiku, while letting you offer a reasonable amount of usage. Once $20 a month Pro is comparable to "just" $100 of usage in your product, it's much easier to close the gap with UX, a better constrained harness, etc.

BoorishBears··on Sonnet 5.5
I know this isn't a model thing, but why do all the labs blow at product outside of models?

Aren't you still getting paid more money than god to write React if you work at Anthropic? I wasted 5 minutes digging into random stupid nooks and crannies in the desktop app to find where I could update: only to find on Linux you need to use apt.

How hard would it be to put a notice where the normal Check For Updates goes that says "This install is managed by [package manager], use [command] to update"

AGI is going to be so awful for product quality on the more basic things. It feels like these are small papercuts that humans would implicitly smooth over, that RL'd models are actually getting worse at dealing with because of their single-mindedness about completing the given task.

BoorishBears··on Opus 5.5 is good at explainer videos
The exercise is trying to reduce the world into significant/not significant efforts and then write some things off as only worth it if you don't care about the former.

Also I honestly thought we were two arrogant people conversing in good faith: you don't find unarrogant people saying things like "If you don't care so much about adding something significant to the world..."

BoorishBears··on Opus 5.5 is good at explainer videos
I used to think like this, but eventually I realized it's mostly just a way to pat yourself on the back. And ironically, I rarely saw people doing significant things waste energy talking like that, since there's no benefit in the exercise.

Very few things can't be reduced to insignificance. MacOS was just cribbing PARC, Facebook is just a glorified PHP forum, Dropbox is just SFTP, etc. etc.

Even in deep-tech, Zipline is just wrapping from deeper-tech (batteries motors etc), GLP-1s were a VA throwaway that dusted off, etc. etc.

LLM wrapper doesn't mean anything, it's an implementation detail. Besides being an LLM wrapper what is a given thing?

The beer app wasn't just a beer app, it was the intersection of the first time accelerometers were doing something that the average consumer could interact with in their pocket, the first time there was something to spend money on for your phone besides a wallpaper, a ton of things.

If anything now when something seems trivial or like a toy, but it has a lot of traction, I want to spend energy figuring out why is it more meaningful than it appears.

BoorishBears··on OpenAI is well positioned to fast-follow Jev
"all of which require everything you've mentioned at minimum"

Sorry who else did everything I mentioned? I think the guy behind Laya tried after noticing Jev's traction... but the site's auth went down and has stayed down for a day now.

"substantially more straightforward than almost any other product in its category"

More straightforward than the spite projects based on constrained decoding? Or Laya with it's couple of days post-training ModernBERT?

-

I have no doubt other teams can build models like this and I've love for a frontier lab to give us an even smarter model with these ergonomics... but in the rush to show Jev what's up, we're mostly getting slop.

PS: I don't know anyone who's done anything of note who uses trivial like that. The commentariat do, and the "I could have done that" crowd do, but I don't pay much attention to them until they actually do the thing.

BoorishBears··on OpenAI is well positioned to fast-follow Jev
... why didn't you link to the actual benchmark which does have Jev at the top?

https://benchmarkheaven.com/jev-models

You linked to some weird subtable that labeled: " Not the default — not the JevBench Score", that can only be reached after you see what I just linked... lmao are you really this hard up about things?

Also every single question (even in the hard set) is single dimensional?: https://github.com/fstandhartinger/jevbench/blob/main/datase...

Jeeze, this is getting sad. I guess after all the mass-psychoses where people thought pointless things are going to change the world, we were due for a mass-psychosis where something interesting just has to be pointless?

BoorishBears··on OpenAI is well positioned to fast-follow Jev
No, all they had to do was come up with a quality post-training recipe, production inference stack that wouldn't fall over, GTM, documentation, schemas, etc. etc.

(also most signs point to this being LLaDA 2.0-adjacent so throw in solving some substantial mid-training)

I think it's 100% a hot take to call what they built trivial. Or at least it used to be.

There was a time when that kind of stuff was something between sour grapes and cluelessness about the gap between an idea and an actual commercial product deployed at scale, but now that's just weirdly normalized.

In fact, if anything I'm the weirdo for repeatedly taking issue with the way people are trivializing it ¯\_(ツ)_/¯

BoorishBears··on OpenAI is well positioned to fast-follow Jev
Expecting a strong zero-shot performer to perform worse in a low data regime?

That only makes sense if you try to rope in data previously used to establish the model's priors, but that wouldn't make sense in this context. That same additional data is what enables things like...

> use generalized models to generate ad hoc specialized classifiers.

BoorishBears··on OpenAI is well positioned to fast-follow Jev
But zero-shot classifiers with this level of intelligence, world knowledge, ergonomics, cost profile, and ease of use are new.

I feel like good engineering doesn't just ignore those things, or at least it didn't before recently. Now I guess social media has added a pressure to reduce everything to a hot take.

BoorishBears··on Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
Finetune and infer: One instance ModernBERT didn't have the learning capacity for a single problem in the shape of my subjective preference task with finetuning, do you not have the basic research taste to realize no conceivable post-training recipe will result in an instance that can zero-shot hundred plus similar questions that vary with each sample?!

You really need to try that to find out?

And again have you actually tried Jev? It has a ton of world knowledge: it's able to infer user personas based on TV show watch histories using fairly recent titles... where the hell do you think that capability is emerging in 395M params?

The irony is if you really want to die on this hill, there are much better angles by focusing on LLMs that've had diffusion heads attached for fast inference with as much of a constrained decoding intelligence penalty: at least that'd put you in the ballpark.

I was being charitable that you know the field and are clueless about Jev, mea culpa for giving you the space to think I'm the one that's missing something.

BoorishBears··on Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
Please see my other reply, I was not exaggerating when I said this feels like asking why ChatGPT is different than a Markov Chain.

You're going to post-train 100s of instances of BERT? Traditional ML had world knowledge more than a fart?

The closest/fairest comparison is still an LLM, but no one has actually chucked enough compute at post-training to make a better Jev yet.

I'm sure in more time that'll happen, and so my excitement is expanded to Jev-like things... but so far most Jev like things are this weirdly reactionary attempts to steal thunder: is it so bad if we have some team actually invest in a quality post-training receipe to compete?

BoorishBears··on Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
I did a GRPO run (multiple now actually) with a per sample rubric that leans heavily on subjective preference judgements that BERT wouldn't have the learning capacity for: not to metion you'd need to finetune hundreds of instances and host them somewhere.

And even if BERT wasn't woefully underintelligent for the task... have 100+ instances of BERT running locally faster than Jev API response times? Sweet rig you must have...

LLMs would not be fast enough without constrained decoding tricks that people fundamentally don't seem to understand make the models much dumber, and sure wouldn't be cheaper or faster.

Again I feel this deep discomfort because presumably you're somewhat intelligent but your opening salvo made it hard not to scream DO YOU EVEN HAVE A SINGLE CLUE WHAT IT DOES instead of giving you my actual answer... yet you're speaking from the chest! If I didn't try it for myself I would have been 100% sucked into you and this ocean of clueless negativity.

-

I apologize if that sounds harsh but it angers me because why should I have to deal with this kind of noise in an already insanely noisy environment? What do you gain from being cluelessly pessimistic?

And dwelling a but more I think it breaks one of my most used filters which was assuming people who know the "old world" of AI/ML are better at judging the "new world" full of hype and noise. Maybe my frustration is also just fear that things moved so quickly that the "old world" is becoming increasingly irrelevant. That'd be really disappointing.

BoorishBears··on Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
Jev is creating a sort of identity crisis for me, because the number of absolutely clueless folks parroting the classifier thing is the first time I've seen this sort of mass psychosis in CS upfront.

Like even 5 minutes of tinkering captures why this isn't anymore like BERT or any past classification model than ChatGPT is like those old Markov Chain generators, yet folks cannot shut up about how this is nothing new.

Absolutely scary and makes me wonder how much of the field is just people super confidently discrediting otherwise promising/interesting directions for development for a cheap dunk!

BoorishBears··on ChatGPT now knows what you do on other websites via ad collector
I think people who make being smart their whole identity are a out to use AI to turn the rest of the population into indentured servants.

The comparison between ELIZA and LLMs is valid you boil it down to "humans evolved for 6-7 million years, had spoken language for 500k years, but have only had something non-human that could generate convincingly novel language well enough to hold a conversation for a few decades".

There's no inherent reason it can't turn out having a non-human generate convincing enough language for conversation isn't a complete evolutionary blindspot the same way the short form feed has pretty much one-shotted society...

BoorishBears··on A Necessary History of the Oddest Letter: W
Why not simple, W: making a mouth shape as-if to say "What", then sharply exhale a puff of air instead
BoorishBears··on Qwen Image 2.1
Qwen's latest image models have a ton of distillation from gpt-image, same with Grok Imagine.

Even the artifacts are getting picked up.

BoorishBears··on Salesforce Global Outage
Looks like a region list to me, maybe just with a lot of regions
BoorishBears··on Salesforce Global Outage
My gut instinct is that this is about Dreamforce with the rickshaws and whatnot
BoorishBears··on German Rheinmetall open-sources its Battlesuite connected weapon system protcol
TIL there are people using DDS for non-realtime
BoorishBears··on Introducing System One Models and Jev
Did you see the video where it plays Doom, it made it click for me
BoorishBears··on Show HN: The bottom 50% of U.S. households are short after essentials (BLS data)
Past a point of wealth, it seems a lot easier to increase wealth through things of no intrinsic value to society.

If I have enough money, I can buy a lot of businesses (or be in a controlling position) and run them poorly (read: efficently) and as wealth concentrates the capital for a new entrant to compete becomes scarcer and regulation on new competition mysteriously materializes, and without any new competition I can raise prices and no one can really do anything.

And now I have more money to do this more times in more places, and wealth is more concentrated, and legislation just seems to go my way more and so on.

That vague playbook is being applied to just about everything today.

Spirit Airlines went through two bankruptcies, and the people who lead the first one were liquidating it while it was still running to pay themselves back for the first one, and turning debt they had forfeited for equity back into debt to increase how much they'd take out of the remains.

I mean you can't even escape that playbook in childrens sports! That playbook is being run into making children sports league a profit squeezing cartel as we speak.

It's just such a simple playbook that if we really go all in on the idea that anything that's technically legal to make money will be allowed, then having more money and basic "high agency" mindset will lead to making money in increasingly perverse ways.

To me the end game seems to be a race between authoritarianism (will these high-agency wealthy types just be able to keep the poorer class under lock and key with TikTok feeds to placate) and collapse (maybe they don't get the poor under deep enough control before the poor realize they've been given a deal that justifies trying to raze it all and start again)(edit: or maybe better framing, they realize they collectively have no meaningful part in the wealth and raze it)

BoorishBears··on Compute-efficient pretraining and scaling to trillion-parameter models
This seems weirdly pessemistic: frontier labs have much stronger pretraining than most open weights models

And reading the release it feels very obvious this is also a ton of aligning their data mix with coding and science: we don't know that this model doesn't have terrible world knowledge or is ruined for anything related to subjective preference

They also repeatedly mention knowledge almost as if they saw that skepticism coming, but then limit knowledge to topics where more understanding of how code/scientific writing looks would produce the same graph as having actual world knowledge maintained.

That's not nefarious (they literally build coding models), but it also means the resulting model isn't necessarily competitive with a frontier model in a broader way.

This feels like the inverse approach to what Thinking Machines did with Inkling (trying to train as "un-spikey" a base model as possible)

BoorishBears··on Ask HN: Why can I only downvote select comments?
I don't think HN staff would ever do that, but what was the topic?
BoorishBears··on GPT-6 Astra
This isn't the gotcha that you think it is: AGI's original definition is being able to do any task that requires human intelligence.

The unlock isn't AGI smart enough to invent quantum mechanics, it's suddenly being able scale human intelligence using grains of sand instead of decades of food and energy and nuturing.

BoorishBears··on GPT-6 Astra
Cool. Being sole proprietor of AGI 15 years ago should result in monuments and religions devoted to you today.

Cancer should be cured, and we should be a post-quantum interstellar fusion-powered civilization.

I wish the AGI crowd would finally shut up now that it's clear no one is even trying for AGI (OpenAI revised that to "$100B in profit")

What we're getting is incredible, where we're headed is incredible, but some people have such a fetish for futuretelling they can't just shut up and enjoy the ride.

BoorishBears··on GPT-6 Astra
After they buy AZ, making this name foreshadowing
BoorishBears··on GPT-6 Astra
Would be surprising if it's not on all 3 major clouds soon enough because that's been their general strategy since said break up
Page 1 of 34Next →