HNHacker News
TopNewBestAskShowJobs

scronkfinkle

314 karma · joined February 9, 2026

submissionscomments
scronkfinkle··on Clef: Open-weight decision models, and new RL fine-tuning platform
> no output price listed

It's weird to think of these kinds of models as having "output tokens". Cross-encoder approaches like Laya add a [MASK] marker per option, but nothing is generated the way an autoregressive transformer generates. It's one bidirectional pass over your input, then a small head scores each option, so you wouldn't really pay for output as much as only input

scronkfinkle··on Ollaya – Ollama for open-source, Jev-style decision models
Yes. JEV generalizes better because they probably have an enormous corpus and trained on it for a long time. Laya's out of the box model is much weaker. However, in the age of LLM's it's incredibly easy and cheap to generate large datasets to fine tune laya for your task, and the training loop is pretty quick and cheap too.

It's so easy that I question why I would ever pay for JEV when eventually I'll have done enough random things that I will also have a large corpus and likely a general model as well.

scronkfinkle··on Programming Tutorials Are Dead
You're correct that writing code by hand has been heavily displaced by AI programming, but understanding how things work isn't really an "antiquated hobby" and this mentality causes a lot of people to incorrectly assume that latent knowledge of how to program is no longer useful.

It's still pretty undisputed that if you know how to write code, you can generate better AI code than someone who doesn't, and do it with less tokens. I think it's more like in university where in other engineering disciplines you learn the math and underlying concepts and are forced to do them by hand. This gives you an understanding of what you're doing, but in the real world you, ya know, just plug that stuff into matlab.

scronkfinkle··on PS5 Linux lead quits: "a bunch of noobs using LLMs" that "they don't understand"
> remember hacktoberfest 2020?

Oh man, if you read the threads about that it's such a time capsule of a different era:

e.g. from this thread: https://news.ycombinator.com/item?id=31628342

> I am honestly surprised how little SPAM there is on GitHub in general. Please don’t take that as a challenge!

scronkfinkle··on Measuring the sloppiness of code
There is some sense of rose-tinted glasses of pre-LLM coding. A lot of human written code, particularly at the enterprise level, was of low quality well before AI automated it.
scronkfinkle··on The Waymo effect: how AI is quietly making research less collaborative
I wanted you to be wrong, and to be able to make this an example of us over-reacting to certain trigger words created by AI, but unfortunately I just scanned the first couple paragraphs with pangram and it reported 100% AI, so you're probably correct.

I get a funny feeling in my stomach over the idea that common and effective means of communication (i.e. it's not X it's Y) have become faux pas to use because of AI. I think it's something about these phrases being taken away from us more-so than the AI inventing them.

scronkfinkle··on Claude is only available to people over 18 years
That's a valid take. The issue I'm wrestling with is the inevitable attempts to point a finger at who is responsible when bad things happen. If you claim it is on the tech companies to know if a child is online, then they will take the safest path for them by forcing ID verification in independent adhoc manners. This is what we are seeing now, and to me this is the worst situation. If you shift the responsibility more to the parents (this child was using a device that didn't send the `PARENT_CONTROL` flag, therefore we assumed they were an adult) it's a completely different conversation. Furthermore, if something happens to a child AND it's obvious from traffic logs that the platform willingly knew a child was being talked to, then that is also a completely different situation than the first.

Leaving it all completely deregulated and/or letting platforms implement it themselves to varying levels of success and personal invasion feels like the worst option to me.

scronkfinkle··on Claude is only available to people over 18 years
I think you may be misunderstanding a bit. There's no forced verification. It'd be more of an RFC that gives parents the ability to communicate their underage child is using the device without revealing or verifying any further information. Or do you suspect that giving any ground will cause the ID verification and centralized behavior?
scronkfinkle··on Claude is only available to people over 18 years
I once heard someone suggest that this should be on the OS level and I'm slowly coming around to the idea. There should be some kind of OS level flag that can easily broadcast to products that a child is using the device. No ID verification required, the parent is responsible for setting it, and all downstream applications are regulated to respect it, If applicable to the product. This removes the incredibly invasive ID verification, and places responsibility on the people responsible for the minor. Companies like meta then can be sued not for failing to detect children (this is, evidently, not working on any platform trying to enforce this currently), but instead can be judged in a black and white manner (i.e. is the child version of meta too predatory to children?)
scronkfinkle··on Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
But I don't want to use your CLI. I already have my own harnesses and workflows. The friction is too high to "just try out" a new model like this. It would be preferable if I can evaluate it over, say, open router like all the other models and then decide from there if it's worth downloading a bespoke tool chain for only 1 lab's models
scronkfinkle··on Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
Please correct me if I'm wrong, but this appears to require Devin to use? I'm disappointed to see I need to use a bespoke platform to interact with this agent, to the point that I probably won't be trying it.
scronkfinkle··on Harvard study predicts most suicide attempts a week in advance
That is incredibly dismissive to people undergoing something that is causing them to temporarily have these thoughts and need intervention
scronkfinkle··on GPT-6 Astra on robot arms
LLM's are a funny technology because on the one hand this is all undeniably impressive at the rate of what's changed from them, and yet despite that I find myself disappointed by the lack of breakthroughs for things I don't find interesting. I like math and programming, and LLM's are pretty good at it, when are they going to get good at folding laundry for me? I think a lot of robotics work promises to solve this category of "boring" breakthroughs, and I'm optimistic we'll be able to achieve it, i just wonder when
scronkfinkle··on GPT-6 Astra
In what way do they have a moat? A cursory look at https://artificialanalysis.ai/models/gpt-6-astra#intelligenc... it lands at 61, only a single point above glm 5.3 while costing significantly more.

The only moat they appear to have is by hoarding compute, and the current trajectory of hardware shows that isn't permanent either for very long

scronkfinkle··on Claude Fable 5.1 and Claude Mythos 5.1
Has anyone been able to get anything substantial done with Fable in the first place? I more or less had totally given up on using it since the alignment checks were so sensitive that it pretty much always threw me back to Opus.
scronkfinkle··on Ubisoft's FOR HONOR will block SteamOS / Linux players starting September 10
There's a couple different ways to look at this. From a charitable point of view to ubisoft, they do not advertise support for Linux. Instead they openly state it's a window's only game. So when a third party (i.e. Valve) makes shim that grants linux functionality, it's not really on Ubisoft to maintain that.

However, from a consumer perspective, this is incredibly frustrating. The average person does not think of "linux vs. windows". They just bought the steam deck and when they bought For Honor it was sold on that "platform". Again, it's hard to blame Ubisoft for this conundrum since they were never necessarily trying to be steam deck compliant to begin with, but it definitely feels wrong to take the game away from a player in this perspective.

All that being said, historically across many games such as Rust, many are quick to blame linux for cheating, but pragmatically banning linux has yet to measurably reduce the amount of cheating that players report experiencing, so I don't agree with the decision at all. It's a footgun disguised as a band aid

scronkfinkle··on Highest-ever ocean temperature measured as powerful El Niño forms
Sometimes I wonder if the anti-immigrant rhetoric that has been growing more popular is a tactical tool used by individuals who don't actually believe in it but instead have accepted that they would rather be wealthy and destroy the planet and are pricing in the consequences of that decision.
scronkfinkle··on There's no reason for software to be slow anymore
It's becoming increasingly evident this is not how things are going to shake out. Even without frontier models, running Qwen 3.8 27B has demonstrated for me and others a "good enough" competency at general programming. Additionally, large open weight models have proven themselves as viable alternatives and we're still in the early years of dedicated hardware
scronkfinkle··on An update on leaving Gmail for Fastmail
Maybe I am misunderstanding but I don't think this is true. Indeed you pay for the inbox, but it's not the primary domain. You can reply from any of the email addresses you created (which doesn't cost extra) and it automatically sets the FROM to be the email it was received on
scronkfinkle··on AI is removing the middle class of software engineering?
I think of it more as "the automation of the stackoverflow engineer". In enterprise software, there has always just been a non-negotiable large volume of code that was required to be written. This has traditionally been offloaded by having seniors do the hard thinking then distill it into a jira ticket which could be handed off to an engineer that'd actually write the code and punch every hiccup into google along the way. This hand off is no longer necessary as that same senior can just kick off an agent and have it handle the implementation for them.

I've heard some refer to this as a "nature is healing" scenario for the industry where if you only signed up for a high paycheck and didn't care to think critically about any of the work you're doing then this will be painful because that previously manual process has been automated. The floor of what's necessary to be considered valuable has been raised.

scronkfinkle··on What sort of maths are LLMs good at?
> So then you need to explain ARC-AGI-3: https://arxiv.org/abs/2603.24621

I don't, we originally had the turing test which was designed to determine human intelligence by its ability to imitate us with natural dialogue, but we've since defeated that. I stated "to me" because it's my personal opinion on a definition whose goalpost will probably never stop being moved.

> Back 1996, EQP automatically solved the Robbins conjecture. But nobody concluded EQP was generally intelligent.

EQP doesn't have the 3 criteria I outlined, which were different than "solving a math problem"

scronkfinkle··on What sort of maths are LLMs good at?
> A good sign that LLMs have reached human level for a much wider class of problems will be if they start proving theorems using methods that, like much of the very best human mathematics, are new and surprising but that with hindsight come to seem beautiful and natural. They should also be methods that are difficult to stumble on by accident. It is hard to say precisely what would count as such a proof, but I think we’ll recognise it when we see it

Agreed. I find that after seeing these results from OpenAI we undeniably have a machine that has:

* General knowledge of nearly every subject humanity has ever learned

* The ability to simulate reasoning (albeit sometimes not very well) with that knowledge

* The ability to reference across the domains of knowledge

To me, this is more or less what I would think "Artificial General Intelligence" is. It's the cumulative knowledge of all general human intelligence, baked into an artificial form, which can then use that knowledge to achieve novel goals.

In many cases of mathematical breakthroughs there is an insight that comes from just happening to know a combination of already existing ideas and then combining them to solve that problem. This is where having that general knowledge seems particularly strong because we can run these machines for weeks on end effectively trying to brute force.

That being said, I could never imagine an LLM in its current form inventing something as elegant as the Fourier transform.

scronkfinkle··on Show HN: Alphabet Soup, a multiplayer game, build the longest word to win
Fun game. Would be nice if there was a single player mode that simply prints the letters and at the end shows the longest word that was missed. That way you can play it offline and place a phone in the center of the table and people can play around it with pen and paper or by just thinking of the word and everyone announces at the end.
scronkfinkle··on Ten advances in mathematics and theoretical computer science
The impressive/surprising thing is your premise because we effectively can spin up an army of mathematicians now
scronkfinkle··on The End of an Era
> and writes entire operating systems from scratch

ehhhh, we're not really there. Not saying it's impossible to reach in the future but large tasks like this are still out of scope for LLM's beyond demoing toy applications. Impressive nonetheless, but I can't tell an LLM to reproduce microsoft word and have it give me a useful enough tool that doesn't contain an immense amount of bugs and require a human in the loop to replace it.

scronkfinkle··on Nvidia, Microsoft, Meta warn against overregulating open-weight models
I don't feel it's meaningful to berate the point anymore about the hypocrisy of the American labs. Now as the sentiment and effort from them to push for regulation increases so does my perception of how pathetic they are behaving. I'm not sure how else to say it other than that it's honestly just embarrassing to watch China and others run circles around us in the USA.
scronkfinkle··on “We have information that Moonshot distilled Fable for the development of K3”
so they distilled one of the best models in the world AND released it for free to everyone. Where can I send them flowers as a thank you?
scronkfinkle··on Help I accidentally a wigglegram
Meanwhile non-frontend folks decide to call one thing "threads" and another thing "strings" and have them be completely unrelated to each other.
scronkfinkle··on Corporate America Is Starting to Ration AI as Cost Skyrockets
On the one hand, organizations are without question using LLM's well beyond what is actually necessary, and as reality kicks in they're forced to scale back accordingly. However at the same time, on intervals counted in months, we're seeing breakthroughs both in hardware and software that dramatically reduce the cost of inference.

Between corporate FOMO and the rapidly decreasing costs of actually running LLM's I'm interested to see at which side of the spectrum these two meet

scronkfinkle··on Microsoft reports AI is more expensive than paying human employees
The title seems misleading, and reading the article explains the reason more clearly. There's nonsense OKR's and objectives at these companies to burn as many tokens as possible. It turns out that when you make a metric out of token usage, it unsurprisingly ends up becoming extremely expensive.

Inference is affordable, and you don't need a SOTA proprietary model to get a lot of use out of this technology. While you likely will still need a human engineer for quite a while longer, I don't agree that some number of humans + an LLM is going to be (or will ever remain) more expensive than just hiring more humans.

Page 1 of 2Next →