HNHacker News
TopNewBestAskShowJobs

c7b

1,729 karma · joined September 3, 2022

submissionscomments
c7b··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Cool idea! Do you happen to have benchmarks for Strix Halo (AMD Ryzen AI Max+ 395)? I take it that Qwen3.8-Flash-Next is not supported?

And a more general question: does your engine detect and optimize for custom setups like multiple (possibly different) GPUs, eGPUs,...? Because if all you have is a stock major system like a Mac or DGX Spark, that's all you're going to care about, and there are a lot of highly optimized single-hardware engines out there that will be hard to beat in the long run. Something that automatically adapts to custom systems that don't have their own subreddits could really fill a gap.

c7b··on Show HN: Jev Plays Pokémon Red
LLMs are like lossy compression of ~all of written text ever produced, with useful recall. To the extent that the corpus contains labelled examples of the given classification task, it's not unreasonable to think that we'll be able to build a decoder for that, just like we already have a useful decoder for next-token prediction. Extend to image classification the same way we already have multimodal LLMs.
c7b··on Show HN: Jev Plays Pokémon Red
A pre-trained universal classifier that can replace specifically-trained ones would have been considered just as much science fiction in the 2010's as the capabilities of modern LLMs. I'm not sure Jev is actually there yet, but at least it sounds theoretically doable today.

That being said, one thing having been unrealistic 10 years ago and just about possible today doesn't mean that it's going to change the world the same way another technically related, previously-impossible thing did. The Jev hype gives me a bit of the "you're still early to crypto" vibes of some later altcoins. I really like the idea, I think it's going to open up possibilities for using classifiers where we wouldn't or couldn't have trained one before. I'm crossing my fingers for an open weights version to drop. But it's still just a classifier, people have built similar things before Jev, the one thing that really stands out about it is their ability to generate hype.

c7b··on Dutch governments builds alternative for Microsoft based on NixOS
Can't speak to it myself, but I'm not sure how up to date your infos are. I read a lot of reviews and it seems that the vast majority of the Steam catalogue should run fine on Linux today. The most important exceptions are games that require kernel-level anti-cheat (so if you're into online multiplayer games that could be a dealbreaker).
c7b··on Dutch governments builds alternative for Microsoft based on NixOS
You're saying this as if emulation (which Wine isn't really, more of a Win32 API compatibility layer; WINE literally stands for 'WINE Is Not an Emulator') was a bad thing. In reality, a lot of games benchmark better on Linux than on Windows today.
c7b··on Ideas on modernizing the open-source desktop
Two projects that I found brought nice, more incremental changes to the desktop were COSMIC and Omarchy. Both use tiling, COSMIC's mixed tiling+floating workflow is amazing and my new daily driver. I personally don't use Omarchy for a variety of reasons, but I'd love to be able to install its launcher and keyboard integrations.
c7b··on Jev in 25 Lines of Python
My source: https://news.ycombinator.com/item?id=49783999#49785351 I don't have an account so couldn't check, a friend couldn't replicate with a one-letter attempt (first letter A 50%) - performed by his agent, though, so not sure what he did exactly.
c7b··on Jev in 25 Lines of Python
You can also just acknowledge that we don't know what they did (and that is because they chose not to tell us). Apparently, until not so long ago Jev would spell out its identity as Q-w-e-n if asked, so we might as well assume that what they did is at least similar to what this blog poster did (who chose to tell us).
c7b··on OpenAI is well positioned to fast-follow Jev
Extremely fast hype and a very short-lived moat, those are the parallels that I see. The OpenClaw guy was smart to cash out at peak hype, and chances are Jev will be overrun by big players and open-weights pretty soon too. I really love the idea, but just don't see any sustainable moat here. People are already coming knocking [0], it should only be a matter of time until we get something like Qwen3.9-Classify (or a classify mode just baked into a multimodal LLM).

[0] https://archerhume.com/posts/jevs-architecture-unmasked

c7b··on OpenAI is well positioned to fast-follow Jev
The OpenClaw vibes are hard to miss here.
c7b··on Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
Technically, Q is picked because it has the highest probability of all letters. But it makes a difference whether the probability for Q is barely above a uniform 1/26~3.8% or whether that one letter concentrates >50%. What I remember from reading the docs is that Jev gives you the full probabilities (and the confidence, which is something like normalized entropy).

But in general, we might be reading too much into this. If I were to build something like this, a Qwen model would be among the first things I'd reach for too. Initially just prompted inside a little harness to guarantee you get the desired output. Next step would be finetuning, finally training your own foundation model, if you can muster the funding. In this fast-moving space, I think it's quite understandable that they'd go public with an MVP asap, so likely not much training on their own. And even if they're finetuning, Qwen's baked-in answer (through Alibaba's finetuning) seems likely to survive unless it was explicitly overridden.

c7b··on Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
Once the first letter is Q, the rest is probably pretty determined. Can you see the confidence for the first letter (don't want to accept the ToS to follow your link)?
c7b··on Pirate Face Rescues LLM Models from Deletion
Interesting. Is there a paper that explains this in a bit more detail, like [0] for abliteration (underlying the Heretic software, afaik)?

[0] https://arxiv.org/abs/2406.11717

c7b··on Brood War Bench
At last something that feels properly orthogonal to pelicans on bicycles.
c7b··on If math is more than proof, we need to better celebrate the rest of it
I mean, there's centuries' worth of mathematical prose to train on. But that's presumably already in the training data, so if it isn't good enough today, it might not get better fast enough to keep track with how fast they'll get better by training on formally verified math. But then again, the prose in Terry's conversation I linked above seemed pretty useful. But it's also a problem requiring famously little advanced mathematics.
c7b··on Asking authors about their own papers
> Separately, our group has been exploring approaches along these lines to make such evaluations more scalable

Actually, that sounds like an interesting idea for peer review in general, to include an interview between referees and authors. If it saves one round of rebuttals/reactions, it needn't even consume a lot more of everyone's time if you're doing those things properly. What it would undermine would be blindness, but something's gotta give, and it was already on its way out.

c7b··on If math is more than proof, we need to better celebrate the rest of it
> we might imagine what it could look like to have an analog of the Millennium Prize Problems for open exposition problems

The core idea seems to me that we should shift the standards for professional evaluation from generating proofs to generating explanations. Makes sense that such a proposal would come from the 3B1B guy, and I actually agree with it, irrespective of AI. But what eludes me is how that could be a defensive mechanism against AI automating humans out of mathematics. AI is likely no less good at producing natural language explanations as it is at generating rigorous proofs. It's telling that even Terrence Tao turned to AI to understand AI-generated results [0]. It seems that the essay doesn't address that issue at all.

[0] https://news.ycombinator.com/item?id=49010345

c7b··on Steam Frame starts at $1059
Well, I hope you're going to be right. My gut feeling, however, is that putting all the little bits that make up the AVP experience together would already have been too much to ask from other manufacturers in the absence of any patents. Just like how we can't seem to get Mac-quality hardware from other manufacturers even at Apple's (somewhat ridiculous) price points (and I say that as someone who doesn't like Apple for its proprietary practices). So the chances that someone is going to deliver a similar quality product, or arguably we'd require something even better (less bulky), in the hope of some gray-zone playbook to work out, don't seem great.
c7b··on Breaking the 1.58-bit Barrier for Ternary LLMs
> We measure the actual symbol distribution of 29 ternary LLM models and find that zeros account for up to 51.5% of all weights. Motivated by this finding, we introduce BITCOS, a simple distribution-adaptive layout

I honestly assumed that's how they already work. I have to admit that I even explained it like that to a friend. Why on earth wouldn't you design it like that from the start (talking about the adaptive, not the measure part; just sacrifice a few bits to clarify your encoding and save a ton of bits)?

c7b··on Steam Frame starts at $1059
In this case however, many of the patents have a hardware element to them. Gaze+tap is powered by an inside camera to track eye movements and a Lidar sensor at the bottom of the rim for motion detection. Not sure your playbook would work out there.
c7b··on Steam Frame starts at $1059
afaik the AVP has a dedicated Lidar sensor for gesture detection. I'd say Apple has managed to patent more basic stuff than that (pinch-to-zoom, rounded corners,...)
c7b··on Steam Frame starts at $1059
The thing is, the AVP has a ton of quality-of-life features, eg relating to gestures, and most of them locked behind patents. Looks like we'll have to wait 20 years before other devices will be able to match it (or whatever the expiry date is on those patents).
c7b··on Steam Frame starts at $1059
Being able to comfortably sit back / lie down and work is an interesting proposition. The Apple Vision Pro is too bulky and expensive to carry that whole category on its own, they need an 'Android equivalent'. But I think Apple messed this up by creating an absolute patent minefield (some 5k patents for AVP features iirc).
c7b··on A Beginning for Mathematics
With formalized math, you only need to validate the problem statement (in theory, in practice agents have already managed to exploit Lean compiler bugs, but the incidence of those should decrease enough to be practically lusable for 'blind' validation of AI proofs in the foreseeable future).
c7b··on Navier-Stokes Announcement
Even funnier that it wouldn't even be the first time that happens: https://en.wikipedia.org/wiki/Grigori_Perelman
c7b··on No Man's Sky Cosmos
I remember first hearing about NMS' concept and wishing that they would take the general idea but downsize it by at least 10 orders of magnitude. A universe with 1-100 millions of planets would still feel absolutely massive and way more than individual players would be able to exhaust, yet it would allow the creators to do a bit better curation, allow players to actually meet, maybe add more hand-crafted experiences. Basically, give up on technical purity in favor of fun.
c7b··on Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
I don't think that follows from the published results. Would have been an interesting hypothesis to add though, and quite easy. Just throw the same setup at some benchmarks.
c7b··on Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
I wasn't aware that we have access to raw reasoning tokens? I thought what you get is a kind of summary. Does the author have some kind of privileged access or was my assumption wrong?

But for the question studied here it probably doesn't matter - overlaps in the publicly available output may be indicative of distillation (or not), regardless of what it is. I would just find it surprising that the Chinese labs would use it so trustingly. The publicly released reasoning trace is the first place where I would suspect some distillation poisoning to be injected.

c7b··on Corporate America is getting hooked on open-source AI
The article clearly distinguishes between open source and open weights models, and the headline and parts of the article state that it's open source models that are having a moment. But it doesn't list any examples. The only model families explicitly named are open weights only.

Could someone clarify whether there are actual open source models that are competitive with the likes of Gemma (mentioned in the article), or is the headline just wrong?

c7b··on Formalizing Fermat's Last Theorem
If you're doing it for fun anyway, why not use the language that gives you the most pleasure?
Page 1 of 20Next →