HNHacker News
TopNewBestAskShowJobs

rao-v

1,253 karma · joined October 26, 2024

v@inferencing.net
submissionscomments
rao-v··on DeepSeek v4.1 Flash
As I also said on Twitter - it really amazes me how fearless Deepseek are. Every single model release is packed with new and crazy clever ideas and somehow, they always commit to training them at near frontier scale.

I know everybody wants the tell all story of the clever ideas that were developed over the last ~3 years at Anthropic and OpenAI, but what I really want to thumb through is DeepSeek's notebook of "brilliant but didn't quite make the cut" ideas.

They must be trying some truely bonkers stuff to be able to land this much architecture novelty in their full releases.

rao-v··on Training a 3.8B LLM to 0.384 CORE for $998
This is really neat! I'd be tempted to try this again targetting ~1B params and the entire cookbook of "current" small model ideas: gated delta nets (or is ~2K context too short to benefit?), per layer or n-gram embeddings, gated residuals etc.
rao-v··on Factoring RSA 260
Devin … now that’s a name I’ve not hear in a long time
rao-v··on The Navier–Stokes Millennium Prize Problem
It’s reasonable to wonder about what chat usage data gets into models (to be honest probably quite little - carefully curating training data and creating higher quality synth data seems to be the current approach) and the implied risk to privacy and creativity (every new patent filed this year probably touched a model before filing).

What I cannot reconcile is the timeline and the concern in this specific case.

I don’t think training pipelines are anything close to the level of continuous training needed to incorporate Aug 15th ideas into a model that generates a breakthrough early Sept. Either OpenAI nakedly had someone with mathematical understanding dig into a specific user’s chats (a massive red flag) or this really is poor handling of a more classic parallel discovery situation (with one party clearly having worked on it longer)

rao-v··on I tested 10 model/harness combinations on the same Three.js task
For what it’s worth rando redditers warns that cannon.js is old.

I found the remarkable https://babylonjs.com/lite-demos/

Both the regular and lite demos are pretty amazing for browserware!

rao-v··on I tested 10 model/harness combinations on the same Three.js task
Does anyone have an efficient, reasonably designed three.js or other web oriented mini game engine? Low poly but modern rendering effects should be a good niche for hobbyist stuff but hard to find.
rao-v··on I tested 10 model/harness combinations on the same Three.js task
Really appreciate this perspective. It’s tempting to stop at being amazed at what these models can do, but my experience matches yours - they need help with details to do well.
rao-v··on There's a new "Google Jail" for independent wikis
As a bandaid solution we probably just need a good base domain for game wikis attached to a somewhat trustworthy foundation (that won’t sellout to Fandom in a week). Agree that wierdgloop.org is not … great
rao-v··on I tested 10 model/harness combinations on the same Three.js task
It’s really interesting how much little choices make the result better or worse. Astra and one of the GLMs added bright lights, and thus looked so much better to my eye.

Genuinely happy with some of the Qwen 3.8 results (especially since I can run that model at Q8).

Interesting to see how much better (at this task) Pi (OMP) is over Opencode as a harness.

I’d love to see a few more with outcomes that are as easy to judge but less subjective.

I’ve got a toy project going to make a fun to watch battle simulator where an LLM (or two if playing vs) has to write programs that control multiple bots (each with their own line of sight and limited battle context) that have to coordinate and fight alongside each other. Goal is to have the LLM update the code based on current situations maybe 5-10 times in a 5 min simulated battle. Exploring even allow the bots to request new programming and score based on number of reprogram steps.

rao-v··on Formalizing Fermat's Last Theorem
I hear you but point me to one novel formalization right now that is not Lean. It’s really becoming a refacto standard. Which is lovely but terrible for pedegogy
rao-v··on Go grandmaster Shin defeats AI KataGo with a two-stone handicap
Your point is pretty fair.

The reason I think the top 100 pro would win is that when I watch pros review they are very quick to catch when someone is playing a non-current or up to date style. I’ve also seen a few middle game situations where the classic famous go move is now “obviously” not good after we’ve seen how AI handles it.

This all adds up to many ways good sound moves to Lee Changho have now obvious responses. Do I believe he would adapt inhumanly fast? Absolutely! Legendary fighting spirit

rao-v··on Formalizing Fermat's Last Theorem
Right? Might be worth another shot
rao-v··on Formalizing Fermat's Last Theorem
An aside on Lean and it's massive library of results: As someone who's put non trivial effort into slowly learning geometric algebra, lie theory and other slightly advanced math topics, I have to say my brain cannot read Lean. It feels so unprocessable.

I've tried the various intros to Lean multiple times (even before Lean 4 came out) and something about the way Lean proofs are written does not align with how I think about proofs. My very brief attempts at Isabelle / RCoq feel more natural.

I think it's a pity that the future of proofs is Lean. I'd love for someone to come up with a more digestable proof language!

rao-v··on Go grandmaster Shin defeats AI KataGo with a two-stone handicap
In a march 2026 interview David Wu (lightvector, Katago’s creator at Jane Street) noted that he doesn’t have a systematic solution for the cyclic group problem, but adding examples to the training set mostly ensures Katago during MCT rollout figures it out. I don’t think there has been a post mid 2024 verified exploit.

https://gomagic.org/david-wu-on-building-katago/

rao-v··on Go grandmaster Shin defeats AI KataGo with a two-stone handicap
The problem is ambient go knowledge. A top 100 player would easily beat time traveling Lee Changho in his first few matchups. Of course give peak Lee Changho a fortnight to prep with Katago and … well that would be something!
rao-v··on Go grandmaster Shin defeats AI KataGo with a two-stone handicap
Could someone sufficiently motivated invest in training Katago to be able to beat Shin Jinseo with 3 stones of handicap? Unfortunately - probably yes.

This in no way detracts from how absurd and remarkable it is that Shin Jinseo can beat KataGo (it gets a LOT of training and architecture refinements https://katagotraining.org/#eloGraphButtons) with 2 stones of handicap.

rao-v··on Go grandmaster Shin defeats AI KataGo with a two-stone handicap
It’s worth understanding that Shin Jinse has been significantly stronger than his nearest human opponents for a while now, more so than Magnus was even at his very peak.

In go ELO like scoring he’s something like 120 points over the next strongest player. No other player has ever broken a 3800 rating let alone 3850. Ke Jie (the previous long time champion) peaked at 3755. Shin Jinseo’s strength graph is the most absurd straight line.

https://www.goratings.org/en/

2 stones is historically the gap between a 9P ranked and a 1P ranked professional player (very roughly the gap between super grandmasters and an almost grandmaster)

That is to say it’s shocking that Katago (almost certainly significantly stronger than AlphaGo) is a mere 2 stones stronger than Shin Jinseo. I suspect it would be 3-4 stones vs any other human pro.

rao-v··on Dwarf Fortress is getting the mother of all magic updates
DF isn't 2D anymore! Back in the day it was literally dig down and dig left right, no Z axis.
rao-v··on Dwarf Fortress is getting the mother of all magic updates
Nostalgic for when Dwarf Fortress was 2D and had an ending (if you dug too deep a catastrophic monster emerged as I recall)
rao-v··on RAG Is Simpler Than You Think
A frontier LLM based query rewriter has atleast a masters in economics, and a pretty good understanding of finance informal language. How long ago did you try this? I'd be curious if you find this still to be the case.
rao-v··on Queryable Executables
Oh I’m sorry to bring such a presentist idea to such a cool project, but if we ever get LLMs inferencing cheaply on consumer hardware and capable of efficient continuous learning, this insane format might be the perfect way to share your unique tamagotchi of expertise in a specific area
rao-v··on Actually Queryable Executables
This is deranged, and perilously close to dumb, which makes it one of the best things I’ve seen on hacker news this year.

Absolutely wonderful stuff.

rao-v··on I Dream of Quieter Computing
Esp32s with eink displays are pretty much this
rao-v··on One night in Uzbekistan: Why was this one data point so influential?
The article suggests it's unreasonable numbers in the original Uzbekistan data source and that other datapoints may have been worse, the authors just didn't correctly execute their basic checks.

"It turns out that Uzbekistan wasn’t even the biggest outlier, but that the version they had published had the axes cropped so you couldn’t see the outliers..."

rao-v··on Bun 1.4 Rust rewrite is not looking good?
What is the right recommendation at this point for a Node alternative? Deno?
rao-v··on Qwen 3.8 27B is excellent, but it defaults to overthinking things
Even Muse Glimmer (as did GPT-OSS I think) does ~4 sliding window attention layers + 1 full attention layer (like Gemma 4). I’m assuming both labs have good reason to think that gated delta nets are not optimal.

Of course it’s possible the labs just stick with the optimal architecture for large models and GDN is best for smaller models.

rao-v··on Qwen 3.8 27B is excellent, but it defaults to overthinking things
I run llama.cpp and specialized forks on 64GB of HBM and I still cannot figure out where to find the final correct guidance on using the Gemma 4 models.

Would appreciate any kind of pointer to the latest!

rao-v··on Ultraviolet Bird Photography
Read this as "Male humans can't tell the difference." and found it comic
rao-v··on Every Fucking Website (2020)
Ha!
rao-v··on Every Fucking Website (2020)
I salute your honesty! But [he hastens to add, looking furtively around] I also frown disapprovingly at your choices
← PreviousPage 2 of 10Next →