HNHacker News
TopNewBestAskShowJobs

rao-v

1,253 karma · joined October 26, 2024

v@inferencing.net
submissionscomments
rao-v··on Gemini 4 Argon
It’s funny that I could tell google was up to something because Gemini chat quality dropped dramatically starting 2ish weeks ago. Agy perf stayed somewhat stable with the odd surprising win (maybe the new model?). I’m a bit sad it was almost impossible to run out of antigravity quota presumably because it was not being used that much).
rao-v··on Show HN: NSL – WSL for Linux
I use mise inside WSL. It's sort of a terminal into "workspace" for me. Should probably spin an orbstsck or whatever to do thisnon my macbook
rao-v··on Show HN: NSL – WSL for Linux
I have to say this is nice packaging. It is weirdly not obvious how nice it is to cleanly use a different “machine” inside your desktop.

I never thought I’d prefer WSL to even my MacBook for working with remote servers and dev but somehow I do.

I of course know the many ways to roll something like this for myself, yes dev containers are better for many things etc. but it’s wierd how good the ergonomics of a WSL like container are.

rao-v··on Show HN: TinyAIArena watch AI agents battle it out
I love this and the aesthetic.

I've been jamming on a sort of Corewars (remember?) / Starcraft hybrid battler where LLMs write sandboxed Lua programs to control bots fighting out 10-100 vs 10-100 tank battles. Every quarter the LLM gets a full view of the situation and can reprogram all the bots to better adapt strategy etc.

It's good fun to watch - excited to share soon.

rao-v··on "As a Language Model": Chat Template Switches LLM Self-Referential Voice
Modern models appear to be much better, at least as proxied by their ability to assess urgency in perhaps a more complex setting: mental health (OpenAI benchmark, so perhaps some skepticism is warranted but the methodology seem reasonable and detailed)

https://openai.com/index/introducing-mentalhealthbench/

rao-v··on Does Georgism work? Five years later
The challenge is that LVT as advocated here pushes for annual updates of the tax based on available information. That easily can move rates up (and down) by “drive people out” amounts in 5-10 years in urban areas.
rao-v··on How I changed teaching after AI managed to do all my homework assignments
Humm doesn’t this perhaps suggest you should do away with the homework, offer optional practice exercises and double down on in-class testing?
rao-v··on Does Georgism work? Five years later
I’ve always been Georgism curious, but didn’t know how to handle some of the more obvious challenges so I really want to thank the author for the work here!

I’m curious about how to think about the dynamics of a LVT:

- I can imagine urbanism creating a flocking behavior that ruins neighborhoods in 5-10 year cycles. Coffee shop draws more commence drawing more affluent people pushing up LVT driving out residents faster than even the spectre of “gentrification”, only to collapse when a nearby area is cheaper for a cool coffee shop to start. (The problem specifically is that you’ve both driven out people and caused an inefficient overbuild in each area)

- Similarly how do you think about new uses for land emerging? If I lease desert land for a data center because it’s so perfect for it, instead of buying it … how (and when) does it show up in LTV increases? What if I trade you some other benefit to keep it out of LVT impacting records?

- Are you just costing the world coordination surplus by forcing high value enterprises to distribute themselves (inefficiently) just far enough apart that they don’t drive up each other’s LVT? That’s a deadweight loss.

I’m a fan of the idea - these are just some of the tricky challenges I don’t have a good answer to yet.

rao-v··on An agent used DNS to reach an external chatbot
Why are we blocking agent access to normal tools without telling them “hey this access is beyond the intended scope of this task”. If I woke up one day and couldn’t reach google.com, I too would start fiddling with tricks to restore access.
rao-v··on Evolving programming languages in the AI era
I’d love for language environments to support encouraging LLMs to specify more when they write code. Why should they first write a complex function or a class and later bolt on a test?

I’d love for a class in this future language to come packaged with tests, invariants, fuzzer parameters, profile targets / performance budgets with realistic inputs (on this 100 element array this should take no more than X clock cycles), race condition stress tests etc.

The compiler (or even linter) should optionally run some / all of these checks and succinctly report back (with knobs so the LLM can manage wall clock time).

Adding each of these should not be follow on steps.

Beyond this, debug hooks should be trivial to set (in code itself), so the LLM can trivially say show me the stack after the 9th time this function is called on this input to the program.

rao-v··on Too AI; Didn't Read
Perhaps:

Did you even read it through carefully once? If you couldn't bother to craft it, why should I read?

rao-v··on The Curious Power of Punctuation
This is a good catch! The choice of beat here to my ear is meant to capture both a sense of urgency (beating a drum to set the pace for rowers) and beating upwind (tacking frequently). The whole article feels a little undercooked and peevishly underinformed for the New Yorker.
rao-v··on LensVLM: Compressing long context as images, expanding only relevant pages
I really like this approach! I sort of think of the vision encoder here as an expensive high fidelity RAG encoder.

The thing I’d love to do with a system like this is train it to be KV cache ordering independent (ie permutation invariant at the page level). Basically each page’s KV cache should be understandable by the model in any ordering - which would allow you to go one step further and treat the KV cache of the vision encoded page as the chunk for the model to reason over.

Then all these zoom in for more detail tricks will extend naturally.

rao-v··on MiMo v2.6
This dashboard is almost certainly built on verl (https://github.com/verl-project/verl), which comes with a bunch of dashboarding capabilities built in (that doesn't look too dissimiliar to these dashboards).
rao-v··on MiMo v2.6
+1 at some point, you need to expect to train a much better base model using everything you've learnt. At the least, you probably want to bring on line the next 10 clever RL environments and ideas your team has been cooking up (which will pipeline into v2.7 etc.)
rao-v··on MiMo v2.6
I might turn this into a blogpost if folks are interested, but my god there is so much clever info in that dashboard.

Here is one really neat bit:

A cutting edge training idea (for agents, it's been used elsewhere for ages) is on-policy RL, basically, it's not enough to say "here is an end to end agentic sequence (including tool calls etc.) that is perfect" you want to say "here is a sequence you might actually have generated that turns out to be correct".

Basically, it's more training efficient to improve models with small tweaks to do more of the right thing they are already doing sometimes than from some perfect oracular "this is the way" answer.

(if you've ever tried to teach humans new skills, you’ve probably noticed this too!)

When you do that, you care about how far the model you are updating (improving) has deviated from the one being used to generate rollouts (agentic rollouts for hard problems can take hours with lots of tool calls, so you can't keep redeploying every slight improvement).

Lo and behold, the dashboard literally has:

partial/avg_staleness (likely the measure of how many micro iterations the "generate answers" model is behind the "improving based on the occasional right answer" model)

train_infer_diff/new_infer/kl (a more direct KL divergence based way of measuring how differently the two models generate tokens)

How cool is that?!

And don't get me started on the clever ideas hiding behind dynsam/avg@n ...

rao-v··on MiMo v2.6
I know we have strong views on what a truly open model is (open weights, open training data, open training code etc.) but I really like how transparent they’ve been about the training of this model.

The realtime dashboard they shared during training (https://mimo.xiaomi.com/rl/) was an incredible learning and teaching tool for me, and they’ve been unusually comprehensive in sharing details about their methodology (check out that tech report - it's got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores (even the stuff they didn’t do well on).

If you’re releasing an open model going forward, please consider offering the community more of this transparency!

rao-v··on I am often wrong
I find myself slightly more hypothesis-driven, so I tend to be more effective putting much of the focus of 2. Gather missing information after 3. Define the problem ... or at least inside the iteration loop.

As others here have noted, it’s very easy to get lost gathering information that isn’t actually helpful in assessing whether the problem is well formed or whether potential solutions are applicable.

The more you force yourself to specify the problem clearly, but treat your initial problem definition as somewhat suspect - potentially missing key dimensions - the more you will be comfortable finding the information to validate / challenge it (or it's implied solutions) and re-shaping both the problem and it's solutions efficiently.

rao-v··on Hister: A private search engine for the pages you visit and the files you keep
I'd love a extension setting to only send tabs that were visible for ~4+ seconds.

I built myself a little extension last year that tracks what information I was looking at, but focused on generating "new info" recaps for the day / week.

I realized that I open / quick view a lot of pages and close them, which is a strong signal that I don't care about that specific page, and it shouldn't be a source of "new insights" that I learnt that day (since I probably don't care about that topic).

I'd love to re-try a simpler version of that project that builds on Hister as a backend actually.

rao-v··on Bend – a language that blocks AI mistakes via proof and runs on GPUs
Hey Victor! Been following you since HVM/Kind, partly because I'm moderately unhappy with the state of out of the box automatic parallelism in modern languages!

Do you plan to invest in profile guided optimization or autotuning in Bend2 - using runtime profiles / cost models to make decisions around SIMD vs. multicore vs. GPU parallelization?

Bend2's model might give you a really nice view into available parallelization. Heck I can imagine integrating an LLM to profile and optimize in an absurdly expensive `-O7` optimization mode one day!

rao-v··on Xiaomi Mimo 2.6 live post-training dashboard
I absolutely love that someone is doing this! Why isn’t IBM for Granite or Google for Gemini?

If you are going to develop a near frontier model, and you don’t think you have special sauce up your sleeve, why not making training runs and RL environment scores etc. visible to the world?

I’m genuinely learning quite a bit just from the dashboard

rao-v··on Backprop Alternative: Augmented Lagrangian Predictive Coding
Turns out this is more compute intensive than regular backprop. The win (if any) would be in training on future highly distributed architectures where neighbouring parts of the model have accesss to very high bandwidth (to each other) but communicating to further away parts is more expensive
rao-v··on Backprop Alternative: Augmented Lagrangian Predictive Coding
I wonder if you could take a traditional backprop trained LLM and apply this approach to finetuning it (presumably needs less memory and compute?). It could be another entry in the spectrum between LORA and full fine tuning.
rao-v··on Benchmark: CadQuery vs. OpenSCAD for agentic CAD work
I tried modifying the provided CadQuery skill to work with build123d with an agent (surprisingly gemini pretty good at this) and it produced very good results (possibly better than the reported CadQuery/OpenSCAD results after taking a little care to avoid leaking the benchmark pass/fail criteria). Worth a try if you are interested in this space.

I like build123d simply because it can export proper STEP (like CadQuery) but has a nicely python friendly design.

rao-v··on A Design Space Exploration of Async/Await
There is better than even odds I'm older than you, so I'd recommend you rethink using phrases like "possibly longer than you've been alive", it's not ... polite regardless of people's age.

The point (and I'd encourage you to find that thread to not retread ground) is that we absolutely can compile most computation heavy code for these different targets reasonably well - what we cannot garentee is that the resulting code is optimal given context. But gosh we can do so much - I’d encourage you to look into in profile guided, target aware, and autotuning optimization etc. (and then of course, there are LLM guided optimizations, but that's a whole other kettle of fish)

rao-v··on A Design Space Exploration of Async/Await
Threads were historically expensive enough that “just spawn a thread” wasn’t a reasonable thing to do in many situations. Thread pools were sort of a last resort, and we ended up with control flow like objects (futures, await etc.) to multiplex concurrency without parallelism.

Go sort of asks why tho and just standardizes on go routines as a good abstraction over both concurrency and parallelism.

This is sorta true elsewhere too. Go rejects a lot of the machinery that OO languages seem to feel obliged to carry around - inheritance hierarchies, explicit interface implementation etc. For what it's worth, I don't write much go, and I don't think it's magical. I just like how clearly it revisited some basics.

Elsewhere on HA you’ll find my extended rant about how strange it is that we don't have a language that elegantly abstracts computation over threads, SIMD, GPUs etc. Compilers can do this sort of thing now, just not optimally.

rao-v··on A Design Space Exploration of Async/Await
I remember being so mad years ago, coming from a pure CS background, when it dawned on me that async await was “mere” control flow and not actual parallelism.

It’s why I feel go (with go routines being the norm) is one of the few imperative languages that was designed vs. filling out a bunch of historical constraints (apologies this is not meant to trigger a language debate, just an idiosyncratic thought)

rao-v··on I spent $220 on Google app ads and 60% of the installs were robots
Isn't this the sort of thing Google is supposed to be doing for us?
rao-v··on Rune is now open source
I like the idea of making it trivial to work across multiple machines but I’d really prefer not to have to trust your coordination server and encryption approach etc.

Could this (optionally) just run over Tailscale (I suppose ssh is always an option)

https://docs.rune.build/learn/network

rao-v··on DeepSeek v4.1 Flash
umm what are you talking about? Basically this crowd (esp. folks like me who run medium models locally) like open stuff and can be a tiny bit unenthused about opaque mysteries handed down from on high. You'll see people delighted with Gemma releases and heck even IBM's Granite models (boring architecturally though they may be) every time they come out. Heck I was chuffed about gpt-oss-120b for weeks. @sama give us another already!
Page 1 of 10Next →