1,253 karma · joined October 26, 2024
I never thought I’d prefer WSL to even my MacBook for working with remote servers and dev but somehow I do.
I of course know the many ways to roll something like this for myself, yes dev containers are better for many things etc. but it’s wierd how good the ergonomics of a WSL like container are.
I've been jamming on a sort of Corewars (remember?) / Starcraft hybrid battler where LLMs write sandboxed Lua programs to control bots fighting out 10-100 vs 10-100 tank battles. Every quarter the LLM gets a full view of the situation and can reprogram all the bots to better adapt strategy etc.
It's good fun to watch - excited to share soon.
I’m curious about how to think about the dynamics of a LVT:
- I can imagine urbanism creating a flocking behavior that ruins neighborhoods in 5-10 year cycles. Coffee shop draws more commence drawing more affluent people pushing up LVT driving out residents faster than even the spectre of “gentrification”, only to collapse when a nearby area is cheaper for a cool coffee shop to start. (The problem specifically is that you’ve both driven out people and caused an inefficient overbuild in each area)
- Similarly how do you think about new uses for land emerging? If I lease desert land for a data center because it’s so perfect for it, instead of buying it … how (and when) does it show up in LTV increases? What if I trade you some other benefit to keep it out of LVT impacting records?
- Are you just costing the world coordination surplus by forcing high value enterprises to distribute themselves (inefficiently) just far enough apart that they don’t drive up each other’s LVT? That’s a deadweight loss.
I’m a fan of the idea - these are just some of the tricky challenges I don’t have a good answer to yet.
I’d love for a class in this future language to come packaged with tests, invariants, fuzzer parameters, profile targets / performance budgets with realistic inputs (on this 100 element array this should take no more than X clock cycles), race condition stress tests etc.
The compiler (or even linter) should optionally run some / all of these checks and succinctly report back (with knobs so the LLM can manage wall clock time).
Adding each of these should not be follow on steps.
Beyond this, debug hooks should be trivial to set (in code itself), so the LLM can trivially say show me the stack after the 9th time this function is called on this input to the program.
Did you even read it through carefully once? If you couldn't bother to craft it, why should I read?
The thing I’d love to do with a system like this is train it to be KV cache ordering independent (ie permutation invariant at the page level). Basically each page’s KV cache should be understandable by the model in any ordering - which would allow you to go one step further and treat the KV cache of the vision encoded page as the chunk for the model to reason over.
Then all these zoom in for more detail tricks will extend naturally.
Here is one really neat bit:
A cutting edge training idea (for agents, it's been used elsewhere for ages) is on-policy RL, basically, it's not enough to say "here is an end to end agentic sequence (including tool calls etc.) that is perfect" you want to say "here is a sequence you might actually have generated that turns out to be correct".
Basically, it's more training efficient to improve models with small tweaks to do more of the right thing they are already doing sometimes than from some perfect oracular "this is the way" answer.
(if you've ever tried to teach humans new skills, you’ve probably noticed this too!)
When you do that, you care about how far the model you are updating (improving) has deviated from the one being used to generate rollouts (agentic rollouts for hard problems can take hours with lots of tool calls, so you can't keep redeploying every slight improvement).
Lo and behold, the dashboard literally has:
partial/avg_staleness (likely the measure of how many micro iterations the "generate answers" model is behind the "improving based on the occasional right answer" model)
train_infer_diff/new_infer/kl (a more direct KL divergence based way of measuring how differently the two models generate tokens)
How cool is that?!
And don't get me started on the clever ideas hiding behind dynsam/avg@n ...
The realtime dashboard they shared during training (https://mimo.xiaomi.com/rl/) was an incredible learning and teaching tool for me, and they’ve been unusually comprehensive in sharing details about their methodology (check out that tech report - it's got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores (even the stuff they didn’t do well on).
If you’re releasing an open model going forward, please consider offering the community more of this transparency!
As others here have noted, it’s very easy to get lost gathering information that isn’t actually helpful in assessing whether the problem is well formed or whether potential solutions are applicable.
The more you force yourself to specify the problem clearly, but treat your initial problem definition as somewhat suspect - potentially missing key dimensions - the more you will be comfortable finding the information to validate / challenge it (or it's implied solutions) and re-shaping both the problem and it's solutions efficiently.
I built myself a little extension last year that tracks what information I was looking at, but focused on generating "new info" recaps for the day / week.
I realized that I open / quick view a lot of pages and close them, which is a strong signal that I don't care about that specific page, and it shouldn't be a source of "new insights" that I learnt that day (since I probably don't care about that topic).
I'd love to re-try a simpler version of that project that builds on Hister as a backend actually.
Do you plan to invest in profile guided optimization or autotuning in Bend2 - using runtime profiles / cost models to make decisions around SIMD vs. multicore vs. GPU parallelization?
Bend2's model might give you a really nice view into available parallelization. Heck I can imagine integrating an LLM to profile and optimize in an absurdly expensive `-O7` optimization mode one day!
If you are going to develop a near frontier model, and you don’t think you have special sauce up your sleeve, why not making training runs and RL environment scores etc. visible to the world?
I’m genuinely learning quite a bit just from the dashboard
I like build123d simply because it can export proper STEP (like CadQuery) but has a nicely python friendly design.
The point (and I'd encourage you to find that thread to not retread ground) is that we absolutely can compile most computation heavy code for these different targets reasonably well - what we cannot garentee is that the resulting code is optimal given context. But gosh we can do so much - I’d encourage you to look into in profile guided, target aware, and autotuning optimization etc. (and then of course, there are LLM guided optimizations, but that's a whole other kettle of fish)
Go sort of asks why tho and just standardizes on go routines as a good abstraction over both concurrency and parallelism.
This is sorta true elsewhere too. Go rejects a lot of the machinery that OO languages seem to feel obliged to carry around - inheritance hierarchies, explicit interface implementation etc. For what it's worth, I don't write much go, and I don't think it's magical. I just like how clearly it revisited some basics.
Elsewhere on HA you’ll find my extended rant about how strange it is that we don't have a language that elegantly abstracts computation over threads, SIMD, GPUs etc. Compilers can do this sort of thing now, just not optimally.
It’s why I feel go (with go routines being the norm) is one of the few imperative languages that was designed vs. filling out a bunch of historical constraints (apologies this is not meant to trigger a language debate, just an idiosyncratic thought)
Could this (optionally) just run over Tailscale (I suppose ssh is always an option)