There's a bunch of other fully open models, including the [Marin](https://marin.community/) series of models out of Stanford and Nvidia regularly releases fully open models.
2,044 karma · joined January 24, 2013
I love receiving emails. If you’re on HN, I would enjoy talking to you. Yes, you.
I'm an AI researcher and (small-time) angel investor.
I’ll happily talk to you about your startup, strategize about how to pitch VCs, or give advice on how to get a job in tech. You are not wasting my time.
There's a bunch of other fully open models, including the [Marin](https://marin.community/) series of models out of Stanford and Nvidia regularly releases fully open models.
I do think that MoEs are clearly the future. I think we will release more MoEs moving forward once we have the tech in place to do so efficiently. For all use cases except local usage, I think that MoEs are clearly superior to dense models.
Where did you try this? On the Ai2 playground?
7B models are mostly useful for local use on consumer GPUs. 32B could be used for a lot of applications. There’s a lot of companies using fine tuned Qwen 3 models that might want to switch to Olmo now that we have released a 32B base model.
1) the company has Nx preferences, for N >1, in which case the company has essentially failed to fundraise or
2) the company sells for less than they raised, which again, is a polite form of failure.
this is basically Inflection 2.0.
you need enough RAM and HBM (GPU RAM) so it’s a constraint on both.
2) It generally tracks pretty well unless the model is gaming the metric (training on the test set, overfit to the specific source of data, etc). The relative rankings will typically match in both.
3) alas, not with the mild winter North America’s having. They only stop below -5C or so. I am lucky though. The woodpecker stopped attacking my house and started attacking my neighbor’s. Even worse, it used to be a downy woodpecker,and it’s now been replaced by a pileated one (think: Woody).
There’s a bunch of recent work that quantizes the activations as well, like fp8-LM. I think that this will come. Quantization support in PyTorch is pretty experimental right now, so I think we’ll see a lot of improvements as it gets better support.
The KV cache piece is tied to the activations imo- once those start getting quantized effectively, the KV cache will follow.
In my experience, PUCT does a lot better than UCT, so you want to also have a prior network.
You don’t have to train a new network, but in my experience, it works much better. I haven’t spent a ton of time using off the shelf networks with MCTS though. Maybe it works great.
very subtle bugs is the MCTS experience. Particularly once parallelism is involved.
I worked at DeepMind on projects that used MCTS. Even with access to the AlphaZero source code, it was very difficult to write an other implementation that got the same results as the original.