57 karma · joined June 19, 2018
Absolutely not. Id bet a lot of money this could be solved with a decent amount of RL compute. None of the stated problems are actually issues with LLMs after on policy training is performed.
o4-mini gets much closer (but I'm pretty sure it fumbles at the last moment): https://chatgpt.com/share/680031fb-2bd0-8013-87ac-941fa91cea...
We're pretty bad at model naming and communicating capabilities (in our defense, it's hard!), but o4-mini is actually a _considerably_ better vision model than o3, despite the benchmarks. Similar to how o3-mini-high was a much better coding model than o1. I would recommend using o4-mini-high over o3 for any task involving vision.
Of course architecture matters in this regard lol. Comparing a CNN to a transformer is like comparing two children brought up in the same household but one has a severe disability.
What I meant in this blog post was that given two NNs which have the same basic components that are sufficiently large and trained long enough on the same dataset, the "behavior" of the resulting models is often shockingly similar. "Behavior" here means the typical (mean, heh) responses you get from the model. This is a function of your dataset distribution.
:edit: Perhaps it'd be best to give a specific example: Lets say you train two pairs of networks: (1) A Mamba SSM and a Transformer on the Pile. (2) Two transformers, one trained on the Pile, the other trained on Reddit comments. All are trained to the same MMLU performance.
I'd put big money that the average responses you get when sampling from the models in (1) are nearly identical, whereas the two models in (2) will be quite different.
You seem to be under the mistaken belief that: 1. Google has competent high-level organization that effectively sets and pursues long term goals. 2. There is some advantage to developing a highly capable LLM but not releasing it.
(2) could be the case if Google had built an extremely large model which was too expensive to deploy. Having been privy to what they had been working on up until mid-2022 and knowing how much work, compute and planning goes into extremely large models, this would very much surprise me.
Note: I did not have much visibility into what deepmind was up to. Maybe they had something.
The only reason convolutions are still used in modern (intelligently designed) ML systems is because it is not known how to build a sparse attention algorithm that achieves 2D and 3D locality and is also compatible with modern accelerators. Swin is an attempt at that, but it is something of a hack.
Think about it: you live your life. You experience things. You experience art, and experience emotions or have interactions with other humans grounded in that art. You form connections with certain styles or techniques.
If you then turn around to create art, you form in your mind a general idea of what you want to create. You then draw on your past experiences to actually create the physical art. What process other than statistical extraction from your mind could it come from?
For sure I believe there are things that we don't understand about the human mind. I think the impact of drug use on art creation is very interesting, for example. It indicates that random chemical processes in our brains can play a large determining role in the actions we take (and in this case, the things that we create).
But to say that humans do not use some sort of inbaked statistical world model in the creative process seems wrong to me.
DeepMind claims (with good cause, IMO), that MuZero can be such an algorithm. Showing that this one algorithm can tackle disparate problems is a way of proving this.
I think the questions that still stand are: is it even possible to build computers that could drive a scaled up MuZero to AGI? And is there a more efficient way to get there? I suspect the answer to both questions is yes.
Still, I think it is pretty incredible that we've managed to build computer programs that can totally adapt to arbitrary datasets and perform arbitrary tasks.
I am not an economist, and I'd love to be proven wrong. I just don't see how this ends in a good way for the economy.
Pests are one such reason. They are extremely difficult to control outdoors, but far simpler to do so inside. Both fertilizer and pesticide treatments (if necessary) can be far more specific - e.g. less wasteful and environmentally damaging) when done indoors. Similarly, inclement weather generally does not affect indoor farms.
There's also my favorite argument: if we're ever going to try to colonize space or other planets, we better be damned good at growing plants artificially. IMO every dollar spent improving this space gets us one step closer to unlinking our future from Earth's.