MuZero’s first step from research into the real world
deepmind.com
deepmind.com
Are you also interested in it? I have found some other people interested in it. Maybe we can start a Discord channel for this?
Also how they evaluated user perceived quality doesn’t seem to be elaborated. That itself is an area of active research last I looked.
Running MuZero online sounds like a fairly computationally expensive prospect...
A cloud stack, from OS kernel settings to TCP/IP to database query optimizers to video codec settings to compiler settings, is made of thousands upon thousands of toggleable options, each of which is usually left at the default because no one on earth understands more than a small fraction of them, much less how to set them all appropriately for each task end-to-end. It's blackboxes on top of blackboxes all the way down. Collectively, inferior options could be giving up an incredible amount of performance. As has been demonstrated by experts in performance tuning, depending on how pessimal the defaults are, you could easily gain orders of magnitude performance by setting them to saner settings, much less truly optimal settings - these sorts of posts turn up routinely on HN, and even in very well-tuned cloud stacks, you have to figure that gains like >10% should be possible.
MuZero here shows that it can work for one piece of the stack. And MuZero is, by design, an insanely general architecture: handles two-player games like chess/Go & handles one-player like ALE, handles continuous action spaces (Sampled-MuZero), reasonably sample-efficient (because it learns an environment model, so using that more is MuZero-Reanalyzed), handles hidden information games against adversaries (Player of Games), and now OP shows self-play in a weird setting. (It still requires problem-specific input layers but even that can be lifted if you're willing to pay for Perceiver inputs which do arbitrary input modalities.)
So you can see the potential here for doing much more of cloud operations (beyond current applications like datacenter cooling control) with DRL agents. Plunk down a MuZero on your entire stack and assign it the goal of optimizing end-to-end for each specific task - DRL is expensive, but cloud-scale is even more so. Needless to say, don't expect any released checkpoints on Github...
I think "parameter optimization" is a better expression here. Optimize can means many things, but certainly it's not optimizing the algorithm/encoding itself. It's all about being smart at the very last mile.
The choice of when to have an I-frame or a P-frame is arbitrary, but the rendered video will look the same. However, too many I-frames can bloat the filesize, and too few can degrade the appearance significantly as errors add up.
They act on a codec parameter related to the I-frame, to pick better rules for good compression without visible errors in the P-frame.
p.s. Meh, any wisdom here?
DeepMind claims (with good cause, IMO), that MuZero can be such an algorithm. Showing that this one algorithm can tackle disparate problems is a way of proving this.
I think the questions that still stand are: is it even possible to build computers that could drive a scaled up MuZero to AGI? And is there a more efficient way to get there? I suspect the answer to both questions is yes.
Still, I think it is pretty incredible that we've managed to build computer programs that can totally adapt to arbitrary datasets and perform arbitrary tasks.