(Actually, there's a surprising dearth of reinforcement learning in general. Very few blog posts or demos or introductory materials. It makes it hard to understand what is new about DQN or how the whole system works on a concrete coding level.)
(Actually, there's a surprising dearth of reinforcement learning in general. Very few blog posts or demos or introductory materials. It makes it hard to understand what is new about DQN or how the whole system works on a concrete coding level.)
The best paper award at CVPR this year (http://www.cv-foundation.org/openaccess/content_cvpr_2015/ht...) has an architecture somewhat like this: instead of a web browser they're using a 3D rendering engine, and searching for pose parameters that cause their rendered image to match an observed image. This is just Bayesian inference using MCMC, where they train a deep net to function as a data-driven proposal distribution.
If your goal were to actually get a working system, you'd probably want to do inference directly on a parameter vector encoding all the relevant quantities - heights, widths, font sizes, colors, etc. of the various boxes, etc. - that you could programmatically ground out into a CSS file. Trying to do inference over the raw text is making things artificially hard since you have to put so much work into even just getting correct syntax. Though maybe that's part of the fun. :-)
This would create a simple way of generating CSS styles for a document, without dealing with the complicated issues of RNNs having limited memory and producing correct syntax.
Then you can use these predictions as a prior probability over what the CSS styles should be. Then you can use some kind of bayesian optimization to find the optimal settings in the least number of experiments.
sum( abs(X_expected - X_actual) *
abs(Width_expected - Width_actual) +
abs(Y_expected - Y_actual) *
abs(Height_expected - Height_actual) )
mapped over all elements ought to do the trick. When you hit 0, it's a perfect reproduction.I thought about trying to do MCMC over a beam search through the rnn output, but ran out of time and patience.