Image GPT
openai.com
openai.com
Oddly, it still uses TensorFlow like the original GPT-2 release despite OpenAI's declared switch to PyTorch, and it has dependency hell so it's not easy to create a wrapper tool for it.
Since it's still the GPT-2 architecture, it might be possible to port the weights to Huggingface Transformers (for the RGB generation), and then write a wrapper to extend it for the image rendering. (filed an issue here: https://github.com/huggingface/transformers/issues/5088 )
Looking at the samples I am amazed by some of the results.
Note the footnote in the paper where evaluations was interrupted by the move to MS Azure. This is relatively old work, and since it's literally GPT-2 but images, it's no surprise they didn't bother to rewrite it in PyTorch. If their next big from-scratch thing (presumably the multimodal GPT-like one-model-to-rule-them-all research I've been looking forward to ever since the TR article) is TensorFlow, then I'll be surprised.
Of course, I would be delighted to be mistaken.
EDIT: v2 of the paper says 2048 TPUs.
I noted the TensorFlow usage because the original GPT-2 release was TensorFlow 1.X, which led to issues when TensorFlow 2.0 was released soon after.
For model training, I strongly recommend using the higher-level Keras APIs for TensorFlow/pytorch-lightning API for PyTorch than the respective base tools.
TF2 is much nicer than TF1 but most of the changes were just to make it more like PyTorch.
For context, this is equivalent to 100 $10,000 GPUs running for 25 days, 24/7.
https://twitter.com/jm_alexia/status/1273349716915470340?s=2...
This is truly a brute force approach.
There is only so much about objects and the world you can learn from 64px thumbnails, and they mention that this is probably a reason iGPT wins on the tiny images like CIFAR but loses to the semi-supervised CNNs on ImageNet: because the CNNs are compute-efficient enough that they can train at the standard 224px and see all of the details and structure that disappears at iGPT's 64px, and learn end-to-end.
$10000 * 2500days / 5yr = $15000 hardware cost
200W * 2500day * (0.10 USD / Whr) = $1.2M in electricity
Roughly, the boxes would be turned on anyway. Might as well put them to work.
And yeah, it draws more power when in use (by a lot). But the data center probably isn’t paying a huge premium on top of what they would already pay for electricity.
So all that’s left is feeling generally queasy about using so much electricity, rather than a practical feeling of “this is bad because it costs a lot.”
Economies of scale have some dramatic effects at the high ends: I doubt the datacenter’s costs are anywhere near linear for the amount of Whr they’re consuming. And even if they are, it’s still a rounding error considering the profitability of GoogSoftBook.
tl;dr: not coal plants :).
I'm pretty sure we're going to close that plant soon though.
Presumably, since:
"Our data center in Eemshaven was the first to be powered by 100% renewable energy from day one."
https://www.google.com/about/datacenters/locations/eemshaven...
https://www.blog.google/around-the-globe/google-europe/dutch...
And they seem pretty serious about removing non-renewables altogether.
https://storage.googleapis.com/gweb-sustainability.appspot.c...
Residential electricity prices in California are more like 0.19 USD/kWh, which is 0.00019 USD/Wh.
Good use of units makes it easy to spot and fix the mistake.
Fwiw, I think 2500 "V100-days" here is an extrapolation of ~1 2048-TPUv3 day. So approximately $8/hr * 24 hours * 2048 => $400k, if you know, you could make use of that TPU Pod for much of the rest of the days of the year :).
But yes, this is still the realm of "Do you have millions of dollars of ML infrastructure" (if you want it quickly). I'm kind of hoping gwern et al. will use Colab to slowly train a variant of this for free :).
function ConnectButton(){document.querySelector("#top-toolbar > colab-connect-button").shadowRoot.querySelector("#connect").click(); console.log("Connect pushed"); } setInterval(ConnectButton,60000);
:) You still get booted every 6 hour or so.(We actually can create a TPU-2048 pod. For a short while, anyway, before it preempts. But we wouldn't need to if we used lucidrain's efficient attention implementations I mention in my other comment.)
> motivated by early color display palettes, we create our own 9-bit color palette to represent pixels. Using this palette yields an input sequence length 3 times shorter than the standard (R, G, B) palette, while still encoding color faithfully.
Makes total sense now. In fact, those images remind me so much of how photos looked on the early internet, since many were palletized.
Even GPT-3 seems to have trouble with world-modeling, it writes convincing text that has all the signs and form of good prose, but the output repeatedly violates physics and common sense in funny ways.
I know just enough about machine learning to have dangerously unrealistic expectations, but I'd like if I could reasonnably hope to see signs of a shared representation or shared knowledge between say image labeling and language modeling. This looks like a very concrete data point to take if you care about generality. Maybe then we can seriously talk about world-modeling.
Could a different NN be used for upscaling?
Far more interesting would be general training optimizations that could train GPT faster on any kind of task.
However, if you were starting this research today, you'd use any of half-a-dozen different self-attention variants which are roughly linear, including OA's own Sparse Transformers (which they did use to generate images, just on a far smaller scale which wouldn't be adequate to show competitive performance with SimCLR etc). With those, it's perfectly possible to do self-attention over whole images at 256px or higher.
(As a matter of fact, Aydao has been working on a StyleGAN which just uses self-attention every layer instead of convolutions, using one of the new attentions; you can see some generated image samples from Flowers here: https://github.com/tensorfork/tensorfork/issues/31 GAN loss, not autoregressive pixel likelihood, but it makes the point.)
Another example: Unicoder-VL https://arxiv.org/pdf/1908.06066.pdf
GPT2 definitely looks better.
It reflects the evolution of NLP models: (RNN) —> (RNN + attention) —> (attention) and the idea that you don’t need the recurrent component. Just look at all elements of the sequence at once, applying varying weights (attention).
Impressive stuff! This performs well even without domain-specific architecture choices.
Transformers are basically learning relations between pairs of input tokens, moving the problem to a more abstract level than predicting directly on tokens. While CNNs excel at benefiting from those two forms of invariance, transformers have permutation invariance, they can predict on sets, graphs and non-euclidean spaces.
Right. The attention layers even learn attention patterns which look like convolution layer kernels! But better, presumably: https://arxiv.org/abs/1911.03584
I think your definition of "understanding" is more the realm of AI, and why so many people in my field sigh when the term "AI" is used when we're really just talking about ML. But then again, you could absolutely in the current day train an ML model to look at images and then produce a natural language "explanation" of the image. It might not be able to make leaps in deductive logic, but it would be able to explain more than just a list of what is in the image. Is this "understanding"? Maybe that question is philosophy.
This work seems so unconstrained in its use of computation, that is almost screams to me that they must be going about it the wrong way.
> One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way are search and learning.
> The second general point to be learned from the bitter lesson is that the actual contents of minds are tremendously, irredeemably complex; we should stop trying to find simple ways to think about the contents of minds, such as simple ways to think about space, objects, multiple agents, or symmetries.
[1]: http://www.incompleteideas.net/IncIdeas/BitterLesson.html
The idea that specialization is not as powerful as computation fails the most basic test of a proactive, rather than retroactive, theory. Can you make proactive claims about what works in any given domain? Is the solution to take the hungriest algorithm and apply it? What about feature engineering, cleaning, parameter tuning, analysis, etc.? Is the most power hungry solution still the most effective? In my opinion, part of the reason humans aren’t just giant computation blobs is that we thrive on constraints (physical, sexual, emotional).