Right now, sitting under my desk is a RTX 2080 Ti GPU which cost around $1000, weighs 3 pounds, draws a maximum of 250 watts, and has a peak speed of 13.4 TFLOPS [1].
We truly live in amazing times.
[1] Not quite a fair comparison: the GPU is using 32-bit floating-point, while ASCI White used 64-bit. But for many applications, the precision difference doesn't matter.
x87 is 80-bit. :)
I hope in 20 years I'll have the equivalent of 160 TPUs under my desk. Hopefully less.
The reason it's not fast enough is that ... there's so much you can do! People don't know. You can't really know until you have access to such a vast amount of horsepower, and can apply it to whatever you want. You might think "What could I possibly use it for?" but there are so many things.
The most important thing you can use it for is fun, and intellectual gratification. You can train ML models just to see what they do. And as AI Dungeon shows, sometimes you win the lottery.
I can't wait for the future. It's going to be so cool.
Seriously though I feel like most of the gains in hardware have been wasted by shittier software both in terms of quality and in the way the software itself acts against the interests of its users.
I thought TPUs are harder to work with because they only support Tensorflow rather than Tensorflow and other high-level frameworks as well as low-level CUDA that are supported by GPUs
TPUs aren't necessarily easier to use – it's about the same – but they're powerful. I've documented some benchmarks in this tweet chain, where I trained GPT-2 1.5B to play chess using a technique called swarm training: https://twitter.com/theshawwn/status/1214013710173425665
The power turned out to be from the fact that every TPU gives you 8 cores at your disposal. I never use the Estimator API. I just scope Tensorflow operations to specific TPU cores. Works great.
In terms of actual performance, I was delighted to discover that TPUs can be faster than GPUs when you use all 8 cores: https://twitter.com/theshawwn/status/1196593451174891520 (solution notebook: https://twitter.com/theshawwn/status/1205914446918492170)
It also gives you flexibility. TPUv2-8 can apparently allocate up to 300GB (!) if you don't scope any operations to any cores. Meaning, you run it in a mode where you only get 1 core of performance, but you get 300GB of flexibility. And then you can connect multiple TPUs together as described in the tweet chain, which quickly makes up the difference.
There is also the question of cost savings. A TPUv3-8 seems about as expensive as a V100. Which one is worth it? Well, it depends. In my experience a GPU is easier to use and quicker to set up if you only need one GPU of horsepower. But suppose you wanted to train a massive model in 24 hours. What's your best option? For us, it was TPUs.
The reason is subtle: It's hard to find any single VM that can talk to 140 GPUs simultaneously. But you can talk to 140 TPUs from a single VM no problem. And since you get 800MB/s to and from the VM, you can average the parameters across all TPUs very quickly.
This is similar to what TPU pods do internally. And while TPU pods are impressive, they are also impressively expensive. A TPUv3 pod will run you $192/hr at evaluation prices. Whereas you can play with a TPUv3-8 for $2.50/hr. You can also play with a TPUv2-8 for free using Colab: https://github.com/shawwn/colab-tricks
Yesterday I used that notebook to port forward Colab's free TPUv2-8 using ngrok, then trained using the new StyleGAN 2 codebase: https://twitter.com/theshawwn/status/1214245145664802817
I think a swarm of TPUs can cost significantly less than a cluster of V100s with less engineering effort.
That said, right now most codebases are designed to work with V100's. It will take time before TPUs widely proliferate. But speaking as someone who was once skeptical of TPUs and who has spent several months trying to discover their secrets, I feel that TPUs can get the job done quicker and easier than a GPU cluster. The hardware is also more accessible, since you can more easily spin up 100 TPUs than 100 V100s. But mainly I like that it's all coordinated from a single machine. It's conceptually simpler to debug and to implement.
If you run into any issues or have any trouble with TPUs, please feel free to ask here or DM me. I love talking about this stuff.
EDIT: In regards to usability, the new Jax library works with TPUs out of the box. Google seems to be heading in the direction of Jax. My initial reaction was "Not another library..." but first impressions were positive. It's not quite the React of ML – an idea which I hope to see soon – but it does seem easier for certain research purposes.
PyTorch also recently gained TPU support, and as far as I know they've put in some serious efforts to make sure things run quickly. As for how you use all 8 cores of a TPU using PyTorch, I haven't looked into it yet. But I'd be surprised if you couldn't. It seems unlikely that they would design an API that would hamstring you to just 1 out of 8 cores.
We had a similar situation at one point. The problem turned out to be that our CPU wasn't generating input data fast enough. So the first step is to confirm that your input pipeline isn't the issue.
The next step would be to break down the problem: Can you extract the smallest part of the codebase into a separate program, and try to make that run under full load?
That's not the technique I used, though. To figure out the multicore stuff, the trick for me was to comment out almost all of the code, until you're left with only a small part that actually runs on the device. Ideally the smallest part.
Basically, change your code so that the model file returns tf.no_op() (or as close to that as possible while still letting your input pipeline run). You want to be in a situation where your training loop is doing an equivalent of while(true) { read_input(); } so that you can verify that your pipeline is able to peg your GPUs to 100% usage.
If you get 100% usage, fantastic! That means you're left with an easy problem: start turning parts of the code back on until you find which part is reducing your performance. Then study that part to figure out why.
If you're not at 100% usage, you're either running into a fundamental limitation (which sometimes happens) or the pipeline isn't designed correctly in some way. I would compare it against other popular codebases such as StyleGAN 2 https://github.com/NVlabs/stylegan2 which is designed to use 8 V100s. The optimizer.py file is pretty insightful: https://github.com/NVlabs/stylegan2/blob/eecd09cc8a067e09e12...
Finally, my biggest tip would be to step back from the problem and think: is there something simple you can do to reframe the problem? When I find myself in a situation where I'm spending a lot of time and energy trying to get a certain thing to work, I can sometimes do X instead for 80% of the benefit. Try to find something like that in this case.
FWIW the TPU profiler was the first tool I reached for. I never got it working. The bag of tricks above ended up giving me effective results on a variety of codebases with no profiler. (A usage graph is pretty crucial, though, which Colab TPUs don't provide.)
So there are a bunch of general tips for solving weird bottlenecks blindfolded.
To answer your question directly:
I am assuming the profiler from Tensorflow works with TPUs in the same way it does with GPUs?
Not really. You're supposed to use cloud_tpu_profiler: https://cloud.google.com/tpu/docs/cloud-tpu-tools
But yeah, if you give specifics (ideally a link to a codebase + dataset + script that runs it) then I can try to look for candidates of what might be the bottleneck.
I looked to use some for solving PDEs, but Google had literally zero documentation on how to cross-compile C to TPUs, launch kernels, etc.
AFAICT, you either use tensorflow or some other product that supports them, and for which the TPU code is not open source, or you can't use TPUs at all.
The ODE solver might be close to what you want.
I use Tensorflow 1.15. The world has been steadily pushing for Tensorflow 2.0 or Jax, but I like the simplicity of the Session model. It's so simple you can explain it in one sentence: it's an object that runs commands. Tell the session to connect to the TPU, and it will run all those commands on the TPU.
Jax is new to me (and to everyone; they just released it). But it looks like Google is pouring some serious R&D into it.
Two things help a lot. One, twitter. You can get a direct line to the people who actually make these beasts. Exploit it when you can. Like you, I dislike using a black box, and I'm intensely interested in the details of how to communicate with a TPU at a low level. I recently asked someone on the jax team about it here: https://twitter.com/theshawwn/status/1213221594052599808
Two, TFRC support has been incredibly helpful. https://www.tensorflow.org/tfrc I don't know who they have working the support channels, but those guys and gals are some of the most helpful and cheerful people I've come across. I often asked them very technical questions and to my surprise, they followed up with an A+ response almost every time, usually the next day.
Pytorch is giving TF a real run for its money, and to be honest I once felt it was a mistake to invest so much time into Tensorflow. But it turned out to be a big advantage due to Google's investment in the overall ecosystem. TPUs are something that only Google has the resources to pull off.
Note that the traditional path towards "just get a TPU up and running and start playing with it" is to use one of their Colab notebooks on the topic. https://cloud.google.com/tpu/docs/colabs I've been implicitly steering you away from these because you seem (like me) to want to know more of the low-level details. Those notebooks are designed to let ML researchers get results quickly, not for hardware enthusiasts to exploit heavy metal. The jax notebooks felt much more satisfying in that regard.
Can u mention how much human Dev time is involved?
We have a stupid-basic single machine Deep reinforcement Self play setup. It takes about 24 hrs to run a full experiment. The NN is the bottle neck. Using Tensor flow. Nothing fancy.
How much dev time for a good enginner (backend, kernel, multi core experience) to get this down to say 1hr ?
Obviously a very general question. Thanks for any input.
The personal ML bots will be a big things. Next step in total automation.
Something I've often wondered, and there are probably good reasons why, is that billionaire tech moguls - even the ones who are outwardly technical (or were in the past - people like Bill Gates, who we know had technical chops in the past) - that none of them (that I'm aware of) haven't ever tried to build "their ultimate computer".
For instance, if I had their kind of money, I've often thought that I would construct a datacenter (or maybe multiple datacenters, networked together) filled with NVidia GPU/TPU/whatever hardware (the best of the best they could sell me) - purely for use as my "personal computer". Completely non-public, non-commercial - just a datacenter I would own with racks filled to the brim with the best computing tech I could stuff into them (on a side note, I've also pondered the idea of such a personal datacenter, but filled with D-Wave quantum computing machines or the like).
What could you do with such a system?
Obviously anything massive parallelism could be useful for - the usual simulation, machine learning, etc; but could you make any breakthroughs with it - assuming you had the knowledge to do such work?
Which is probably why none have done it - at least as a personal thing.
I mean, sure, I would bet that people who own large swathes of machines in a datacenter, or those who outright own datacenter (like Google or Amazon) - their founders and likely internal people do run massively parallel experiments or whatnot on a regular basis, ad-hoc, and "free" - but it's a commercial thing, and other stuff is also running on those machines...
But a single person is probably unlikely to have or think of problems that would require such a grand scale before they would just "start a company to do it" or something similar; because in the end, just to maintain and administer everything in such a datacenter, if one were built, would require (I would think) the resources of a large company.
Of course, then I wonder if such companies - especially ones like Google and Amazon, which own and run many datacenters around the world, and also sell the resources of them for compute purposes - weren't started in some fashion (even if only in the back of their heads) by their founders with that idea or goal in mind (that is, to be able to own and use on their whim "the world's largest amount of computing power"...?
https://www.pcworld.com/article/3313424/inside-seattle-livin...
If anyone were to own a secret HPC cluster, it'd probably be a finance billionaire. Or the owner of a think-tank who made their money as a subcontractor for state intelligence agencies.
«A single GPU card like the AMD Radeon MI60 has more computing power than year 2000 supercomputer ASCI Red (fastest supercomputer in the TOP500 list of June 2000):
• MI60: 7.4 TFLOPS (FP64)
• ASCI Red: 3.2 TFLOPS (FP64) »
https://mobile.twitter.com/zorinaq/status/112491212518746521...
[1] https://en.wikipedia.org/wiki/Thread_block_(CUDA_programming...
A few years before that, at another univ, they put an ancient IBM mainframe in place with a crane, temporarily removing the roof of the building.
According to Wikipedia it was launched in 1997 so it does line up. If I remember correctly, the system was bought from Cray after SGI bought the rest, so I didn't know they even had it at Sun in 1995.
Also, the original model of the E10k supported 64 GB RAM.
On one beautiful day during the weekend, only a two weeks after delivery, our system admin noticed that the machine went offline. Logging in remotely didn’t work at all. No ping either.
He drove to work and ... the machine was gone.
Thieves had used a crane to lift it out of the building through a window onto a truck.
Sun told him that this wasn’t the first time such a thing had happened and that somewhere in the chain from order to delivery, an insider tipped off the thieves about where to find the latest.
[1] GPUs are another kettle of fish of course, but have their own wellknown problems that prevent widespread use outside graphics
I am tempted to build something similar but AMD's wishy-washy ECC guarantees as well as Linux-specific issues make me unsure.
ASUS PRIME TRX40-Pro motherboard
8 sticks Kingston KSM26ED8/16ME (16GB, 2666Mhz DDR4, ECC)
Dual boot Windows 10 / Ubuntu 18.04
2x Samsung EVO 970 NvME M.2 1TB SSDs in RAID0 configuration.
Running in a Coolermaster Cosmos case with a stupidly big air cooler at the moment, to be replaced with a decent liquid cooler (it works, the Cosmos is a huge case because it had to be to hold a hacked Supermicro dual Xeon server board in it before (I really wanted ECC for my workstation))
nVidia 1080Ti+ GPU.
The ECC is detected and claims it is working although I've yet to see it correct an SBE. I haven't been running non-stop memory tests either though so.
I may end up removing the Linux partition since WSL2 works so well on this box.
BTW, Century Micro has the only unbuffered ECC modules at 3200MHz native speed (at least that was the case on the 39x0X release day). I don't know if they can be sourced outside Japan, though.
Fun story, I actually botched my order and got 2666MHz ones... but on closer inspection, it turned out the chips on the modules were actually native 3200MHz ones. With the SPD EEPROM saying they are 2666MHz. So I ended up overclocking them at their actual native speed. And I tweaked the timings to be a little shorter than what the 3200MHz modules were advertized for.
BTW I seen 64GB DDR4 sticks on amazon...
Just got my 3950X w/ 64gb ram, not sure that I'd be able to practically use any more compute than this for what I play with, which is mostly multiple back-ends and some container orchestration for local dev and occasionally video re-encodes (BR-Ripping for NAS).
Some think $4k for this CPU is too much... considering the shear performance that you can get these days for under $10K there's never been a better time to build or buy a computer. My only regret is wasting time and money on aRGB that I cannot configure in Linux.
I do wish for exascale in the medical field though.
I even-more-wish that the energy world can see ~similar improvements in efficiency.
The amount of processing power we each carry in our pockets (even the cheapest throw-away smart phones) would have been almost unthinkable 30 years ago; it's akin to the difference of an Altair of the 1970s vs what was available just 10-20 years prior. What took up a room now sat on a desk and could be purchased for the price of a car.
Now, what took up a room now sits in your pocket, and almost could be given away in a box of cereal its so cheap.
Heck - think about what's available in the embedded computing realm for pennies (or just a few dollars in single quantities) - it's mind boggling to an extent.