AMD’s 64-Core Threadripper 3990X, only $3990 Coming February 7th
anandtech.com
anandtech.com
A few years before that, at another univ, they put an ancient IBM mainframe in place with a crane, temporarily removing the roof of the building.
Right now, sitting under my desk is a RTX 2080 Ti GPU which cost around $1000, weighs 3 pounds, draws a maximum of 250 watts, and has a peak speed of 13.4 TFLOPS [1].
We truly live in amazing times.
[1] Not quite a fair comparison: the GPU is using 32-bit floating-point, while ASCI White used 64-bit. But for many applications, the precision difference doesn't matter.
x87 is 80-bit. :)
I hope in 20 years I'll have the equivalent of 160 TPUs under my desk. Hopefully less.
The reason it's not fast enough is that ... there's so much you can do! People don't know. You can't really know until you have access to such a vast amount of horsepower, and can apply it to whatever you want. You might think "What could I possibly use it for?" but there are so many things.
The most important thing you can use it for is fun, and intellectual gratification. You can train ML models just to see what they do. And as AI Dungeon shows, sometimes you win the lottery.
I can't wait for the future. It's going to be so cool.
Seriously though I feel like most of the gains in hardware have been wasted by shittier software both in terms of quality and in the way the software itself acts against the interests of its users.
I thought TPUs are harder to work with because they only support Tensorflow rather than Tensorflow and other high-level frameworks as well as low-level CUDA that are supported by GPUs
TPUs aren't necessarily easier to use – it's about the same – but they're powerful. I've documented some benchmarks in this tweet chain, where I trained GPT-2 1.5B to play chess using a technique called swarm training: https://twitter.com/theshawwn/status/1214013710173425665
The power turned out to be from the fact that every TPU gives you 8 cores at your disposal. I never use the Estimator API. I just scope Tensorflow operations to specific TPU cores. Works great.
In terms of actual performance, I was delighted to discover that TPUs can be faster than GPUs when you use all 8 cores: https://twitter.com/theshawwn/status/1196593451174891520 (solution notebook: https://twitter.com/theshawwn/status/1205914446918492170)
It also gives you flexibility. TPUv2-8 can apparently allocate up to 300GB (!) if you don't scope any operations to any cores. Meaning, you run it in a mode where you only get 1 core of performance, but you get 300GB of flexibility. And then you can connect multiple TPUs together as described in the tweet chain, which quickly makes up the difference.
There is also the question of cost savings. A TPUv3-8 seems about as expensive as a V100. Which one is worth it? Well, it depends. In my experience a GPU is easier to use and quicker to set up if you only need one GPU of horsepower. But suppose you wanted to train a massive model in 24 hours. What's your best option? For us, it was TPUs.
The reason is subtle: It's hard to find any single VM that can talk to 140 GPUs simultaneously. But you can talk to 140 TPUs from a single VM no problem. And since you get 800MB/s to and from the VM, you can average the parameters across all TPUs very quickly.
This is similar to what TPU pods do internally. And while TPU pods are impressive, they are also impressively expensive. A TPUv3 pod will run you $192/hr at evaluation prices. Whereas you can play with a TPUv3-8 for $2.50/hr. You can also play with a TPUv2-8 for free using Colab: https://github.com/shawwn/colab-tricks
Yesterday I used that notebook to port forward Colab's free TPUv2-8 using ngrok, then trained using the new StyleGAN 2 codebase: https://twitter.com/theshawwn/status/1214245145664802817
I think a swarm of TPUs can cost significantly less than a cluster of V100s with less engineering effort.
That said, right now most codebases are designed to work with V100's. It will take time before TPUs widely proliferate. But speaking as someone who was once skeptical of TPUs and who has spent several months trying to discover their secrets, I feel that TPUs can get the job done quicker and easier than a GPU cluster. The hardware is also more accessible, since you can more easily spin up 100 TPUs than 100 V100s. But mainly I like that it's all coordinated from a single machine. It's conceptually simpler to debug and to implement.
If you run into any issues or have any trouble with TPUs, please feel free to ask here or DM me. I love talking about this stuff.
EDIT: In regards to usability, the new Jax library works with TPUs out of the box. Google seems to be heading in the direction of Jax. My initial reaction was "Not another library..." but first impressions were positive. It's not quite the React of ML – an idea which I hope to see soon – but it does seem easier for certain research purposes.
PyTorch also recently gained TPU support, and as far as I know they've put in some serious efforts to make sure things run quickly. As for how you use all 8 cores of a TPU using PyTorch, I haven't looked into it yet. But I'd be surprised if you couldn't. It seems unlikely that they would design an API that would hamstring you to just 1 out of 8 cores.
We had a similar situation at one point. The problem turned out to be that our CPU wasn't generating input data fast enough. So the first step is to confirm that your input pipeline isn't the issue.
The next step would be to break down the problem: Can you extract the smallest part of the codebase into a separate program, and try to make that run under full load?
That's not the technique I used, though. To figure out the multicore stuff, the trick for me was to comment out almost all of the code, until you're left with only a small part that actually runs on the device. Ideally the smallest part.
Basically, change your code so that the model file returns tf.no_op() (or as close to that as possible while still letting your input pipeline run). You want to be in a situation where your training loop is doing an equivalent of while(true) { read_input(); } so that you can verify that your pipeline is able to peg your GPUs to 100% usage.
If you get 100% usage, fantastic! That means you're left with an easy problem: start turning parts of the code back on until you find which part is reducing your performance. Then study that part to figure out why.
If you're not at 100% usage, you're either running into a fundamental limitation (which sometimes happens) or the pipeline isn't designed correctly in some way. I would compare it against other popular codebases such as StyleGAN 2 https://github.com/NVlabs/stylegan2 which is designed to use 8 V100s. The optimizer.py file is pretty insightful: https://github.com/NVlabs/stylegan2/blob/eecd09cc8a067e09e12...
Finally, my biggest tip would be to step back from the problem and think: is there something simple you can do to reframe the problem? When I find myself in a situation where I'm spending a lot of time and energy trying to get a certain thing to work, I can sometimes do X instead for 80% of the benefit. Try to find something like that in this case.
FWIW the TPU profiler was the first tool I reached for. I never got it working. The bag of tricks above ended up giving me effective results on a variety of codebases with no profiler. (A usage graph is pretty crucial, though, which Colab TPUs don't provide.)
So there are a bunch of general tips for solving weird bottlenecks blindfolded.
To answer your question directly:
I am assuming the profiler from Tensorflow works with TPUs in the same way it does with GPUs?
Not really. You're supposed to use cloud_tpu_profiler: https://cloud.google.com/tpu/docs/cloud-tpu-tools
But yeah, if you give specifics (ideally a link to a codebase + dataset + script that runs it) then I can try to look for candidates of what might be the bottleneck.
I looked to use some for solving PDEs, but Google had literally zero documentation on how to cross-compile C to TPUs, launch kernels, etc.
AFAICT, you either use tensorflow or some other product that supports them, and for which the TPU code is not open source, or you can't use TPUs at all.
The ODE solver might be close to what you want.
I use Tensorflow 1.15. The world has been steadily pushing for Tensorflow 2.0 or Jax, but I like the simplicity of the Session model. It's so simple you can explain it in one sentence: it's an object that runs commands. Tell the session to connect to the TPU, and it will run all those commands on the TPU.
Jax is new to me (and to everyone; they just released it). But it looks like Google is pouring some serious R&D into it.
Two things help a lot. One, twitter. You can get a direct line to the people who actually make these beasts. Exploit it when you can. Like you, I dislike using a black box, and I'm intensely interested in the details of how to communicate with a TPU at a low level. I recently asked someone on the jax team about it here: https://twitter.com/theshawwn/status/1213221594052599808
Two, TFRC support has been incredibly helpful. https://www.tensorflow.org/tfrc I don't know who they have working the support channels, but those guys and gals are some of the most helpful and cheerful people I've come across. I often asked them very technical questions and to my surprise, they followed up with an A+ response almost every time, usually the next day.
Pytorch is giving TF a real run for its money, and to be honest I once felt it was a mistake to invest so much time into Tensorflow. But it turned out to be a big advantage due to Google's investment in the overall ecosystem. TPUs are something that only Google has the resources to pull off.
Note that the traditional path towards "just get a TPU up and running and start playing with it" is to use one of their Colab notebooks on the topic. https://cloud.google.com/tpu/docs/colabs I've been implicitly steering you away from these because you seem (like me) to want to know more of the low-level details. Those notebooks are designed to let ML researchers get results quickly, not for hardware enthusiasts to exploit heavy metal. The jax notebooks felt much more satisfying in that regard.
Can u mention how much human Dev time is involved?
We have a stupid-basic single machine Deep reinforcement Self play setup. It takes about 24 hrs to run a full experiment. The NN is the bottle neck. Using Tensor flow. Nothing fancy.
How much dev time for a good enginner (backend, kernel, multi core experience) to get this down to say 1hr ?
Obviously a very general question. Thanks for any input.
The personal ML bots will be a big things. Next step in total automation.
Something I've often wondered, and there are probably good reasons why, is that billionaire tech moguls - even the ones who are outwardly technical (or were in the past - people like Bill Gates, who we know had technical chops in the past) - that none of them (that I'm aware of) haven't ever tried to build "their ultimate computer".
For instance, if I had their kind of money, I've often thought that I would construct a datacenter (or maybe multiple datacenters, networked together) filled with NVidia GPU/TPU/whatever hardware (the best of the best they could sell me) - purely for use as my "personal computer". Completely non-public, non-commercial - just a datacenter I would own with racks filled to the brim with the best computing tech I could stuff into them (on a side note, I've also pondered the idea of such a personal datacenter, but filled with D-Wave quantum computing machines or the like).
What could you do with such a system?
Obviously anything massive parallelism could be useful for - the usual simulation, machine learning, etc; but could you make any breakthroughs with it - assuming you had the knowledge to do such work?
Which is probably why none have done it - at least as a personal thing.
I mean, sure, I would bet that people who own large swathes of machines in a datacenter, or those who outright own datacenter (like Google or Amazon) - their founders and likely internal people do run massively parallel experiments or whatnot on a regular basis, ad-hoc, and "free" - but it's a commercial thing, and other stuff is also running on those machines...
But a single person is probably unlikely to have or think of problems that would require such a grand scale before they would just "start a company to do it" or something similar; because in the end, just to maintain and administer everything in such a datacenter, if one were built, would require (I would think) the resources of a large company.
Of course, then I wonder if such companies - especially ones like Google and Amazon, which own and run many datacenters around the world, and also sell the resources of them for compute purposes - weren't started in some fashion (even if only in the back of their heads) by their founders with that idea or goal in mind (that is, to be able to own and use on their whim "the world's largest amount of computing power"...?
https://www.pcworld.com/article/3313424/inside-seattle-livin...
If anyone were to own a secret HPC cluster, it'd probably be a finance billionaire. Or the owner of a think-tank who made their money as a subcontractor for state intelligence agencies.
«A single GPU card like the AMD Radeon MI60 has more computing power than year 2000 supercomputer ASCI Red (fastest supercomputer in the TOP500 list of June 2000):
• MI60: 7.4 TFLOPS (FP64)
• ASCI Red: 3.2 TFLOPS (FP64) »
https://mobile.twitter.com/zorinaq/status/112491212518746521...
[1] https://en.wikipedia.org/wiki/Thread_block_(CUDA_programming...
I am tempted to build something similar but AMD's wishy-washy ECC guarantees as well as Linux-specific issues make me unsure.
ASUS PRIME TRX40-Pro motherboard
8 sticks Kingston KSM26ED8/16ME (16GB, 2666Mhz DDR4, ECC)
Dual boot Windows 10 / Ubuntu 18.04
2x Samsung EVO 970 NvME M.2 1TB SSDs in RAID0 configuration.
Running in a Coolermaster Cosmos case with a stupidly big air cooler at the moment, to be replaced with a decent liquid cooler (it works, the Cosmos is a huge case because it had to be to hold a hacked Supermicro dual Xeon server board in it before (I really wanted ECC for my workstation))
nVidia 1080Ti+ GPU.
The ECC is detected and claims it is working although I've yet to see it correct an SBE. I haven't been running non-stop memory tests either though so.
I may end up removing the Linux partition since WSL2 works so well on this box.
BTW, Century Micro has the only unbuffered ECC modules at 3200MHz native speed (at least that was the case on the 39x0X release day). I don't know if they can be sourced outside Japan, though.
Fun story, I actually botched my order and got 2666MHz ones... but on closer inspection, it turned out the chips on the modules were actually native 3200MHz ones. With the SPD EEPROM saying they are 2666MHz. So I ended up overclocking them at their actual native speed. And I tweaked the timings to be a little shorter than what the 3200MHz modules were advertized for.
[1] GPUs are another kettle of fish of course, but have their own wellknown problems that prevent widespread use outside graphics
BTW I seen 64GB DDR4 sticks on amazon...
Just got my 3950X w/ 64gb ram, not sure that I'd be able to practically use any more compute than this for what I play with, which is mostly multiple back-ends and some container orchestration for local dev and occasionally video re-encodes (BR-Ripping for NAS).
Some think $4k for this CPU is too much... considering the shear performance that you can get these days for under $10K there's never been a better time to build or buy a computer. My only regret is wasting time and money on aRGB that I cannot configure in Linux.
I do wish for exascale in the medical field though.
I even-more-wish that the energy world can see ~similar improvements in efficiency.
The amount of processing power we each carry in our pockets (even the cheapest throw-away smart phones) would have been almost unthinkable 30 years ago; it's akin to the difference of an Altair of the 1970s vs what was available just 10-20 years prior. What took up a room now sat on a desk and could be purchased for the price of a car.
Now, what took up a room now sits in your pocket, and almost could be given away in a box of cereal its so cheap.
Heck - think about what's available in the embedded computing realm for pennies (or just a few dollars in single quantities) - it's mind boggling to an extent.
According to Wikipedia it was launched in 1997 so it does line up. If I remember correctly, the system was bought from Cray after SGI bought the rest, so I didn't know they even had it at Sun in 1995.
Also, the original model of the E10k supported 64 GB RAM.
On one beautiful day during the weekend, only a two weeks after delivery, our system admin noticed that the machine went offline. Logging in remotely didn’t work at all. No ping either.
He drove to work and ... the machine was gone.
Thieves had used a crane to lift it out of the building through a window onto a truck.
Sun told him that this wasn’t the first time such a thing had happened and that somewhere in the chain from order to delivery, an insider tipped off the thieves about where to find the latest.
(I'm just as guilty, having spent over two grand on MBPs multiple times!)
I haven't priced out any recent macbooks to know if that's still true though. Glancing at the new 16" macbook pro, it seems like it might be reasonably priced for what you're getting.
However, the costs go up a lot if you spec out a custom config (+$400 just for 32GB RAM). Then the MBP starts looking quite a bit more expensive. Overall I don’t think they’re a bad buy though if you want macOS.
The Dell XPS [1] with a comparable spec cost $1650 compare to MBP 16" $2399. In the old days Apple would have priced it closer to $2199 or slightly lower.
Somewhere along the line they started making Mac same margins as iPhone.
[1] https://www.dell.com/en-us/shop/deals/new-xps-15-laptop/spd/...
I seriously doubt 5-10%. You are talking about minimum $50 - $100+ dollar difference, that has never happened. Mac has always been roughly 20-30% more expensive than a laptop with comparable specs. So $1000 comparable spec laptop, Apple will sell you one for $1300, ( But with more expensive upgrades )
The 30% has been fine for years, the quality and finishing as well as macOS was well worth the price tag. But in recent years it hasn't been 30% at all.
So I decided to give Apple a chance, and I am still using that laptop today, it has really been great value for my money.
From wikipedia:
"Intended as a replacement for the Portable, the 140 series was identical to the 170, though it compromised a number of the high-end model's features to make it a more affordable mid-range option. The most apparent difference was that the 140 used a cheaper, 10 in (25 cm) diagonal passive matrix display instead of the sharper active matrix version used on the 170. Internally, in addition to a slower 16 MHz processor, the 140 also lacked a Floating Point Unit (FPU) and could not be upgraded. It also came standard with a 20 MB hard drive compared with the 170's 40 MB drive."
If people are buying Apple products at those prices then why should Apple lower their prices? The answer is they shouldn't.
https://en.wikipedia.org/wiki/Apple-designed_processors#Appl...
https://web.archive.org/web/20100217014904/http://xnu-dev.go...
Although still not without issue, my workflow has aligned so much with Linux and it's finally reached a good enough point for my day to day use... not having to use the VM based mac or windows docker has been really nice (WSL isn't good enough imho).
1) Other companies that are willing to sell you workstation and laptop computers that you can run other operating systems on such as one of the Linux variants, Windows, BSD etc. Nobody is forcing you to buy a computer from Apple.
2) There is a thriving second hand market of Apple machines just look at ebay, craigslist, gumtree etc.
If a new Apple machine isn't worth it to you, you are free to buy alternatives.
I am going to buy one of the newer Lenovo Thinkpads as I don't think the MacBook pro is worth it to replace my ageing Macbook Pro.
Can you provide a citation for this? Having worked with literally thousands of engineers, I have never seen a hackintosh in real life.
YMMV.
It’s a respectable hobby, though, like iOS jailbreaking or emulation, and for persistent people it does let them run MacOS on more powerful hardware than they could afford.
Google Trends suggests that searches for 'Hackintosh' peaked around 2009 and have been steadily declining down to around a half since then. Searches for 'Ubuntu' dominate so much it makes the Hackintosh graph look flat by comparison, but Hackintosh seems about as popular as 'Manjaro' (Linux Distribution) currently is, fwiw:
https://trends.google.com/trends/explore?date=all&q=hackinto...
I gave OSX a test drive and found it much more simple compared to Windows. As I already had an iPhone/iPad it made sense to switch. Come upgrade time I bought a Macbook Pro and have been on OSX since.
What you're saying is the fact people are buying their machines means they shouldn't change price or value proposition... because there are in fact purchases; which, is quite frankly, baffling, to me, a humble idiot.
Mercedes must think their EQC is positioned perfectly in the market with 55 sales? I now imagine Magic Leap will be leaping to raise prices with their next version?
Obviously.
> What you're saying is the fact people are buying their machines means they shouldn't change price or value proposition... because there are in fact purchases; which, is quite frankly, baffling, to me, a humble idiot.
If Apple are selling the machines in sufficient quantities at whatever they are priced at (I haven't cared to look) then obviously Apple's customers think they are worth it. It isn't really more complicated than that.
> It's okay to be completely full of shit, just don't market it as truth.
I know you think you are being big brained but not everything is "you must provide a citation". It is pretty obvious Apple knows the market well, knows exactly what they can and can't charge for certain products. You pretending otherwise because I haven't provided you with a citation is a complete joke, it is like asking someone to cite for evidence that the sky is blue.
In conclusion, you think no company should adjust their prices, ever? Because the current price is the market price which is obviously the right price because it's the market price?
It's circular reasoning, which can be applied to any sales situation - and if it explains everything, it explains nothing.
Obviously not. I am saying they have no incentive to change the price if the sales are inline or above with what they would have forecasted.
> Because the current price is the market price which is obviously the right price because it's the market price? > > It's circular reasoning, which can be applied to any sales situation - and if it explains everything, it explains nothing.
Again you don't seem to understand basic market economics. Your product is only worth what people are willing to pay for it. There is the odd exception to the rule (Head and Shoulders Shampoo being one of them, which is priced far higher than they originally intended because people assumed it didn't work because it was cheap).
Generally if there are two or more companies producing product X (in this case Computer Workstations and Laptops) then the market will coalesce around a particular price point for a particular specification. Sure there are those that will always stick to a brand, but the vast number of consumers won't be loyal.
Whether or not the company makes a profit on each unit sold is irrelevant to its market price. If they price their product higher than their competitors people will look at the alternatives.
e.g. I bought a MacBook Pro in 2015 because Apple's machine was cheaper than Lenovo, Dell for the same spec and had a better screen than any of the competitors machines.
This really isn't complicated stuff. I think that personal bias seems to cloud people to some basic truths.
No. There is usually no such thing as _a_ market price. There's a distribution of prices for purchases of the same item (and that's when we ignore the cases of transactions involving more than just the transfer of money).
> the equilibrium of both supply and demand.
Supply and demand for a specific products are more the _result_ of socio-economic processes and phenomena rather their _causes_.
If you're talking to me about socio-economic processes when we're talking about something simple, it pretty clearly shows you're you're not very educated in Economics.
https://en.wikipedia.org/wiki/Supply_and_demand#Criticism
If you like your charts and economic formalisms, and believe in "market prices", perhaps you should take the time to read the Candide-like "Production of commodities by means of commodities" by Piero Sraffa.
So they definitely couldn't switch completely, and supporting both would be expensive for apple due to doubling mobos + testing + drivers etc, and confusing for the consumer because of the differing max RAM capabilities.
I am not the audience for a Mac Pro though, maybe it's fine?
I mean, I wouldn't re-write code to support hardware-accelerated video transcoding on Ryzen+GPU if I knew that in 2-3 years I was moving from x86-64 to ARM64.
Sure, there are a lot of neat tricks ARM can do with special instructions and hardware accelerators in very well controlled use cases. But, for the average creative professional who doesnt have time or patience to play with hyper-optimizing their workflow, having an x86 monster that can chew through any arbitrary workload (optimized or otherwise) is going to provide the best experience for the foreseeable future.
> Photoshop 1
> Photoshop 1 (1990.01) requires a 8 MHz or faster Mac with a color screen and at least 2 MB of RAM. The first release of Photoshop was successful despite some bugs, which were fixed in subsequent updates. Most users ended up using version 1.07. Photoshop was marketed as a tool for the average user, which was reflected in the price ($1,000 compared to competitor Letraset’s ColorStudio, which cost $1,995).
> Photoshop 1.x requires Mac System 6.0.3, 2 MB of RAM, a 68000 processor, and a floppy drive.
(Unfortunately I can no longer google the link of the video)
Which is not to say they will move to ARM. I am still skeptical of it.
QuickSync is only just a Hardware Video Encoder. Which AMD has as well, it is called VCN [1]. Not to mention Apple hasn't been using QuickSync for as long as they have been shipping T2, where Apple uses their own Video Encoder within T2. ( T2 is just a rebadged A10 )
USB 4 is out, the Thunderbolt 3 spec has been out for quite a long time as well, and yet we dont even have a single announcement with USB 4 controller.
At nearly $4000 for just the CPU, is that still consumer territory? I assume only huge companies would spend that much money on a single CPU.
More or less like consumer goods companies are using the "pro" keyword on pretty much any somewhat evolved product but the other way around.
This isn't the same experience I've had with the consumer threadripper at all, so I don't think these numbers make for a simple comparison.
My chips are pretty well cooled.
It might also be that "consumer" workloads will run the cores at the high speed more infrequently than a server which might be running full tilt 24/7. Just a thought
Which uses more energy at Max load?
[1] https://twitter.com/yiningkarlli/status/1204564015113895936
The Mac Pro is a specialty item, not the norm.
Those who can afford it, will buy it.
edit: yes, it is definitely overkill for most if not all avaiable games, but in a certain scene "overkill" is considered awesome
https://www.intel.com/content/www/us/en/products/processors/...
Most games are not particularly CPU intensive (although there are exceptions like Ashes of the Singularity which actually does usually get bottlenecked by the CPU unless you have a really high end one)
Intel chips are still competitive for single-threaded performance, from what i can tell https://www.pcworld.com/article/3453946/amd-threadripper-397...
With the next console gen being based on Ryzen instead of the much less efficient Jaguar architecture, maybe 8 cores might be better used.
However, Lots of the youtube influencer ruling class will buy it, in part because of what you said ("those who can afford it..").
The real legitimate consumer base for this are people whose work productivity is held back by compute loads that are embarrassingly parallel. If you spend a lot of your time waiting for a (well threaded) compiler to finish, or blender to render something out, or whatever.
(at least for now, I don't work in that industry but it may just be a code quality issue)
Twitch streamers need a "streaming" solution. They play live, and instantly react to the crowd. If someone pays for an emote or something, the Twitch-streamer is expected to look on camera and say thank you to the donor (and maybe repeat the message that the donor paid for).
This means that a Twitch streamer's computer MUST encode the gameplay live. Traditionally, Twitch streamers would buy two computers, one to play video games, and a 2nd computer to process the video stream and upload it to Twitch.
With the advent of 16+ core computers, Twitch streamers have begun to simply buy one computer, lock 8-cores to the video game, and then lock 8-cores to the Twitch encoder.
Presumably, something like a Threadripper (24+ cores) could process the video stream for better quality and lower bandwidth. Maybe live VP9 encoding, for example (8-cores for the video game, 16-cores or more for the encoder)
PS: Although FLOPS is not a good way to measure these stuff, it's a good indication of possible upper bound for deep learning related computation.
CPU’s do general computing. They’re super flexible but if you have a specific workload you might be able to use a different piece of silicon to get more performance.
GPU’s do more parallelized computing but they don’t do as many different operations. They’re really good at doing a small not super complex fast but massively parallelized (like updating an array of pixels on a screen, for example).
TPU’s are even more parallelized but the operations they do are even more specific and often simpler than the operations GPUs do.
Those things make CPUs faster at sequential execution.
In effect: CPUs are latency optimized. GPUs are bandwidth optimized.
The more specific you get on the circuit the less flexible it is and the more bandwidth you get at a specific task.
The tradeoff is flexibility for application specific performance. CPUs can do hella stuff but they can’t do a specific thing faster than specialty hardware.
The thing is, GPUs have horrible latency characteristics compared to CPUs. Whenever a GPU "has to wait" for RAM, it only switches to another thread. In contrast, CPUs will search your thread for out-of-order work, speculative work, and even prefetch memory ("guessing" what memory needs to be fetched) to help speed up the thread.
--------
Consider speculative execution. Lets say there is a 50% chance that an if-statement is actually executed. Should your hardware execute the if-statement speculatively?
Since CPUs are latency optimized, of course CPUs should speculate.
GPUs however, are bandwidth optimized. Instead of speculating on the if-statement, the GPU will task switch and operate on another thread. GPUs have 8x to 10x SMT, many many threads waiting to be run.
As such, GPUs would rather "make progress on another thread" rather than speculate to make a particular thread faster.
---------
What problems can be represented in terms of a ton-of-threads ? Well, many simple image processing algorithms operate on 1920 x 1080 pixel entries, which immediately provides 2,073,600 pixels... or ~2-million items that often can be processed in parallel.
When you have ~2-million items of work (aka: "CUDA Threads") waiting, the GPU is the superior architecture. Its better to make progress on "waiting threads" than to execute speculatively.
But if you're a CPU with latency-optimized characteristics, programmers would rather have that if-statement speculated. The 50% chance of saving latency is worth more to a CPU programmer.
Hmm, I think I see what you're trying to say, but maybe more precise language would be better here.
GPU cores have extremely efficient thread-barrier instructions. NVidia PTX has "barrier", while AMD has "S_BARRIER". Both of which allow the ~256 threads of a workgroup to efficiently wait for each other.
-------
The other aspect, is that "Waiting on Memory" (at least, waiting on L2 memory) is globally synchronized in both AMD and NVidia systems. Waiting for an L2 atomic operation to complete IS a synchronization event, because L2 cache has a total memory ordering on both AMD and NVidia platforms.
Tying L2 cache to higher levels allows for memory coherence with the host CPU, or other GPUs even. That is to say: "dependencies" are often turned into memory-sync / memory-barrier events at the lowest level.
Synchronizing threads is one-and-the-same as waiting on memory. (Specifically: creating a load-and-store ordering that all cores can agree upon).
---------
I think what you're trying to say is that dependency chains must be short on GPUs, and that there are many parallel-threads of dependency chains to execute. In these circumstances, an algorithm can run efficiently.
If you have an algorithm that is an explicit dependency chain from beginning to end (ironically: Ethereum hashing satisfies this constraint. You work on the same hash for millions of iterations...), then that particular dependency chain cannot be cut or parallelized.
But Ethereum Hashing is still parallel, because there are trillions of guesses that can all work in parallel. So while its impossible to parallelize a singular ETH Hash... you can run many ETH Hashes in parallel with each other.
* NVidia RTX 2070 Super is 9 TFlops for $500
True, Radeon VII and RTX 2070 are "consumer" GPUs... but Threadripper is similarly a "consumer" CPU and commands a lower price as a result.
"Enterprise" products cost more. EPYC costs more than Threadripper, V100 costs more than RTX 2070 Super. If you're aiming at maximum performance at minimum price, you use consumer hardware.
Similarly, Threadripper loses on RDIMMs, LRDIMMs, and have 1/2 the memory channels and 1/2 the PCIe lanes. Most people don't need that either.
Of course consumer products have fewer features than enterprise products. The chip manufacturers need to leave some features for "enterprise". The general idea is to extract more wealth from the people who can afford it, while providing consumers the features they care about at a lower cost.
-----
Really, the feature enterprise GPUs need seems to be SR-IOV, or other PCIe-splitting technologies. This allows a singular GPU to be split over many VMs.
Double-precision floats is niche to the scientific fields, also "enterprise" but I don't think most enterprise customers use double-precision floats.
In particular, the 3990x, 64-core Threadripper will have 256MBs of aggregate L3 cache, and 512kBs L2 cache per core (32MBs of L2 cache). Highly-optimized kernels may fit large portions of data within L3 cache and rarely even touch DDR4!
Note: Each L3 cache is only 16MBs between 4-cores. It will take some tricky programming to split a model into 16MB chunks, but if it can be done, Threadripper would be crazy fast.
True, GPUs have really fat VRAM to work with, but CPUs have really fat L3 cache and L2 cache to work with. And the CPU caches are coherent too, simplifying atomic code. GPUs do have "shared memory" and L2 caches, but they're far smaller than CPU-caches.
More accurately, Cerebras has 9.6 "bullshitobytes" per second. If you can't verify this, it doesn't exist. You could claim insane "bandwidth" by considering your register file to be your "memory". But that doesn't make it so.
Chips like these are for a specialty market, for those who are using applications and workloads that can actually take advantage of all those cores and aren't running a personal datacenter. You're not going to see this offered in any of the rack/blade systems offered by the likes of Dell, HPE, SuperMicro, Lenovo, etc where organizations are actually going to be purchasing EPYC chips.
When running with many cores is there some way to get the Linux kernel to run them at a fixed frequency?
I've been doing multicore work lately on AWS - which works pretty well but you don't get access to event counters so sometimes I'm having trouble zeroing in on what the performance bottleneck is. At higher core counts I get weird results and I can never tell what's really going on. Running locally I have concerns about random benchmark noise like thermal throttling, "turbo" and other OS/hardware surprises (I know on modern chips single thread stuff can run really different from when you load all the cores). I've been thinking of getting something a bit dumber like an old 16 core Xeon (I'm on a bit of a budget) and clocking it down - or maybe there is some better solution?
just a wild thought: have you looked at 'cpupower frequency-set' ? it _might_ help.
it's a big chunk of metal, but it would make life easier for some stuff
The other option if it's an occasional thing is you can change it in your BIOS and lock them (I believe) to a specific speed.
Since you're not doing it for higher performance just consistent performance you can lock it to your base clock frequency and it should be stable.
seldon ~ # cpupower frequency-set -f 1.86GHz
Setting cpu: 0
Setting cpu: 1
Setting cpu: 2
Setting cpu: 3
Setting cpu: 4
Setting cpu: 5
Setting cpu: 6
Setting cpu: 7
seldon ~ # cpupower --cpu all frequency-info | grep -E '^analyzing CPU.*|current CPU frequency'
analyzing CPU 0:
current CPU frequency: 1.86 GHz (asserted by call to hardware)
analyzing CPU 1:
current CPU frequency: 1.86 GHz (asserted by call to hardware)
analyzing CPU 2:
current CPU frequency: 1.86 GHz (asserted by call to hardware)
analyzing CPU 3:
current CPU frequency: 1.86 GHz (asserted by call to hardware)
analyzing CPU 4:
current CPU frequency: 1.86 GHz (asserted by call to hardware)
analyzing CPU 5:
current CPU frequency: 1.86 GHz (asserted by call to hardware)
analyzing CPU 6:
current CPU frequency: 1.86 GHz (asserted by call to hardware)
analyzing CPU 7:
current CPU frequency: 1.86 GHz (asserted by call to hardware)
this an oldish Intel(R) Core(TM) i7 CPU. just tried it on an odroid, and it seems to do the right thing there as well.I don't know about a nice high-level linux interface, there may well be one, but if you need to access these settings from a running machine there are MSRs you can poke.
Plenty of PCIe space to afford a few GPUs to experiment with your 64-core CPU. You probably can shove 2x GPUs + 1x FPGA into one box.
Threadripper is only 300W+ under load. Idles for the CPU are ~100W for the whole system. (https://www.kitguru.net/components/cpu/luke-hill/amd-ryzen-t...)
GPUs are similar. GPUs take ~20W while idling, but 300W or 500W while under load. FPGAs are also similar.
-------
The total system idle of a Threadripper + 2x GPU + FPGA rig probably is under 200W.
If you happen to utilize the entire machine, sure, you'll be over 1000W. But you'll probably only utilize parts of the machine as you experiment and try to figure out the optimal solution.
A "mixed rig" that can handle a variety of programming techniques is probably what's needed in the research / development phase of algorithms. Once you've done enough research, you build dedicated boxes that optimize the ultimate solution.
Your earlier comment makes sense now that I understand what you're saying.
now whether its worth selling a body organ to kit out..
I don't think Windows 10 has this limitation anymore but it's a very valid question. Given their weird calculation for the number of licenses you need to run Windows server, it's probably good to be cautious about licensing weggeweest it comes to running Windows on high performance chips like Threadripper.
There still is. And the Threadripper is just one CPU, so you should be ok.
With pro edition of Win 10 there should be support for up to 2 CPUs, with whatever amounts of cores they have.
Edit: according to this, the normal windows 10 pro edition would suffice. I feel like I’ve read conflicting things though, so take it with a grain of salt: https://answers.microsoft.com/en-us/windows/forum/windows_10...
IIRC windows's scheduler wasn't as good as linux's at managing these kind of paralel workloads.
In embedded/automotive, majority of the tooling does not have a linux version. Compiling is a bitch. Still, you're probably right about the 99% comment.
Although, I’m not sure how the 64 cores are presented to the OS. If they’re presented as, say, 2 32 core CPUs, Windows will complain.
like: zone1: 62 cores, zone2: 2 cores
the OS was also rather unstable.