64 Core Threadripper 3990X CPU Review
anandtech.com
anandtech.com
Microsoft deliberately does these limitations in order to force people to pay more for its sofware. It's a shame, really.
Tesla doesn't and they still dot it.
[0]: https://www.linqpad.net/Purchase.aspx (Purchase site has the version comparison)
edit: I’d actually go as far as saying that by your definition almost every good software tool has excessive market power.
That's the "affirming the consequent" fallacy. Specifically, siblings start from the true statement that: people buy software with locks.
They then affirm the consequent, that software with locks is the correct market solution.
(of course they do also bin based on demand/sales, which is scummy.)
Why is this scummy? It's a free market and it's not a life saving drug, they can ask for whatever they want, nobody is forcing you to buy it.
That's like saying "Software engineers are paid according to demand/market which is scummy."
The best business strategy for companies and professionals alike is to find a niche with less competition where you are a leader(cough, monopoly, cough) and charge whatever the market will bear. If Apple/Microsoft/AMD/Intel can get away with that it means the market can bear it. The fact that you don't like it, is another issue.
Of course, in practice, it can be hard to tell whether intentional "crippling" is going on. Yield problems are common and might manifest in the exact same way, after all - for example, chips that don't "make timing" for their design frequency can be sold for lower specs.
No, it's deliberate discounting of economy products!
What's the difference? I don't know either. Price discrimination is a very real thing that exists virtually everywhere. It's why the market will bear Beats headphones in the same shops that sell generics at 20% the price.
At the end of the day, a manufacturer doesn't "owe" you something at cost. The idea of capitalism is that competition in the market will drive prices down instead, and it does. But one side effect is that customers and sub-markets that are able to bear higher markups on their versions of stuff are asked to pay it even when it doesn't make sense at the level of per-part costs.
So it's true, that that Beats headset costs only $4 to make and "should" be available much cheaper than it is. But it isn't, and the reason is simply that people are willing to buy them at $200. If you're not, that seems unfair, I agree. But... what's the alternative?
That's a very nice way and non-scummy way of putting it, yes. If that deep discount compared to the "full-featured" price was made clearer, I assume that many people would not even find it scummy that these discounts are conditioned to being sold a "crippled" version of the product. So, depending on the "framing" you pur on this, you can nicely account for either of these perspectives.
intel/amd could not afford to sell hardware at consumer prices if buyers could turn around and put those consumer parts in servers at scale. in the alternate universe where certain features aren't disabled for consumer parts, your i7 costs as much a a xeon; you don't get xeon features for i7 money.
Sure, they are paying more per person. But in exchange they are getting more room, more attention from attendants, boarding first, etc.
If the plane was completely full of people in coach and had no first class at all, I don't think that means the airlines would lose money.
Also, if there were no people in coach, then the first class tickets would be astronomical.
That is because for the vast majority of transactions, there is a very large gap between the highest price a consumer is willing to pay and the lowest price a producer is willing to accept. The canonical example is for essentials like food: people are willing to pay every single penny they have for food for their family. OTOH, farmers are willing to sell wheat in bulk for under 10 cents a pound.
For obvious reasons farmers cannot set different prices for each pound of wheat to get "all the money". It's seen as deeply unfair when the vast majority of industries are price takers while a select few can charge the maximum a customer can bear. If every industry could effectively price differentiate could we'd be spending all our money on food with no money left to buy CPU's and our whole free market system would collapse.
Of course people need food, but there are many kinds of food and it's competitive between the different foods.
Lifesaving drugs are something where it's often difficult to substitute.
That's pretty much the only argument _for_ doing it; it's the equivalent of defending your argument as free speech - just because you _can_ say it, doesn't mean you should.
It's scummy and inefficient from an obvious human perspective to intentionally cripple a superior product in order to target a cheaper market, but that's the natural result of market actors being incentivized to maximize their own return, they don't really have any reason to care about the human side of that argument.
you know, I hadn't actually thought of that. If you did this, presumably it would involve the cost of setting up another manufacturing line, (colossal), and you'd _still_ have binning issues. Presumably the tolerances would be looser however; if you can build within 5% power usage of X, you should be able to _easily_ build within 5% power usage of (x/2).
Interesting thoughts.
Maybe the other 70 have some major issue with the special components, so they are only 99% functional and therefor can't be used in the high end segment, but can be used in the consumer chips.
Not every chip is perfect, which is why there exists different product classes. For some chips, an i5 is just an i7 that didn't fully pass validation.
Googles it...
$84 a year per user. Or $168 for E5 per year.
That just doesn't strike me as Microsoft being bad?
I would bet that many other bits of software also have this limitation because they too thought that using a 64bit value for a processor mask would be sufficient.
----------
Furthermore, your Threadripper 3990x should ONLY really be pushing affinities for across a 8c / 16-thread die anyway, at least as far as the Windows scheduler is concerned.
Windows programmers should use multiple thread groups for different 16-thread CCX. Because its extremely costly to move all of your local state from one die's L3 cache into another.
https://images.anandtech.com/doci/15318/amd_rome-678_678x452...
Look at the chip, like physically look at it. AMD has 9x chips here (1x I/O die for memory, and 8x compute chips, each with 8x cores). You really only want to be moving threads within those 8x cores, because moving a thread across chips is less efficient.
So the ideal world, people would understand Windows's scheduler and work around it. Instead of complaining about how it is different from Linux's. Windows's Thread groups being 64-sized is reasonable for the job it wants to do... and there are other parts of the API that allow you to work across Window's 64-sized processor groups. In particular, to increase your affinities across to a 2nd processor group.
It should be fixed by a professional coder who doesn't impose limitations, especially in an OS, based on wanting to fit bit flags into a single variable.
Windows should fix this -- there should be nearly unlimited cores per group. They can still have limitations per Windows version so they can price segment their market.
Someone is going to run an OS from 2020 on a brand new machine in an age where 64 threads is common place ?
These things get segmented according to market - home is tailored to suit commonplace.
Then 8 jumped to 128GB.
In this case it seems pretty justified to include this feature in the enterprise edition only.
A stolen or misappropriated product key is almost as bad as a virus, though. It will work great... but only at first, until Microsoft notices that 50,000 customers are using it.
I've bought numerous Windows 7 and a couple of Windows 10 keys, plus a few MS Office keys, never had any issues.
Sometimes, the keys would require phone activation, but that was ages ago. Nowdays they just work.
The providers give you a key that you can use to activate, and a link for the ISO, but that part is optional. The keys work with untouched ISOs, too. FWIW, I once downloaded the seller's ISO and compared the hashes, and they matched up.
I may have to add, I live in Germany, which is part of the EU, and the courts here have decided key resale is completely legal and MS has to tolerate it.
["Volume license software may be sold individually"]: https://translate.google.com/translate?sl=auto&tl=en&u=https...
[original]: https://www.golem.de/news/bgh-urteil-software-aus-volumenliz...
This program replaced it, but I have zero info besides knowing it replaced BizSpark so might not be comparable:
Yes, it could be a config option as well instead of a fixed choice.
The main reason this example is not relevant anymore is the "alternatives" system Linux uses nowadays, which dynamically patches the kernel code according to the hardware features (in this case, it detects it's running on a single core and removes the SMP locking code).
I can find a billion blog posts about notepad supporting Linux and Windows line returns. I struggle to find anything about improvements for new hardware.
I wish there was a blog (albeit unfashionable these days) or similar which told these stories. I often feel the lack of their existence is a taint on data driven decisions: these customers are likely insignificant, but their influence on others is significant (even if their advice isn’t “buy a 64 core CPU”, it will be “Windows 10 is great, use that!”).
I’ve largely been a MSFT fanboy since a young age, and I kind of get being mainstream, but tossing a bone to the enthusiastic supporters would be nice.
They called out this problem between Threadripper 3990X and Windows 10 a bit ago.
This was his video on the 3990X recently for those interested: https://youtu.be/1LaKH5etJoE
A kernel developer is going to be 3-4 steps removed from those decisions. In many ways they are just a consumer of a product like this too and work on their subsystem. They are unlikely to be worried because it's the product they build and might even feel that it's inexpensive for the level of engineering they invest.
We mostly operate sTR4s as HEDTs, but we also have X399 server boards with IPMI and 10GbE in a board configuration designed to be installed racked. To us, there's only a minor functional difference in role.
A pair of used 7601s, motherboard, ram, power supply and chassis would still be cheaper for a SOHO setup.
Would you like to chat about opportunities to do analagous work for GPU instruction sets?
I work at a startup making high-performing GPU software and compilers, incidentally including a regex engine! (We can't quite match HyperScan on the CPU, but support capture groups and very high throughput on GPUs.) We also have several other interesting projects, and would like to start a superoptimizer at some point.
P.S. After reading your blog, one of our engineers said: "We see you like PSHUFB. So do we."
Run Crysis only on the CPU.[0]
I work at a video production company, with a good mix of Linux, FreeBSD, and Windows for the Adobe torture.
Threadripper: incredible. The CPU is so strong, we struggle (NVMe RAID0) to keep up with disk IO.
So I thought some of our more parallel tasks could be moved to Epyc.
In the end 80-90% of the cores sit idle. So it got repurposed to a VMHost, and it does that good well enough.
Kind of off topic, but it’s the kind of nasty surprise I wouldn’t want to get after deciding to buy, so hope this helps someone.
https://windowsserver.uservoice.com/forums/295047-general-fe...
To be fair to the Windows team, AMD in the data center / pro desktops wasn't really viable for a very long time, so its understandable that it wasn't prioritized.
Sounds like it’s coming.
> 1080p60 HEVC at 3500 bitrate with "Fast" preset - 319 fps
Ok, why were these parameters chosen? What's the application? I recommend everyone to look at 1080p60 video footage encoded with h265 with the "fast" encoder preset at 3500 bitrate. Calling it terrible would be a compliment. Unless you encode really slow and visual easy motion, which brings up the question why you would need 60 fps in the first place. Even at the "medium" preset with 1080p60 you should - regardless of application - be at least in the 5000+ range with your bitrate. And even that comes with a lot of trade offs, because that's just where live streaming starts.
Fun fact: There's no point in running x265 on the fastest presets unless you absolutely need to have HEVC; x264 on slow is faster _and_ gives better quality per bit. See the second graph on https://blogs.gnome.org/rbultje/2015/09/28/vp9-encodingdecod... (a few years old, the situation is likely to look similar but not identical).
x265 has made massive improvement over the years, 2015 x265 wasn't even considered good; despite all of its hype, or another way to think about it is how well x264 managed to squeeze every last bit of detail possible.
If I am investing USD 4000 on a CPU, I'd probably go for the EPYC part for 500 more and twice the memory bandwidth. It'd be interesting to run these benchmarks under perf to see how many L3 cache misses happen and how much they cost in cycles.
btw, the cost for server component would probably be more expensive too right?
Video games and pointer-chasing care more about that latency figure than bandwidth. I'm sure there are bandwidth-bound tasks, but any latency-bound task would prefer highly-clocked, over-volted (1.35V) DDR4 UDIMMs at 3200 MT/s CAS16 or faster.
Server RDIMMs are naturally slower, due to the register (and LRDIMMs are probably even slower). RDIMMs and LRDIMMs are designed for capacity more so than latency: you can have 1TB of LRDIMMs RAM on one machine but all the RAM runs slightly slower (3200MT/s maybe CAS22 or slower) as a result.
is that like twice as slow?
[0] https://www.anandtech.com/show/15359/trx80-and-wrx80-dont-ex...
If you want a system that can be a workhorse 128-thred monster on the working days, but reach 4.5GHz boost on the weekends for CS:GO and other video games... you'll want a Threadripper, not an EPYC.
The low clock speeds, and the RDIMMs / LRDIMMs of EPYC add latencies which slow down video game (mostly single-threaded) performance.
--------
For those who don't play video games: there are a variety of single-threaded tasks still in various workplaces. A surprising amount of 3d graphics work remains single-thread bound.
In particular, modeling is typically single-thread bound (the GUI thread, where the user is clicking menus and such). Most custom scripts are single-thread bound, and 3d modelers need a LOT of custom scripts. Those scripters aren't necessarily optimization masters who know how to take advantage of multi-core architectures.
3d Rendering is of course multithreaded. But the 3d artist still needed to click on a lot of menus and scripts to get the mesh to look right.
For web applications one usually has a software chain like web server <-> wsgi server <-> dozens of python instances.
Standalone processes just implement threading, which also is fairly easy (as far as theading itself can be easy).
Scientific libraries like scipy can use parallel processes automatically in the background (using things like BLAS), as long as the data is modelled correctly.
The current state of threading and parallel processing in Python is a joke. While they are still clinging to the GIL and single-core performance, the rest of the world is moving to 32 core (consumer) CPUs.
Python's performance, in general, is a crappy[1] and is beaten even by PHP these days. All the people that suggest relying on multiprocessing probably haven't done anything that's CPU and Memory intensive because if you have a code that operates on a "world-state" each new process will have to copy that from a parent. If the state takes ~10GB each process will multiply that.
Others keep suggesting Cython. Well, guess what? If I am required to use another programming language to use threads, I might as well go with Go/Rust/Java instead and save the trouble of dabbling with two languages.
So where does that leave (pure-)Python? It can only be used in I/O bound applications where the performance of the VM itself doesn't matter. So it's basically only used by web/desktop applications that CRUD the databases.
It's really amazing that the machine learning community has managed to hack around that with C-based libraries like SciPy and NumPy. However, my suggestion would be to drop GIL and copy the whatever model has been working for Go/Java/C#. If you can't drop GIL because some esoteric features depend on that, then drop them as well.
[1] https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
I think multi-interpreting is the way to go, but that still would require a framework for ensuring safe memory access.
Speaking of Go, I always thought it would be neat to write a python implementation in Go, but leverage Go's GC, and implement the 'go' keyword/function for easy parallelism. But you still have the problem of scoping and memory safety. Or similar idea but with Rust. Something tells me that isn't a trivial undertaking, especially if you want all the libraries, which is 75% the point of python.
This is wrong, there are multiple ways of python threads working on shared data.
> It's really amazing that the machine learning community has managed to hack around that with C-based libraries like SciPy and NumPy.
Well, the main implementation of the whole language is c-based. I can't see how that implies hackiness.
> If you can't drop GIL because some esoteric features depend on that, then drop them as well.
There have been multiple python implementations without a GIL available for way over 10 years, for example pypy, ironpython and jython. Yet these never went mainstream, which strongly implies the GIL problem actually isn't that much of a problem in the real world.
I think this is a problem
[0] https://docs.python.org/3.8/library/multiprocessing.html
It seems to be working fine for me on kernel 5.5.0 on a 3800x and on a 3950x...
[1] https://www.phoronix.com/scan.php?page=news_item&px=K10temp-...
https://stackoverflow.com/questions/22597324/what-cache-inva...
It's possible that Tesla does something different that makes sense in their specific situation, but general purpose CPUs - as far as I know - don't use NNs for this.
Multi-megabyte NNs for a normal branch predictor would be impossible; way too high latency to be useful.
That's just a fancy way of saying matrix multiplication :)
[1] https://stackoverflow.com/questions/37070/what-is-the-meanin...
[2] https://en.wikipedia.org/wiki/Control_register
[3] http://man7.org/linux/man-pages/man3/posix_madvise.3.html
One simple policy would be to try to ensure your application is as small as reasonably possible. If you can fit your entire executable image in a fraction of your L3 you are probably sitting in a really good spot.
You may want to avoid the issue by giving up 10% if your performances, less so by giving up 50%.
For what it is worth, I have an AMD Ryzen 2700X Eight-Core Processor that I got in 2018, and I keep SMT off. I do some light gaming with it, and I am happy. I did not notice a big drop in performance, but I did not truly measure the difference.