213 karma · joined November 27, 2016
For example, the electricity costs of the lab in which the research is run would typically be paid for by the university and would be considered overhead. It's not "administrative bloat". Most of the particularly gross administrative bloat is on the undergraduate side of things where higher tuition costs have paid for more "activities".
The Steele dossier includes information from Russian sources. It was not "paid for" by Russia.
No one has argued that Hunter Biden was not being paid because he was a Biden. The question was whether it led to improper behavior on the part of Joe Biden. That email proves nothing. It's entirely plausible that Hunter Biden got this guy an invite to some social function where shook hands with the Vice President. This is also entirely ignoring the fact that the firing of the prosecutor wasn't some sort of one off idea that Biden came up with on his own. There are independent speeches by all sorts of people from 6 years ago that complain about the corruption of the prosecutor in question, which makes this whole point moot. The prosecutor was fired for not investigating the company that hired Hunter Biden.
Like most liars, Trump isn't inventing everything out of thin air, but he misrepresents so often and so much that you can't draw any useful conclusions from anything he says.
The types of workloads run on GPUs typically like very high memory bandwidth and are usually willing to live with higher memory latency to get it. Onboard GPU memory is usually built with this in mind (trade off capacity and latency for increases in bandwidth). This is generally speaking the opposite of what you want in a CPU where people often want very high memory capacity and lower latencies, but may not limited by memory bandwidth, so simply sticking a GPU on die and giving it access to a memory subsystem that was not designed to feed a GPU is not going to make anything better.
Yes, if you insert a single 512-bit FMA that runs every so often in your code you will get a 15% performance hit from the lower frequency, but that's much less likely than the old case where people who were trying to use AVX-512 for memcpy and the like would slow down scalar code.
From later in the same post:
> Here, we have the worst case scenario of transitions packed as closely as possible, but we lose only ~20 μs (for 2 transitions) out of 760 μs, less than a 3% impact. The impact of running at the lower frequency is much higher: 2.8 vs 3.2 GHz: a 12.5% impact in the case that the lowered frequency was not useful (i.e., because the wide SIMD payload represents a vanishingly small part of the total work).
Interestingly enough, this is another feature that is supposed to have been improved on server Icelake. The frequency transition halt time is now pretty much negligible. The "core frequency transition block time" goes from ~12 us on CLX (similar to the number quoted above) to ~0 us on ICX.
(Slide with frequency transition info: https://images.anandtech.com/doci/15984/202008171754441.jpg)
If the only issue with AVX-512 is thermal downclocking because you end up using more power, it's almost definitely because you are getting more work done per time. A few AVX-512 instructions in a mostly scalar workload is not going to significantly increase power dissipation and therefore should not induce thermal downclocking, while a heavily utilized AVX-512 kernel will burn power, but should also be doing work twice as fast per instruction.
(Intel Hotchips slide: https://images.anandtech.com/doci/15984/202008171757161.jpg)