Graviton 3: First Impressions
chipsandcheese.com
chipsandcheese.com
That's not correct. AWS sells Graviton 3 based EC2 instances at a higher price than Graviton 2 based instances!
For example a c6g.large instance (powered by Graviton 2) costs $0.068/hour in us-east-1, while a c7g.large instance (powered by Graviton 3) costs $0.0725/hour [1]. Both instances have the same core count and memory, although c7g instances have slightly better network throughput.
I believe that is pretty unusual as, if my memory serves me right, newer instances family generations are usually cheaper than the previous generation.
I don’t think the point was that they would increase the cost of existing instance types, only that over time and generations the price will trend upwards as more workloads shift over.
Nearly every business seeks to maximize profit. Right now AWS is in growth phase - why wouldn't they raise rates in the future?
I bet quite a lot of migrations involved changing hardcoded instance types, docker image shas etc. Would be hard to revert back to quickly.
I think the take is pretty good one.
Offering a new option is not a price increase. You can still do all the same things at the same prices, plus if the new thing is more efficient for your particular task you have an additional option.
Universally everyone understands "raising prices" to be - "raising prices without any customer action".
As in you consider your options, take into consideration pricing, design your architecture, you deploy it, and you get a bill. Then suddenly, later, without any action of your own, your bill goes up.
THAT is raising prices, and it is something AWS has essentially never done.
What you're describing is a situation where a customer CHOOSES to upgrade to a new generation of instances, and in doing so gets a larger bill. That is nowhere near the same thing.
I'm sure once hardware settles it'll go down in price. Given this was announced like 6 months ago, I'm sure they wanted to launch something.
It doesn't even have spot pricing yet.
[0] https://www.phoronix.com/scan.php?page=article&item=aws-grav...
So I believe that using Graviton 3 at these prices is still a much better deal than using Graviton 2.
This seems ... wrong? I haven't tried it but according to the link below SVE2 intrinsics are supported in GCC 10 (and Clang 9):
https://community.arm.com/arm-community-blogs/b/tools-softwa...
Moreover, starting with the 8.1 version, gcc began to use SVE in certain cases when it succeeded to auto-vectorize loops (if the correct -march option had been used).
Nevertheless, many Linux distributions are still shipped with older gcc versions, so SVE/SVE2 does not work with the available compiler or cross-compiler.
You must upgrade gcc to 10.1 or a newer version.
However, it wasn't exactly a raging success, with I think the predicted amazing compiler tech not materialising, but maybe it is the right answer, but the implementation was wrong? I'm no CPU expert...
I do think a big part of the problem is that people want to distribute binaries that will run on a lot of CPUs that are physically really different inside. But nowadays there's JIT compilation even for JavaScript, so you could distribute something like LLVM, or even (ecch) JavaScript itself, and have the "compiler scheduling" happen at installation time or even at program start.
There have been a small number of attempts since Itanium, like NVIDIA's Denver, which make for much better baselines. I don't think those are anywhere close to optimal designs, or really that they tried hard enough to solve in-order issues at all, but they at least seem sane.
Can branch prediction be turned off on a compiler or application level? If you're optimizing for energy use that is. Disclaimer: I don't actually know if disabling branch prediction is more energy efficient.
Or is the worry on the other side; that processors have gotten so out-of-order that only huge dedication to guesswork can keep the beast sated? I don’t see this as a million miles from software techniques in JiT compilers to optimistically optimize and later de-deoptimize when an assumption proves wrong.
I think you might be right to be nervous if you wrote programs that took fairly regular data and did fairly regular things to it. But, as Itanium learned the hard way, programs have much more dynamic, emergent and interesting behaviour than that!
The data cache memory is one of the solutions to avoid the extremely long latency of loading data from a DRAM memory.
The alternative to a data cache memory is to have a hierarchy of memories with different speeds, which are addressed explicitly.
The latter variant is sometimes chosen for embedded computers where determinism is more important than programmer convenience. However, for general-purpose computers this variant could be acceptable only if the hierarchy of memories would be managed automatically by a high-level language compiler.
It appears that writing a compiler that could handle the allocation of data into a heterogeneous set of memories and the transfers between them is a more difficult task than designing a CPU that becomes an order of magnitude more complex due to having a hierarchy a data cache memories and a long list of other hardware mechanisms that must be added due to the existence of the data cache memory.
Once it is decided that the CPU must have a data cache memory, a lot of other hardware design decisions follow from it.
Because there is an inverse relationship between the load latency and the data cache memory size, the cache memory must be split into a multi-level hierarchy of cache memories.
To reduce the number of cache misses, data cache prefetchers must be added, to speculatively fill the cache lines in advance of load requests.
Now, when a data cache exists, most loads have a small latency, but from time to time there still is a cache miss, when the latency is huge, long enough to execute hundreds of instructions.
There are 2 solutions to the problem of finding instructions to be executed during cache misses, instead of stalling the CPU: simultaneous multi-threading and out-of-order execution.
For explicitly addressed heterogeneous memories, neither of these 2 hardware mechanisms is needed, because independent instructions can be scheduled statically to overlap the memory transfers. With a data cache, this is not possible, because it cannot be predicted statically when cache misses will occur (mainly due to the activity of other execution threads, but even an if-then-else can prevent the static prediction of the cache state, unless additional load instructions are inserted by the compiler, to ensure that the cache state does not depend on the selected branch of the conditional statement; this does not work for external library functions or other execution threads).
With a data cache memory, one or both of SMT and OoOE must be implemented. If out-of-order execution is implemented, then the number of registers needed to avoid false dependencies between instructions becomes larger than it is convenient to encode in the instructions. so register renaming must also be implemented.
And so on.
In conclusion, to avoid the huge amount of resources needed by a CPU for guessing about the programs, the solution would be a high-level language compiler able to transparently allocate the data into a hierarchy of heterogeneous memories and schedule transfers between them when needed, like the compilers do now for register allocation, loading and storing.
Unfortunately nobody has succeeded to demonstrate a good compiler of this kind.
Moreover, the existing compilers have frequently difficulties in discovering the optimal allocation and transfer schedule for registers, which is a simpler problem.
Doing efficiently the same for a hierarchy of heterogeneous memories seems out-of-reach for the current compilers.
And last not least, unknown memory latency is not the only source of problems, branch (mis)predictions are another. And they have the same remedies as cache misses: multithreading and speculative execution.
So if you wanted to get rid of branch prediction as well, you could come up with something like the CRAY-1.
However, for this, fine-grained multi-threading is enough. Simultaneous multi-threading does not bring any advantage, because the thread with the mispredicted branch cannot progress.
Out-of-order execution cannot be used during branch mispredictions, so like I have said, both SMT and OoOE are techniques useful only when a data cache memory exists.
Any CPU with pipelined instruction execution needs a branch predictor and it needs to execute speculatively the instructions on the predicted path, in order to avoid the pipeline stalls caused by control dependencies between instructions. An instruction cache memory is also always needed for a CPU with pipelined instruction execution, to ensure that the instruction fetch rate is high enough.
Unlike simultaneous multi-threading, fine-grained multi-threading is useful in a CPU without a data cache memory, not only because it can hide the latencies of branch mispredictions, but also because it can hide the latencies of any long operations, like it is done in all GPUs.
Fine-grained multi-threading is significantly simpler to implement than simultaneous multi-threading.
Alloc/GC: https://github.com/dotnet/runtime/blob/main/docs/design/core...
Reg alloc: https://github.com/dotnet/runtime/blob/main/docs/design/core...
The interesting probabilities are all decided at runtime.
Now we have AI workloads there is a place for a big lump of dumb compute again, but not in general purpose code.
Since processors are expensive and hard to change, they do tricks to allow themselves to be used more efficiently in common cases. That seems like a reasonable behavior to me.
I’ve wondered why Apple Silicon made the trade off decision to not include SVE support yet, given that support for lower precision FP vectorization seems like it could have made their NVidia perf gap smaller.
In my limited and sheltered experience, the only viruses I've gotten in the past decade or so was from dodgy pirated stuff or big "download" button ads on download sites.
I agree however that certain AWS services are disproportional expensive.
(Unless they are doing something like putting profits towards some sort of carbon maximization scheme)
Speaking as someone who did sys admin for a small independent cloud provider, it definitely isn't virtually 100% of operating costs
https://www.gwern.net/docs/ai/scaling/hardware/2021-jouppi.p...
One bit of feedback for the author: the sliding scale is helpful, but the y axes are different between the visualizations so you cannot see the apples to apples comparison needed. Suggest regenerating those.
There is no point to add more cores if they can't cooperate.
How come I'm the only one pointning this out?
I think 4 cores will max out the memory contention, so keep on piling these 128 core heaters on. But they will not outlive a simple Raspberry 4!?
[1] https://travisdowns.github.io/blog/2020/07/06/concurrency-co...
I know M4 had much better multicore shared memory perf. than M3, but now both of those are old and I don't have users to test anything now.
Writing us not, but respecting the single writer principle is usually rule zero of parallel programming optimisation.
If you mean reading/writing to the same memory bus in general, then yes, the bus need to be sized according to the need of the expected loads (i.e. the machine need to be balanced).