This. The talks about 30% (or even 300%!) impact based on a graph without any details or methodology, often itself based on a microbenchmark or a tweet that got viral is really not helping.
This. The talks about 30% (or even 300%!) impact based on a graph without any details or methodology, often itself based on a microbenchmark or a tweet that got viral is really not helping.
I especially liked the AWS forum post being breathlessly circulated which is from someone claiming that their business critical workload no longer fits on a single m1.medium (1 vCPU, <4GB RAM), joined by someone else who was actually swapping and hitting the OOM killer.
There is a real impact but this is going to be the lazy IT worker's new plausible excuse of the day for months…
It matters not what the workload was running on before or accusations of laziness, the key takeaway from that post was that there was an impact on performance caused by the Meltdown/Spectre mitigations.
The fact another commenter barged into the thread with an unrelated issue simply shows that AWS forum operates within the Law of Forums (which states all threads are to be cross pollinated with unrelated issues from at least one other user).
EDIT: I should say that if this fellow is the only affected user from all the computer users of the world, then the mitigations applied can be called successful. That seems unlikely :-)
In this case, the thread made it clear that the real problem was that they didn't have a good deployment story — notice how the original poster was mentioning needing to move everything to an m3.medium manually? I would be quite surprised if they haven't had other problems in the past (e.g. what happens when a system update kicks off if that server is already running at 90+% utilization? or when they get more users and/or the existing users start doing slightly more) but hadn't wanted to deal with the hassle of migrating.
This.
Sometimes this, sometimes that, it depends really.
> The talks about 30% (or even 300%!) impact based on a graph without any details or methodology, often itself based on a microbenchmark or a tweet that got viral is really not helping.
Of course it's helping, everyone learned about CPU architecture, speculative execution, caches, and started trading AMD and INTC all of the sudden.
But to be serious, doesn't "details and methodology" apply to any benchmarks we see on Twitter. Or do people automatically start buying more servers when they read that tweet. Anyone here did that?
The point of that number is warn people to go and measure their own workload. Google measure and yay! not much of an impact. It's great news. Someone who does telephony and maybe is sending lots of tiny little UDP packets back and forth might be impacted. Should they start crying and running around, no. They should measure because it might affect them.
If the usage is split 50/50 kernel/userspace. And the userspace changes slow the app down 10%. Then the user space changes slowed the userspace code down 20%. Likewise if the kernel changes slow the app down 10%, they slowed the kernel code down by 20%.
So the total impact of the changes is 20% in your example.
The high performance of the physical CPU was achieved through means that compromise security. The top of the speed range was bullshit, basically. Now that the top is being lopped off, you experience the true speed of that CPU while operating in a secure manner.
So the proper execution is slower than the secure execution, right? Ergo, it is 'slower'.
Can you share more?
[ADDED: It's back. Was pulled back to make some changes.]
> Measureable: 8-12% - Highly cached random memory, with buffered I/O, OLTP database workloads, and benchmarks with high kernel-to-user space transitions are impacted between 8-12%. Examples include Oracle OLTP (tpm), MariaBD (sysbench), Postgres(pgbench), netperf (< 256 byte), fio (random IO to NvME).
> Modest: 3-7% - Database analytics, Decision Support System (DSS), and Java VMs are impacted less than the “Measureable” category. These applications may have significant sequential disk or network traffic, but kernel/device drivers are able to aggregate requests to moderate level of kernel-to-user transitions. Examples include SPECjbb2005 w/ucode and SQLserver, and MongoDB.
It talks about a significant slowdown, 8%-19% for OLTP Workloads (tpc), sysbench, pgbench, netperf (< 256 byte), and fio (random I/O to NvME). 3%-7% for Database analytics, Decision Support System (DSS), and Java VMs
As mentioned in the article, there are 3 different (identified) vulns, each with its own set of patches, and sometimes multiples different ways to patch. Each of those patch will have its own impact, and work is currently in progress to find new ways to patch that will have a lower impact on performance.
Also, this means we will need to be used to patch, as new vulns of this category will definitely be found, and new changes to kernels/cpu/compilers will be designed and will need to be applied.