Intel Confronts Potential ‘PR Nightmare’ With Reported Chip Flaw
bloomberg.com
bloomberg.com
I know that bugs happen and that there was nothing intentional on this one, but at times like this is hard to held at bay the temptation of claiming for a class lawsuit against Intel....
This isn’t an excuse for Intel consistently having terrible verification practices and shipping horrendous hardware bugs. From 2015: https://danluu.com/cpu-bugs/ There have been more since then.
I’ve talked to multiple people who work in intel’s testing division and think “verification” means “unit tests”. The complexity of their CPUs has far surpassed what they know how to manage.
Found a quote:
"We need to move faster. Validation at Intel is taking much longer than it does for our competition. We need to do whatever we can to reduce those times… we can’t live forever in the shadow of the early 90’s FDIV bug, we need to move on. Our competition is moving much faster than we are".
Overall it’s a depressing story of predictable market failure as well as internal misbehavior at Intel, if true. Few buyers want to pay or wait for correctness until a sufficiently bad bug is sufficiently fresh in human memory. And if you do want to, it’s not as if you’re blessed with many convenient alternatives.
The same reason could have been used to give the NSA some legroom for instance, but tell everyone that's why they won't do so much verification in the future.
As other comments suggest, there might be a third stage, completely forgetting how to design and validate chips properly.
Furthermore, I just a read an article (can't find the link) that certain ARM Cortex cores have this same issues as Intel.
More likely "good enough" is much lower because ARM users aren't finding the bugs. The workloads that find these bugs in Intel systems are: heavy compilation, heavy numeric computation, privilege escalation attackers on multi-user systems. Those use cases barely exist on ARM: who's running a compile farm on ARM, or doing scientific computation on an ARM cluster, or offering a public cloud running on ARM?
Vendor, in conversation: "We're pretty sure we can make the next version do cache coherency correctly."
Me (paraphrased): "Don't let the door hit you in the ass on the way out."
Management chain chooses them anyway, I spend the next year chasing down cache-related bugs. Fun.
(I should remark that there are good reasons for this effort. Such as: It boots in under 500ms, it's crazy efficient, doesn't use much RAM, and your company won't let you use anything with a GPL license for reasons that the lawyers are adamant about).
So now you get to find all the places where the vendor documentation, sample code and so forth is wrong, or missing entirely, or telling the truth but about a different SOC. You find the race conditions, the timing problems, the magic tuning parameters that make things like the memory controller and the USB system actually work, the places where the cache system doesn't play well with various DMA controllers, the DMA engines that run wild and stomp memory at random, the I2C interfaces that randomly freeze or corrupt data . . . I could go on.
It's fun, but nothing you learn is very transferrable (with the possible exception of mistrust of people at big silicon houses who slap together SOCs).
There are hardware manufacturers that are better than others at being open and providing documentation. My minimal level of required support and documentation right now is mainline linux support.
Can you document your work publicly, or is there something I can read about it? I'm very interested in alternative kernels beside Linux.
When you buy an SOC, the /contract/ you have with the chip company determines the extent and depth of their responsibility. On the other hand, they do want to sell chips to you, hopefully lots of them, so it's not like they're going to make life difficult.
Some vendors are great at support. They ship you errata without you needing to ask, they are good at fielding questions, they have good quality sample code.
Other vendors will put even large customers on a tier-1 support by default, where your engineers have to deal with crappy filtering and answer inane questions over a period of days before getting any technical engagement. Issues can drag on for months. Sometimes you need to get VPs involved, on both sides, before you can get answers.
The real fun is when you use a vendor that is actively hiding chip bugs and won't admit to issues, even when you have excellent data that exposes them. For bonus points, there are vendors that will rev chips (fixing bugs) without revving chip version identifiers: Half of the chips you have will work, half won't, and you can't tell which are which without putting them into a test setup and running code.
My favorite ARM experience was where memcpy() was broken in an RTOS for "some cases". "some cases" turned out to be when the size of the copy wasn't a multiple of the cache line size. Scary stuff.
> The AMD microarchitecture does not allow memory references, including speculative references, that access higher privileged data when running in a lesser privileged mode when that access would result in a page fault.
Out-of-order processors generally trigger exceptions when instructions are retired. Because instructions are retired in-order, that allows exceptions and interrupts to be reported in program order, which is what the programmer expects to happen. Furthermore, because memory access is a critical path, the TLB/privilege check is generally started in parallel with the cache/memory access. In such an architecture, it seems like the straightforward thing to do is to let the improper access to kernel memory execute, and then raise the page fault only when the instruction retires.
Agreed!
(We should probably also stop overgeneralizing about the nature of computational workloads.)
Incorrect. It also affects interrupts and (page) faults.
Any usermode to kernel and back transition.
Hosting on bare metal will become more attractive. Too bad you can't long OVH and Hetzner.
What does that even mean?
Also Hetzner just introduced some AMD Epyc server.
As opposed to "shorting" a stock, which means making a bet that it will go down in value.
HN doesn't let you do this to new comments to avoid back-and-forth commenting that is typical in flamewars.
You can reply anyway, but you have to click on the timestamp ("X minutes ago") to do it.
The other benchmark that has generated some consternation is running 'du' on a nonstop loop.
Both of these situations are pathological cases and don't reflect real-world performance. My guess is a 5-10% performance hit on general workloads. Still significant, but nowhere near as bad as some of the numbers that are getting thrown around.
And, databases are the worst case scenario, most real-world applications are showing 1% performance impact or less.
https://www.computerbase.de/2018-01/intel-cpu-pti-sicherheit...
https://www.hardwareluxx.de/index.php/news/hardware/prozesso...
Your last link is all gaming benchmarks, which as the article mentions are not affected much.
[1] http://lkml.iu.edu/hypermail/linux/kernel/1801.0/01274.html [2] http://lkml.iu.edu/hypermail/linux/kernel/1801.0/01299.html
Not to be mean, but that's not what is being changed.
You're right on the bug - userlevel code can now read any memory regardless of privilege level. However the fix isn't to manually check the privileges on each access - that would be extremely slow and wouldn't actually fix the problem.
The fix is to unmap the kernel entirely when userspace code is running. Because the kernel will no longer be in the page-table, the userspace code can no longer read it. The side-effect of this is that the page-table now needs to be switched every-time you enter the kernel, which also flushes the TLB and means that there will be a lot more TLB misses when executing code, which slows things down a lot.
So, to be clear, it is not accessing pages that is being slowed down, it is the switch from the kernelspace to the userspace.
All your points are right though. Page access times will in general be slower because of all the extra TLB flushes, leading to more TLB misses when accessing memory.
And how often the kernel services interrupts.
Or am I completely off the mark?
If people who received written assurance from Intel that their hardware is 100% bug free can form a legal class, sure. I highly doubt there is even a single one such customer.
Yes and no. Yes, Intel would get a chance to claim that the case should be dismissed out of hand. To do that, they have to prove that, even assuming all the claimed facts are true, the people suing still don't have a valid case. That's a high bar. It can be reached - there's a reason that preliminary summary judgment is a thing in court cases - but it takes a really flawed case to be dismissed in this way.
How flawed? SCO v. IBM was not completely dismissed on preliminary summary judgment, and that was the most flawed case I've ever seen.
> It may very well be a question of who has the better legal team.
Well, Intel can afford to hire the best. A huge class-action suit can sometimes attract the best to the other side as well, though. (There's not just one "best", so there's enough for both sides of the same court case.)
IANAL, but it looks to me like there's at least the potential for a valid court case. CPUs are (approximately) priced according to their ability to handle workloads; if they can't provide the advertised performance, they didn't deserve the price they sold for.
What I meant was that the presence of the bug itself is not a valid cause, for example you can't claim that due to the error you lost 1 trillion dollars via a software hack - even if it's true. If Intel can prove they acted ethically when disclosing the bug and that they replaced / compensated users up to the value of the CPU, they are in the clear.
The question is did anyone receive performance assurance from Intel? Probably not.
Some cloud providers or compute grids just lost a lot. Maybe they will find an angle to claim compensation.
Well, if I rent a VPS with x performance, I still expect x performance after this flaw is patched. The company providing the virtual machine will perhaps have to pay 30% more to provide me with the same product I've been getting.
Since most VPS offerings arbitrage shared resources, this will not increase costs of providing VPSes by the full performance penalty.
So you may suddenly find that your own performance requirements, that were previously satisfied by "2x m5.xlarge" are no longer being met by that configuration, and I doubt AWS will just provide you with more resources at no additional charge.
Are there any providers that state you will get x performance? Most that I've seen say you will m processors, n memory, and p storage but don't make any guarantees about how well those things will perform.
On the other hand, shrinking Intel's market share due to bad PR and thus adding some competition into the industry could actually foster that progress.
The bigger issue is for things that don't scale easily. That sql server that was at 90% capacity is suddenly unable to handle the load. Sure that could've happened organically, but now it happens (perhaps literally) overnight for everyone all at once.
Expect a bunch of outages in the next few weeks as companies scramble to fix this.
It seems you need root or physical access to the system as a prerequisite for the attack.
Where that gets tricky is when everyone's using cloud hosting solutions where the physical machines are abstracted away, and a given physical server may be running multiple virtual servers for different customers.
Think of it like this:
* Somewhere in a data center at a cloud provider is a physical server, wired up in a rack..
* That server runs virtualization software, allowing it to host Virtual Server 1, Virtual Server 2, and Virtual Server 3.
* Virtual Server 1 belongs to Customer A. Virtual Servers 2 and 3 belong to Customer B.
* Normally, Virtual Server 1 can't access any memory allocated to Virtual Servers 2 and 3.
* BUT: Customer A can now use Meltdown to read the entire memory of the physical server. Which includes all the memory space of Virtual Servers 2 and 3, exposing Customer B's data to Customer A.
That's the threat here.
Just wanna point out that a 30% performance hit means a 43% cost increase.
For those confused: the math here is a 30% decrease puts you at 70%. To go from 70% back to 100%, 30% only gets you to 91% (0.70*1.3). 1/0.7 = 1.43 means you need 43% to recover.
Our ElasticSearch nodes all had 32GB of ram and we had 10 of them and they were all being pushed to the max.
Something like this would be a massive hit, requiring a lot more work into identifying new bottlenecks and scaling up appropriately.
It's, however, really bad if you sell CPU cycles for a living. You just lost between 5 and 30% of your capacity. If you have a large building, you just lost part of your parking lot to the Intel Kernel Page Problem building.
Honestly, I'd just make sure the server firewalls are super tight and not take in the future patches. At least for now.
I know I’m not alone.
Then again. Think of microservices, Kubernetes for instance; Network requests are system calls.
(In my case anyway)
30% overhead might be inscentive to revisit the assumption we can’t rewrite it for Linux.
I don’t know very much about computing on that scale, but I wonder if all the people selling off Intel stock are thinking this story through.
It's possible that the patches applied to fix this bug will cause some single-threaded benchmarks to change from Intel being the fastest to AMD being the fastest.
Not trying to kill expectations. This decision isn’t mine alone. You know the old saying “nobody got fired for buying Cisco” that applies to Intel too.
That's a good description of basically every cloud environment out there, from AWS on down.
In other words they are extremely common.
We'll start to get conscious about the number of syscalls we use on each operation, start using large buffers, start buffering stuff user-side...
Good security is about layers. No one layer can be assumed to be watertight, but with enough layers you hopefully get to a good place.
What do you mean by "compressible"?
Kernel ABIs will eventually reflect that and crop up higher level expensive calls that replace groups of currently cheap syscalls (that will become expensive after the fix).
And Intel will profit handsomely from next generation CPUs that'll get an instant up-to-30% performance boost for fixing this bug.
Who really sells CPU cycles? Cloud providers sell instances priced per core. So the real hit is by the customers since they have to shell out for more instances for the same amount of computing power.
The hit I see is by providers of 'serverless' computing, since they charge per request and have their margins reduced.
AWS, Azure, and GCP all bill serverless with a combination of per-request fees and compute (GB-seconds), so I'd expect the entire hit to be passed on to the user since this will cause increased compute time for each request. N requests that used to average 300ms each will now be N requests that average, say, 400ms, so the per-request billing remains the same and the compute billing will increase by approximately 30%.
30% is a big hit. I'm wondering if that isn't a bit exaggerated, or perhaps the consequence of a poorly optimized workarounds that will rapidly improve. I recall seeing figures on the order of 3% only a few days ago.
How big it will be for your workload is a function of what your workload is. Benchmark if it is important to you.
Also a 30% decrease is also equivalent to setting Moore's law back 7 months. A 5% loss is only setting it back 1 month. I know that's a bit of a naive calculation. But the point is computing power has long operated in an exponential domain. So big differences in absolute numbers aren't necessarily a big deal.
Maybe a lot more now?
Well... Everyone who bought AMD. Some people managed to see beyond the hype and go for the optoon that made sense.
What hype are you referring to? Are you suggesting the people who bought AMD knew this was a problem for Intel?
True, sometimes you will leave boxes at low utilisation for various reasons, e.g. to deal with traffic spikes. But those reasons have not gone away. So now instead of heaving a predictable increase in CPU cost, you have an unpredictable increase in performance snafus.
The only good news is that the real performance hit will be less than 30% on many workloads. Especially once the providers start juggling and optimising.
TL:DR - queue behaviour gets nonlinear as you approach the theoretical max load. If you are running your processors at a high load, even a small change in code throughput makes a huge difference to real world behaviour.
If it, like it seems, is just an attack on OS kernels and PV hypervisors, you can simply turn off the mitigation, since nowadays kernel security is mostly useless (and Linux is likely full of exploitable bugs anyway, so memory protection doesn't really do that much other that protecting against accidental crashes, which isn't changed by this).
Even if it's an attack against hypervisors any large deployment can simply use reserved machines and it won't have a significant cost.
sudo cat /dev/mem
?I'm having a hard time understanding why this is worse than any other local root escalation bug except for the consequences of the necessary patch.
EDIT: I see that /dev/mem is no longer a window on all of physical RAM in a default secure configuration. Is it true that there's no way for root to read kernel memory in a typical Linux instance? If so, the severity of this issue makes more sense to me.
> I'm having a hard time understanding why this is worse than any other local root escalation bug except for the consequences of the necessary patch.
It's not, as far as I'm aware. The fact that the patch has perf consequences is why it's such a big deal.
a) It would allow any non-root process to read full memory, including the kernel and other processes, or
b) It would allow one cloud VM to read full memory of other cloud VMs on the same physical machine, or
c) With enough cleverness, it would allow even sandboxed Javascript on a web page to read full memory of the computer that it is running on.
The first "instruction" of your program is the last address on the stack, in the list of addresses you pushed to the stack.
You are executing code, but you did not inject any executable code, you did not need to modify any existing code pages (which are probably read only), you did not need to attempt to execute code out of a data page (which is probably marked non executable).
Address Space Layout Randomization is a way to prevent the "return oriented programming" attack. When a process is launched, the address space is randomly laid out so that the attacker cannot know which address in memory the std C lib printf function will be located at -- in this process.
Now let's think about the kernel. If you could know all of the addresses of important kernel routines, you could potentially execute a "return oriented programming" attack against the kernel with kernel privileges. Without modifying or injecting any kernel level code. These hardware vulnerabilities allow user space code to deduce information about kernel space addresses.
Now that's a lot of hoops to jump through in order to execute an attack. But there are people prepared to expend this and even more effort in order to do so. Well funded and well staffed adversaries who would stop at nothing in order to access more and better pr0n collections.
> If you could know all of the addresses of important kernel routines, you could potentially execute a "return oriented programming" attack against the kernel with kernel privileges. Without modifying or injecting any kernel level code.
The user <-> kernel transition is mediated (on x86-64) with the SYSCALL instruction, which jumps to a location specified by a non-user writable MSR. How does return-oriented programming work in that case?
The crucial point here being that there must already be an existing overflow vulnerability in the kernel. Knowing all the addresses is no use if you can't force execution to go to them.
AIUI, the present circumstances are:
- there exists a public PoC from some researchers of side-channel leaking kernel address information into userland via JavaScript which may be unrelated
- there exists a Xen security embargo that expires Thursday that might be unrelated
- AWS and Azure have scheduled reboots of many things for maintenance in the next week, which seems unlikely to be unrelated to the Xen embargo
- a feature that appears to be geared toward preventing a side-channel technique of unknown power has been rushed into Linux for Intel-only (both x86_64 and ARM from Intel)
- a similar class of prevention technique has been landed in Windows since November for both Intel and AMD x86_64 chips (no idea about ARM)
- the rush surrounding this, and people being amazingly willing to land fixes that imply a 5-30% performance impact, strongly suggest that unlike almost every major CPU bug in the last decade, you can't fix or even work around this with a microcode update for the affected CPUs, which is _huge_. The AMD TLB bug, the AMD tight loop bug that DFBSD found, even the Intel SGX flaws that made them repeatedly disable SGX on some platforms - all of them could be worked around with BIOS or microcode updates. This, apparently, cannot. (Either that or they're rushing out fixes because there's live exploit code somewhere and they haven't had time to write a microcode fix yet, but O(months) seems like they probably concluded they outright can't, rather than haven't yet.)
EDIT: This post[2] discusses the specific speculative execution cache attack and claims there is a JavaScript PoC (but doesn't cite a source for that claim)
[1] https://www.youtube.com/watch?v=ewe3-mUku94
[2] https://plus.google.com/+KristianK%C3%B6hntopp/posts/Ep26AoA...
Also, RUH-ROH. https://twitter.com/brainsmoke/status/948561799875502080
- Intel issued a press release saying they planned to announce this next week after more vendors had patched their shit, which lends me more cause to believe that the Xen bug might be the same one [1]
- Intel claims in the same PR that "many types of computing devices — with many different vendors’ processors" are affected, so I'll be curious to see whether non-Intel platforms fall into the umbrella soon
- macOS implemented partial mitigations in 10.13.2 and apparently has some novel ones coming up in 10.13.3 [2]
- someone reasonably respected claims to have a private PoC of this bug leaking kernel memory [3]
- ARM64 has KPTI patches that aren't in Linus's tree yet [4] [6] ([6] is just a link showing the patches from 4 aren't in Linus's tree as of this writing)
- all the other free operating systems appear to have been left out of the embargoed party (until recently, in FBSD's case), so who knows when they'll have mitigations ready [5]
- So far, Microsoft appears to have only patched Windows 10, so it's unknown whether they intend to backport fixes to 7 or possibly attempt to use this as another crowbar to get people off of XP 2.0
- Update: Microsoft is pushing an OOB update later today that will auto-apply to Win10 but not be forced to auto-apply on 7 and 8 until Tuesday, so that's nice [7]
[1] - https://newsroom.intel.com/news/intel-responds-to-security-r...
[2] - https://twitter.com/aionescu/status/948609809540046849
[3] - https://twitter.com/brainsmoke/status/948561799875502080
[4] - https://patchwork.kernel.org/patch/10095827/
[5] - https://lists.freebsd.org/pipermail/freebsd-security/2018-Ja...
[6] - https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
[7] - https://www.theverge.com/2018/1/3/16846784/microsoft-process...
Seems that Google/Project Zero felt the need to go ahead and break embargo. Worth adding to the above list of news sources.
If you read the article you quoted:
> We are posting before an originally coordinated disclosure date of January 9, 2018 because of existing public reports and growing speculation in the press and security research community about the issue, which raises the risk of exploitation. The full Project Zero report is forthcoming (update: this has been published; see above).
Just from public Gooogling, I believe it may have been the Register who tried to get in on the scoop, and broke the embargo:
https://www.theregister.co.uk/2018/01/04/intels_spin_the_reg...
My checking doesn't show any of those three explicitly listed in Apple's security updates up through 10.13.2/2017-002 Sierra.
Are those kernel logical addresses?
ASLR, PIC (position independent code: chunks of the binary move around between executions), and RELRO (changing the order and epermissions of an ELF binaries headers: a common ROP pattern is to set up a fake stack frame and call a libc function in the ELFs Global offset table) are all mitigations against ROP, but none solve the underlying problem.
The reason ROP exists is that x86-64 use a Von Neumann architecture, which means that the stack necessarily mixes code (return addresses) and data. The only true solution is an architecture that keeps these stacks separate, such as Harvard architecture chips.
As for bypassing the aforementioned mitigations...
ASLR: Only guarantees that the base address changes. Relative offsets are the same. So to be able to call any libc function in a ROP chain, all you need is a copy of the binary (to find the offsets) and to leak any libc function address at runtime. There are a million ways for this data to be leaked, and they are often overlooked in QA. Once you have any libc address, you can use your regular offsets to calculate new addresses.
PIC: haven't yet dealt with it myself, but you can use the above technique to get addresses in any relocated chunk of code, but I think you'll need to leak two addresses to account for ASLR and PIC.
RELRO: This makes the function lookup table in the binary read only, which doesn't stop you from calling any function already called in the binary. Without RELRO, you can call anything in libc.so I think, but with RELRO you can only call functions that have been explicitly invoked. This is still super useful because the libc syscall wrappers like read() and write() are extremely powerful anyway. Full RELRO (as opposed to partial RELRO) makes the procedure linkage table read only as well, which makes things harder still.
If this is the kinda thing that interests you, I heartily recommend ropemporium.com which has a number or ROP challenge binaries of varying difficulty to solve. If you're not sure where to start, I also wrote a write-up for one of the simpler challenges [1] that is extremely detailed, and should be more than enough to get you started (even if you have me experience reversing or exploiting binaries)
Disclaimer: I'm just some dipshit that thinks this stuff is fun, if I've made a mistake in the above please let me know. I also haven't done any ROP since I wrote the linked article, so im probably forgetting stuff.
[1] https://medium.com/@iseethieves/intro-to-rop-rop-emporium-sp...
This seems very wrong. I'm not aware of any privilege isolation in Windows relying on the secrecy of any value. Security tokens have opaque handles for which "guessing" makes no sense. Are you aware of anything?
1. Read the root ssh private key from the openssh deamons kernel pages maintaining the crypto context and ssh into the system
2. Read a sudo auth key generated for someone using sudo and then use that to run code as a root user
3. Read the users password's whenever a session manager asks the users to reauth
4. If running in AWS/GCP inside a container/vm meant to run untrusted code, read the cloud provider private keys and get control on account
5. RCE to ROP powered privilege escalation exploit seems reasonable...
6. Rowhammer a known kernel address (since you can now read kernel memory) to flip some bits to give you root
Also remember running JS is basically RCE if you can read outside the browser sandbox, ads just became much more dangerous...
Incidentally, this seems to indicate that zero-copy I/O is actually a security improvement as well, not just a performance improvement?
I am not really sure how/if zero copy may/may not solve this problem.
If this bug only allows reading kernel pages, zero copy may actually help if the unprivileged user can't read your pages, but from the small amount of available description it looks like it can read any page, but kernel pages are more interesting because thats a ring lower and which is why all the focus is on that.
I am fairly certain there is more protection against being able to read memory owned by process on a lower ring level so zero copy may be a bad idea for security critical data.
And based on the disclosure that google published, looks like any memory can be read
Based on all the hoopla around the linux kernel patches the thinking is : yes it can. Or VM escape. Or both.
I'm cursed when it comes to timing. It's like when I bought that house in 2007, held onto it waiting for the market to recover, then tried to sell it only to find out my tenants had been using it to operate a rabbit-breeding business for years and completely trashed the place (thank you, useless property manager), forcing me to sell it at a loss anyway (6 months ago).
Also, I hate rabbits now. And I veered off topic, sorry.
Well I guess you're not the right person to without about a great ninja-rockstar position at our new RaaS startup.
/one has to joke sometimes to avoid crying over taking a 30% hit in costs... over a stupid CPU bug
The thing is, AMD probably very narrowly just missed this one --- if they did more aggressive speculative execution, they would be the same.
>>attack that would be almost expected in a processor with speculative execution unless special measures were taken to prevent it.
if you're going to put in features with expected attacks you should definitely be putting in features to prevent it , and if it is an expected attack it shouldn't be special measures it should just be an inherent part in introducing the feature.
Doesn't multi-user timesharing and virtualization predate every modern CPU and OS though?
At first, computers were very expensive, and so were shared between many users. Mainframes, UNIX, dumb terminals, etc.
Then computers became cheap. Users could each have their own computer, and simply communicate with a server. Each business could have their own servers co-located in a datacenter.
Then virtualization got really good, and suddenly cloud servers became viable. You didn't have to pay for a whole server all the time, and if demand rapidly increased you didn't need to buy lots of new hardware. And if demand decreased you didn't get stuck with tons of useless hardware.
The second stage (dedicated servers) was the case when speculative execution was implemented. We're currently in the third stage, but Intel haven't changed their designs.
Indeed, this reminds me of cache-timing attacks, which probably can be done on every CPU with any cache at all --- and they've never seemed to be much of a big deal either.
I’m wondering whether ARM chips are affected if they are whether they are uniformly affected or whether it depends on vendor implementation choices.
Ironically, in human populations it produces the opposite effect.
https://news.ycombinator.com/newsguidelines.html
Edit: since https://news.ycombinator.com/item?id=16063749 makes it clear that you're using HN for that purpose, which is not allowed here, I've banned this account. Would you please not create accounts to break the site guidelines with?
I'm also wondering if/hoping for a fix that involves increased memory usage instead of the speed.
(I apologise if this is blindingly obvious for somebody well versed in low-level programming.)
https://www.phoronix.com/scan.php?page=article&item=linux-41...
And while databases try to minimize the number of syscalls they still end up doing a lot of them for read, writeout, flush.
But this will require you to have the right kind of flash storage, right kind of fs, right kind mount options, and probably a different code path in userspace for DAX vs traditional storage.
So we're a little ways away from this.
That doesn't move anything from kernel land into userspace, certainly not in the app's process in userspace anyway.
There's no talk in the DAX information about how this results in a zero-syscall filesystem API, and I'm not seeing how that would ever work given there would then be zero protections on anything. You need a handle, and that handle needs security. All of that is done today via syscalls, and DAX isn't changing that interface at all. So where is the API to userspace changing?
This work is experimental but you can mmap a single file on a filesystem on this device using new DAX capabilities. Most access will not longer require a syscall.
This comes with all the usual semantics and trappings of mmap plus some additional caveats as to how the filesystem / DAX / hardware is implemented. Most reads/writes will not require a trip to the kernel using the normal read()/write() syscalls. Additionally, there is no RAM page cache baking this mmap instead the device is mapped directly at a virtual address (like DMA).
Finally, flush for these kinds of devices is at the block level implemented using normal instructions and not fsync. Flush is going to be done using the CLWB instruction. See: https://software.intel.com/en-us/blogs/2016/09/12/deprecate-...
LWN.net has lots of articles and links in their archives from 2016/2017. It's a really good read. Sadly I do not have time to dig more of them up for you. Do a search for site:lwn.net and search for DAX or MAP_DIRECT.
As in, if you call read/write instead of using mmap you're still getting a syscall regardless of if DAX is supported or not. Not everything can use mmap. mmap is not a direct replacement for read/write in all scenarios.
The reply to that comment is accurate: that's a pathological case. Probably an order of magnitude off.
https://www.postgresql.org/message-id/20180102222354.qikjmf7...
Real-world use cases introduce much more latency from other sources in the first place.
I'm sticking with an expectation in the 2%-5% range.
In the real world, where code is doing real things besides just entering/exiting itself all day, I think it's going to be a stretch to see even a 5% performance impact, let alone 10%.
But overall, yeah.
The reality is that OLTP databases execution time is not dominated by CPU computation but instead of IO time. Most transactions in OLTP systems fetch a handful of tuples. Most time is dedicated to fetching the tuples (and maybe indices) from disk and then sending them over network.
New disk devices lowered the latency significantly while syscall time has barely gotten better.
So in OLTP databases I expect the impact to be closer to 10% to 15%. So up to 3x over the base case.
The first set of numbers isn't actually unrealistic. Doing lots of primary key lookups over low latency links is fairly common.
The "SELECT 1" benchmark obviously was just to show something close to the worst case.
Latency through loopback on my machine takes 0.07ms. Latency to the machine sitting next to me is 5ms.
We're actually (and to think, today I trotted out that joke about what you call a group of nerds--a well, actually) talking multiple orders of magnitude through which kernel traps are being amplified.
Uh, latency in local gigabit net is a LOT lower than 5ms.
> We're actually (and to think, today I trotted out that joke about what you call a group of nerds--a well, actually) talking multiple orders of magnitude through which kernel traps are being amplified.
I've measured it through network as well, and the impact is smaller, but still large if you just increase the number of connections a bit.
Intel has already dropped and AMD is up. Maybe there's more to move, but first-order effects are at least partially priced in already.
But what about second-order effects? Seems like virtualization should be vulnerable (VMWare and Citrix), but maybe they actually benefit as customers add more capacity.
Software-defined networking and cloud databases should also suffer though it's unclear how to trade these.
AWS, Google Cloud and Azure might benefit as customers add capacity but there's no way to trade the business units. So what about cloud customers where compute costs are already a large percentage of total revenue?
Netflix should be OK but Snap and Twilio could get squeezed hard. Akamai and Cloudflare might have higher costs they can't pass through to customers.
And where's the upside? Who benefits? If the performance hit causes companies to add capacity, maybe semiconductor and DRAM suppliers like Micron would benefit.
Add on top of that the fact that a lot cloud customers over provision (there's good scientific papers on how much spare CPU capacity there is). Cloud service providers that sell things on a per request / real CPU usage model (vs reserved capacity) prob benefit more.
Also, you can't just separate trading in AWS or GCE from the rest of the core business.
Potentially business units of DELL, HP, IBM, ... should do better as people use this as a justification to upgrade overdue hardware they should cover 5% to 10% lower performance (needing more units to cover that).
I owned Intel back in 2010 when they bought McAfee for 7.8 billion. They said the future of CPU's and chip tech was embedding security on the chip. the real answer was mobile and gpus.
Not only did I immediately know this was a horrendous deal, it clearly showed that the CEO and management had no clue on their own market's desires and direction. At the time, I was hoping they were going to buy Nvidia, it would have been a larger target to digest at 10 bil, but doable by Intel at the time.
The MacAfee purchase turned out to be one of the worst large cooperate purchases in history. Had they invested the 7.8 billion $ blindly into an sp500 index fund, their investment would be worth ~19-20 billion.
At least we can play our sorrows away.
This sounds like it’s positively evil for outfits that rely heavily on virtualisation also.
There isn't a guarantee that will compensate for that any more than if you updated some piece of your software infrastructure to a new version that just got slower.
--Oracle's marketing tomorrow, probably
(to their credit, SPARC does fully isolate kernel and user memory pages, so they were ahead of the curve here... for all 10 of their users who run anything other than Oracle DB on their systems.)
This is a all hands on deck kind of situation. Apple doesn't usually do well with security firedrills like this.
The most obvious issue with this benchmark is that Phoronix is testing the latest rcs, with all of their changes, against the last stable version [EDIT: I misread or this changed overnight, see below] that doesn't have PTI integrated, instead of just isolating the PTI patchset. The right way to do this would be to use the same kernel version and either cherry-pick the specific patches or trust that the `nopti` boot parameter sufficiently disables the feature. That alone makes the test worthless.
There is no way this causes a universal 30% perf deduction, especially not for workloads that are IO-bound (i.e., most real-world workloads). This is a significant hit for Intel, but it's not going to reduce global compute capacity by 30% overnight.
EDIT: Looking at the Phoronix page, the benchmark actually appears to use 4.15-rc5 as "pre" and 4.15-some-unspecified-git-pull-from-Dec-31-that-isn't-called-rc6 as "post". I thought I had read 4.14.8 there last night, but may not have. Regardless, the point stands -- these are different versions of the kernel and the tests do not reflect the impact of the PTI patchset.
I'm saying that it's not a reliable measurement of the impact of the PTI patchset. There was a PgSQL performance anecdote [0] (actually tested with the real boot parameters instead of entirely different versions of the kernel) that showed 9% performance decrease posted to LKML, which Linus described as "pretty much in line with expectations". [1]
Quoting further from that mail:
> Something around 5% performance impact of the isolation is what people are looking at.
> Obviously it depends on just exactly what you do. Some loads will hardly be affected at all, if they just spend all their time in user space. And if you do a lot of small system calls, you might see double-digit slowdowns.
So in general, the hit should be around 5%, and "[y]ou might see double-digit slowdowns" seems like the hit on a worst-case workload is hovering closer to the 10% range than 30%. That's also what the anecdote from LKML shows, unlike Phoronix which shows 25%-30% or worse.
This is more of an attrition thing than a staggering loss. With people saying MS patched this in November, it would be interesting to see if people saw a similar 5-10% degradation in Windows benchmarks since that time.
>How often do companies release performance downgrades of that scale?
I don't know which "company" you're referring to here, but substantial changes in kernel performance characteristics are pretty common during the Linux development/RC process, and yes, definitely some workloads will often see changes +/- 10% between the roughly bi-monthly stable kernel releases.
If you're surprised that Linux development is so "lively", you're not alone. That's one of the selling points of other OSes like FreeBSD.
I'm sure there are frantic emails claiming that AMD shouldn't be punished for Intel's mistake.
EDIT: actually the fix will go out with 4.14.12 and 4.15rc7, both `X86_BUG_CPU_INSECURE` and AMD's addendum to be protected from the collateral damage.
One thing I can definitely respect about the kernel developers is that they don't seem to make any effort to be nice about the fact that they need to deal with undocumented nonsense from vendors all round.
--- a/arch/x86/include/asm/processor.h
+++ b/arch/x86/include/asm/processor.h
+ * On Intel CPUs, if a SYSCALL instruction is at the highest canonical
+ * address, then that syscall will enter the kernel with a
+ * non-canonical return address, and SYSRET will explode dangerously.
+ * We avoid this particular problem by preventing anything executable
+ * from being mapped at the maximum canonical address.
+ *
+ * On AMD CPUs in the Ryzen family, there's a nasty bug in which the
+ * CPUs malfunction if they execute code from the highest canonical page.
+ * They'll speculate right off the end of the canonical space, and
+ * bad things happen. This is worked around in the same way as the
+ * Intel problem.I'm having trouble finding a good reference right now.
Seems like very little got through to the media about the details regarding this flaws effects and costly workaround.
I actually think it will - it would be easy to give more accurate details that cause many readers to glaze over.
> How to ELI5 the risk in desktop PC? A piece of JavaScript in some 0x0 pixel iframe in a tab you're not even looking at stealing your passwords and SSH keys
(Although nothing is proven in the latter regard, I wouldn't be per se surprised to see something in that direction once the exact nature of the issue is more widely known)
People keep repeating this claim because it sounds dramatic, but I'm not sure it's a fair description. The original source appears to be a single snide tweet from @grsecurity [1] referencing this comment [2].
It's far from obvious that the comment was even "redacted" at all. It seems more likely that "stay tuned" is either a reference to the more detailed comments elsewhere in the patch (in arch/x86/kernel/ldt.c), or a reflection of the fact (which is clearly spelled out in the commit message) that future patches are likely to change the location of the LDT mapping.
I've skimmed through the commit messages and comments from the latest patchset [3] and couldn't find anything else that even hinted at redaction, nor could I find any mention of redactions on the linux-kernel mailing list.
Furthermore, it's worth bearing in mind that @grsecurity has been involved in numerous public feuds with the Linux security folks. So in the absence of concrete evidence, I'm not particularly inclined to assume his tweet was made in good faith.
[1]: https://twitter.com/grsecurity/status/947147105684123649
[2]: https://github.com/torvalds/linux/commit/f55f0501cbf65ec41cc...
Excuse me while I go and invest in a company making rubber underwear.
Writes someone who's never seen an errata.
I think you took it a bit too far. May be, I am missing something. Is it really that bad?
[0]: https://www.phoronix.com/scan.php?page=article&item=linux-41...
This is an issue, but laypeople are overblowing the effect on their everyday computing.
[0]https://gist.github.com/woachk/2f86755260f2fee1baf71c90cd653...
It's mostly inelastic. People don't buy resources unless they need to do work. Alternatives, such as on-prem hardware, are affected as well.
The only work that would be affected would be the work that becomes unprofitable at a 10-30% hardware cost increase and this is probably marginal enough to be ignored.
I work on a system (at Google) with significant hardware cost. Fundamentally, the work my system is doing needs to be done. But the time we spend improving efficiency is hugely elastic. I look at a CPU profile or request flow, have an idea to improve it, look at how much machine resources we're spending, guesstimate how long it will take us to implement and maintain, and use a chart to see if my idea is worth my team's time or not.[1]
If the resources get more expensive or simply impossible to acquire,[2] we'll optimize more.
[1] There are other considerations (will it make the code more complex, thus increasing risk of a security/privacy flaw? what about opportunity cost given that it's hard to hire more engineers and scale a team?) but that's a reasonable mental model.
[2] There's a global RAM and SSD manufacturing crunch already, and if lots of CPUs are replaced due to these vulnerabilities or simply are no longer enough, that's gonna be a big crunch as well. If you're one of the few biggest cloud providers, I think you can't just replace / add tons more hardware than planned without dramatically increasing the price per unit for everyone.
In the context of the above posts, "elastic" is an economic term used to refer to the demand curve. There is no supply/demand curve when it comes to internal technical decisions for optimization =)
I would say your work is highly correlated to the price of hardware per some unit of performance AND the amount of work that you need to complete AND the amount of hardware already available AND the cost of labor needed for optimization.
I imagine in your case, since these systems are tightly controlled, you can probably run unpatched without taking a substantial risk.
Cloud is probably a fraction of Google’s, IBM’s, Oracle’s and MS’s profits. It’s probably a significant part of Amazon’s profits but Amazon doesn’t trade on earnings.
If your cloud provider lowers its performance/efficiency by 30%, isn't it up to its customers to switch? But where to switch too? There's no choice among the big cloud providers. They all use Intel and AMD in varying amounts. As a cloud user you are stuck with Intel or AMD chips. Demand doesn't change, supply is lowered.
Looking at it like this, it is more likely that demand for cloud resources will rise, and as such revenue of the cloud providers will as well.
Those 5--7% will make me feel like I'm powered by Atom and it will be even worse for single-core bounded workloads in the cloud.
Virtualization IS real-world-usage. This is going to damage Intel where it will hurt the most, the datacenter (which is largely virtualized using VMware, or Hyper-V). The Xeon CPU's have some of the best profit margins for Intel. If they erode away to Epyc (which is finally becoming available) this could be pretty good for AMD's, espically since AMD has said the the past few years there strategy is to go after the datacenter market.
Maybe why its a favorite of /r/wallstreetbets
Guilty as charged.
>I'd be panic buying / selling every day to adjust which seems stressful...
Just treat it as numbers on a screen rather than money. Keeps the emotions at bay.
Research suggests otherwise:
https://personal.vanguard.com/pdf/ISGDCA.pdf
Diversification is a more effective strategy for dealing with volatility.
Sometimes it's wiser to take into account human psychology, even if academics are out there with data to convince us otherwise.
I'm personally with the CEO of Vanguard on this one. A short entrance into the market, especially in today's situation is probably the way to go. People can do whatever they want, and if they put their money where their mouth is and go in with a lump sum, then I'll respect it.
Otherwise, as someone who is bringing a lot of money myself onto the market now, I'm DCAing.
If you can train yourself to do the opposite you could make a lot off this stock. Buy fear, sell greed. Never panic.
Losses are never realized if you don't sell, and just wait it out until it is back near 15 again.
Wrong way to invest. Don't buy an individual company's stock unless you are ready to (1) hold for the long-term and (2) ignore (short-term) unrealized losses. Especially in the first year after investing in an individual stock it is typical to see an unrealized loss, because gains require time to accumulate.
The stock market is not a slot machine. Investing requires patience and discipline.
Also, historically the stock market as a whole has outpaced inflation, so even just investing in randomly chosen companies should have gains over time (there is even evidence that a portfolio of randomly chosen companies performs about the same as a portfolio of companies chosen by analysts and advisers).
Edit: I'm (also) wondering about Windows, in case anyone knows yet.
+#ifdef CONFIG_PAGE_TABLE_ISOLATION
+# define DISABLE_PTI 0
+#else
+# define DISABLE_PTI (1 << (X86_FEATURE_PTI & 31))
+#endif
PS - MSFT has not published relnotes, so we do not know yet. We'll find out soon enough. +void __init pti_check_boottime_disable(void)
...
+ ret = cmdline_find_option(boot_command_line, "pti", arg, sizeof(arg));
+ if (ret > 0) {
+ if (ret == 3 && !strncmp(arg, "off", 3)) {
+ pti_print_if_insecure("disabled on command line.");
+ return;
+ }
+ if (ret == 2 && !strncmp(arg, "on", 2)) {
+ pti_print_if_secure("force enabled on command line.");
+ goto enable;
+ }
+ if (ret == 4 && !strncmp(arg, "auto", 4))
+ goto autosel;
+ }
+
+ if (cmdline_find_option_bool(boot_command_line, "nopti")) {
+ pti_print_if_insecure("disabled on command line.");
+ return;
+ }
+
+autosel:
+ if (!boot_cpu_has_bug(X86_BUG_CPU_INSECURE))
+ return;Just patch.
No, it very much is.
> and risk
No, there is no risk. I already run everything as admin.
> just to have a little faster system.
5-30% is not "a little".
> Maybe you don't care about this patch
Indeed. And I expect many other power users also don't (but regardless, this is irrelevant).
> but you will need others that are dependent on it.
Well when that actually becomes a problem I will act accordingly. If more patches like this pop up I obviously won't install any of them. If there's a patch for a drive-by browser exploit depending on this, that will obviously be a different story.
> Just patch.
Hell no. My patching this makes my computer slower while providing exactly zero benefit to anyone.
I don't mean to be rude but if you are having to ask how to disable automatic updates then you probably aren't someone who keeps up to date with all the latest issues. When it becomes problem, you just won't know.
Let your OS vendor do all this for you. They are good at it.
...Wow. First of all, that's not what I asked. I asked how to disable or block this patch. Blocking "automatic updates" is neither equivalent to disabling this patch (post-install) nor to blocking it (pre-install). Second of all, I'm running Windows 8.1, on which I can actually block updates easily. I don't know if I can be picky about which patches I block on 10 because I have barely used it, but I will have to start using it soon and I really don't want to waste time installing the update only to find out I can no longer uninstall it.
And third of all, you're really spewing nonsense. I've done security work in the past which I don't care to post details about here anonymously. I still keep up with security news regularly and I actually look into the update details before installing them (which should be obvious if you read my previous comment on how I said what I do depends on the actual updates). None of which you need to believe (and I really don't care if you don't), except for the minor caveat that if you're trying to be convincing, this holier-than-thou attitude moves you well in the opposite direction.
If I understand well, there may be a serious bug in some x86 CPUs but nothing is known publicly. Presumably, all current Intel CPUs are affected and none by AMD but we can't really be sure, it is still a secret.
It is impressive how a simple, yet to be justified patch has so much influence. It opens up new ways of manipulating the market...
But if you scroll down to "Windows-Benchmarks: Anwendungen" you can see that most applications do not have any performance hit with the Windows patch.
Only M.2 SSD seem to be affected.
To me, this sounds like unnecessary work if Intel is coming up with a microcode patch within a few months.
As it threatened the company's existence, the President refused to consider guaranteeing a loan because the company had been a contributor, and any help could potentially appear as corruption.
How quaint such considerations seem today...
Talking about chromeOS, is there any speculation about the impact of bug? Does it's hardened sand-boxing techniques put it in a better position even if KASLR is compromised?
It was just stupid through and through. Even today we see ARM coming to full versions of Windows, but Google is still kicking it with Intel CPUs.
KASLR is/was the cover for the kernel patches, to avoid disclosing the real bug.
I have no real context for this, but is 7.2% considered a "soar"? And a 3.8% decrease seems like kind of not a lot, considering what's fucking happening here.
It's absolutely impossible for the most visible executive of one of the largest firms to engage in insider trading in such an obvious fashion and get away with it.
At this level, there is always a paper trail of who knew what when. There are internal and external audits if any suspicions arise.
Plus there are sever penalties, both civil (in the employment contract) as well as criminal. Intel's CEO is without a doubt in the 9- or 10-digit range of personal wealth. Risking time in jail to avoid a 10% loss on his stock holdings would be a terrible decision even if they considered the chance of being caught was low.
[in 2016] the FEC Charged 78 parties in cases involving trading on the basis of inside information.
A number of these cases involved complex insider trading rings which were cracked by Enforcement’s innovative uses of data and analytics to spot suspicious trading.
For example, brought insider trading cases against:
- two hedge fund managers and their source, who was a former employee of the U.S. Food and Drug Administration
- a former Goldman Sachs employee
- a former senior employee at Puma Biotechnology Inc.
I didn't have a chance to read the 2 papers so I would appreciate a TL;DR.
I am a SW developer working with high level languages; security and OS development are not my specific fields so while I don't need an ELI5 I would appreciate a sufficiently "layman's terms" explanation.
* https://googleprojectzero.blogspot.com.au/2018/01/reading-pr...
* https://www.theregister.co.uk/2018/01/02/intel_cpu_design_fl...
* https://www.amd.com/en/corporate/speculative-execution
* https://newsroom.intel.com/news/intel-responds-to-security-r...
TL;DR is that mitigation for Spectre on an OS level is not very expensive in terms of performance while Meltdown mitigation, which affects only Intel, will have a performance penalty between 5% and 30%
The SEC could look at internal emails he got about this flaw, and if he got them (and especially if he replied to them) then it's pretty clear he knew about it.
They doubtlessly have other investigative tools they can use for this as well.
The sell could also have been scheduled in advance, which would - AFAIK - not violate any insider trading regulations.
https://en.wikipedia.org/wiki/SEC_Rule_10b5-1#A_possible_loo...
It gets filed with the SEC so should be suitably permanent
Edit: Here's trustworthy reporting on that action that also predates this announcement: https://www.fool.com/investing/2017/12/19/intels-ceo-just-so...
Intel has a rule where c-suite and up execs who have been at the company 5+ years must hold a minimum amount of stock. For the CEO, that number is 250,000. The CEO recently exercised options and immediately sold them, holding onto the absolute minimum required (250k).
This is all available on the SEC's website, trades like this are reported to the SEC and are considered public information.
Apparently, Microsoft has been releasing patches for NT since November, so it isn't exactly predating the bug.
https://finance.yahoo.com/quote/INTC/insider-transactions?p=...
The marvel is that somehow the news was kept from hitting the media until he had finished his trade, and then was revealed immediately after the transaction was complete. I wonder how that coincidence happened.
At the same time, noone is doing eight figures transactions which require reporting to the SEC without talking to a lawyer. Right? Right? They really dislike insider trading, it's one of the few things where even rich people can get imprisoned -- and the typical jail sentence has been steadily climbing up for decades now.
Insider trading, like many white-collar crimes, exists primarily for its value as a weapon. There is nothing actually illegal about the act of selling a stock; it's all about casting aspirations as to intent and who-knew-what-when.
In other situations, intent is usually an aggravating factor, enhancement, or affirmative defense. It is not the thing that qualifies an otherwise 100% legitimate act as a bad thing.
My anecdotal, unsubstantiated perspective is that insider trading is unlikely to be an issue for anyone who hasn't made enemies, and that it may suddenly become an issue for anyone naive enough to make enemies recklessly. cf. Martin Shkreli, who couldn't be linked to a specific "bad trade" so was brought on generic "securities fraud" instead.
Not playing ball with the people wielding these powers seems to be the dangerous thing.
Yes, mens rea is an important consideration, but it's nuanced. Check the Model Penal Code [0], which identifies four differing types of mens rea, including negligence and recklessness; that is, a "guilty mind" (mens rea), for criminal purposes, does not necessarily require what would be conventionally considered bona fide malice or intent to harm.
What I meant when I said an "enhancement" or "aggravating factor" is that usually you have an objectively asocial actus reus, like theft, and should mens rea come into play, it's generally a defensive thing seeking to exculpate the accused, as in "I didn't know it belonged to someone else" (the affirmative defense), not to deny the act.
But with insider trading and other instances of nuanced malum prohibitum [1], the usual relationship between actus reus and mens rea is inverted. To make the crime, one starts with the mens rea, the bad intent, and must identify (or, if necessary, manufacture) an apparently-normal act to register as the external offensive conduct that harmed society and warrants legal action.
That is a scarier proposition because if your daily business involves technical acts that can be converted into actus reus, there's obviously going to be ample opportunity for people to assign and rationalize their preferred ideas about your thought process there and convince themselves that you're a criminal based on their personal level of dislike or offense. If this gets brought in court, your defense will amount to convincing the jury to believe you instead of the prosecutor, which is a straight-up likability and performance contest.
Whereas, with better-defined crimes, there is a physical, independent actus reus that people recognize as objectively bad and probably intentional. If you didn't steal the thing, if they can't show that you stole the thing, that's now the ground that you're fighting over, and that's much better for the defendant because it's much less fickle.
Essentially it makes every defense necessarily affirmative because the conduct is not otherwise unlawful. The government must dislike you enough to assume bad faith first.
[0] https://en.wikipedia.org/wiki/Model_Penal_Code#Mens_rea_or_c... [1] https://www.law.cornell.edu/wex/malum_prohibitum
And what the SEC can prove and what they can't, I do not claim to be an authority of.
I doubt the CEO is exactly shaking in his boots or sold for that reason. I'd even be willing to bet Intel will provide a fix on the hardware and continue to use the same sockets so cloud providers can just change them out without further hardware changes.
Intel's stock will probably trade down to 40, fill the gap and continue its uptrend, largely because the market is very bullish on the chip space right now.
Its impossible for him to not know something that is this serious.
If he actually did knew about this, then this is a proof of some serious incompetence.
http://www.nasdaq.com/symbol/intc/insider-trades
Most of trades are "automatic sells" which should mean the trades were setup days in advance. And the selling was reported on the back of the Intel issue only.
I assume the system calls to interact with the GPU, or to do any sort of I/O, are going to incur the performance overhead. So rendering frames, reading/writing from the network, and loading assets from the disk could all cause issues.
And this is just for gaming. Anyone using a cloud provider that's running on Intel needs to worry about similar things.
What a nightmare...
Though even if I lose some percentage, it will still outrun AMD on single core applications (and probably also multi core ones).
If I were Intel, I'd offer free replacements for at-par performance, and potentially tiny cash payment for upgraded performance. Assuming the marginal cost to produce chips, especially older/slower ones, is very low, the only real cost to them is losing out on potential upgrade sales which would have happened organically, for a while.
However, doing this keeps Qualcomm/ARM and AMD from making massive inroads into the market. As well, it would be a great way for Intel to accelerate adoption of their newer technologies, causing even greater lock-in (you could assume a much higher percentage of users will have a feature)
This all works great for socketed CPUs (still common for servers/cloud). For embedded, where CPU is more likely non-replaceable even if socketed) CPU peak performance probably doesn't matter as much -- maybe do a discount coupon? Or work with equipment vendors to subsidize upgrades.
Laptops and non-technical end users (who couldn't swap their own CPU) probably don't care as much but also don't have the ability to upgrade. A rebate/upgrade program would work, or a substantial cash payment. Doing it as e.g. $50 cash or $250 toward your next Intel-CPU laptop would be interesting.
Intel is rich, the market leader, incumbent across multiple segments, etc., so they really should go overboard on their response.
Assuming there is a software mitigation which has a performance impact, sophisticated users would be capable of adding more capacity (if it's a horizontal scale type workload), upgrading early (if they had extra capacity for futureproofing), or spending money, potentially subsidized by Intel, to upgrade immediately. If there's no mitigation, upgrade early, or rearchitect application (moving away from shared security domains on single boxes, etc.)
http://erlang.org/pipermail/erlang-questions/2018-January/09...
I will eat my hat if that 30% number holds up.
I don't keep up on these kinds of metrics, but I'm under the impression that Intel still dominated CPU benchmarks berfore this issue, so if this question is answered in the negative then I doubt it will affect Intel very much.
Spectre affects every modern CPU, and performance impact will be marginal. Meltdown affects Intel-only, and the performance impact will be significant. Here's Fortnight's preliminary server results[0].
Once everything[1] is patched up it's just going to get worse for Intel. Ryzen was already fast though and had superior SMT (Hyperthreading) efficiency than Intel prior to this bug (~10%). This is going to take a serious toll on Intel's single-threaded performance advantage (~10% from IPC, 25% clock frequency advantage).
My advice is to stick with Ryzen/Threadripper/Epyc.
[0]https://www.epicgames.com/fortnite/forums/news/announcements...
[1]https://gist.github.com/woachk/2f86755260f2fee1baf71c90cd653...
If not, then the fix may expose many driver bugs (where driver accessed user space directly instead of through copy_from_user()), making the whole thing even more painful.
If so, clearing all caches upon a failed privilege check sounds like something within the capabilities of microcode and without unreasonable performance penalties.
Unfortunately, that would not explain the complex in-kernel fix...
[1] https://plus.google.com/+KristianK%C3%B6hntopp/posts/Ep26AoA...
EDIT: what remains in cache is not "speculatively read privileged data" but more "unprivileged data whose address is correlated to speculatively read privileged data". Retrieving later such address allows one to infer what the privileged data was. Still, the point about clearing all caches as countermeasure holds...
Not to mention it's pretty misleading about technical stuff, saying e.g. "Chip design errors are exceedingly rare. More than 20 years ago, a college professor discovered a problem with how early versions of Intel’s Pentium chip calculated numbers."
So, they imply that chip design errors are once in a few decades kind of things; that's just complete nonsense. Chip design errors are common most just aren't this bad.
Furthermore, they imply this stock change is anything more than temporary volatility; given the tiny, tiny changes so far that's just premature. Perhaps the stocks will adjust more in the future - we'll see. But as of now, this piece is financially and technically pretty misleading.
Welcome to the world of business reporting.
To be clear; I'm not saying this increase won't stick, just that bloomberg is pretending they're reading this from the numbers, not merely expecting it to occur. It'd be fine to say you might expect the stock to trend higher. But pretending you can look at that noisy line and say the current 7% increase is statistically significantly higher in a humanly relevant way is just hogwash.
Obviously you might expect the stock to soar. It may well happen. It just isn't visible in the data they present; the article is simply click bait (or worse, market manipulation).
I didn't have options activated till this morning, because i hadn't gotten around to it. Had I been able to activate them yesterday afternoon, I'd be in on INTC and AMD, both sides. Anyways, I have options in play on the intel side but not the amd side, so while "Talk is cheap" - I'm in.
"All computers with Intel chips from the past 10 years appear to be affected... The security updates may slow down older machinery by as much as 30 percent, according to the report."
"discovered a problem with how early versions of Intel’s Pentium chip calculated numbers... Intel had to recall some chips and took a charge of more than $400 million."
Worse, they understated the noisiness of the chip error "signal" but they also understate the noisiness of the stock value signal. And hey presto; if you conveniently present only the data you want all of the sudden a correlation looks really meaningful!
Furthermore, even the qualification "major" isn't really deserved quite to the extent that bloomberg implies; specifically there have been worse chip errors much more recently than 20 years ago (e.g. amd's somewhat similar phenom issues). What makes this one worse isn't just the bug, it's how chips are used: cloud computing makes this much more relevant than similar bugs not too many years ago.
Disabling a whole feature (TSX) across an entire generation of processor and early steppings of the next? That's pretty major, given it's a marketed feature.
Many chip design errors are patched with microcode updates. There is speculation that this one cannot.
What do you think, is this realistic?
Tl/DR: Intent matters.
Doesn't seem to be much of a rumor.
Are you doing a lot of really CPU intensive tasks that make a lot of syscalls? Most people aren't. If you are, you'll probably feel that pain. And if you're not, you probably won't notice a difference. And if you don't know you can either research if what you're doing typically is going to be affected, or wait and see.
Do we know that? There is a lot of speculation but the embargo is not lifted yet afaik. There are obviously elements of this that are remaining intentionally obscured pending embargo expiry (redacted comments), so it could be that a smaller contingent of chips are affected and no one is bothering to correct the damage/limit the patch only to applicable components so that the cat doesn't get out of the bag too soon (side benefit: so that an extensive emergency fix like this can be widely tested before it's applied to a relatively limited set of hardware).
Seems just as valid as any other speculation to me.
https://www.google.com/search?q=intel+stock
Edit: AMD 8% up: https://www.google.com/search?q=amd+stock
This bug is a true security bug that has been lurking around for 10 years... and it could have been already abused. This is double scary because this attack does not leave any traces behind and it allows any application to pretty much enter God Mode (this is the worst case scenario of a security situation (hence the name meltdown (like nuklear meltdown is the worst case situation in a nuclear power plant)).
This bug is also bad because in certain situations the fix of this bug causes severe peformance penatlty.
You are only fucked if you don't update your OS.
I wonder how this will play out since it seems that everyone, apart from Intel, view this issue as a "flaw".
No, you can just read data you shouldn't be able to.
> Intel is committed to product and customer security and is working closely with many other technology companies, including AMD, ARM Holdings and several operating system vendors, to develop an industry-wide approach to resolve this issue promptly and constructively.
Let's try to drag AMD and ARM into the mud with us.
> However, Intel is making this statement today because of the current inaccurate media reports.
"Fake news!"
Well you can't fault the PR folks for trying to spin it. They had to say something.
Dragging AMD's and ARM's names into this is completely inexplicable to me however.
> Advanced Micro Devices Inc. surged [...] after a report that Intel Corp., its only remaining rival in the market for personal computer processors
Ahh, so Intel is the only holdout against the dominant wave of AMD processors, right?
facepalm.jpg