Linux 5.8 Set to Optionally Flush the L1d Cache on Context Switch
phoronix.com
phoronix.com
Or where the payoff (due to limits/quotas) implies a cap on the effort that's still worth it. It's like Backups: you do as much as you want to pay for, considering the amount of risk you'll be left with.
Now pray that there is not another exploit that can be used. :)
Like many things; security is an onion, there are layers to it and removing some of those layers can be fine because the outer layers may protect you, but ultimately you just increase risk.
Don’t get mad about the fact it didn’t work, I’m glad it didn’t work, but you can’t walk away with the knowledge that something like that will _never_ work.
I’m not sure what you were trying to prove, but you can’t tell people to prove you wrong without any knowledge of what browser you’re running with, or what version or what you’re expecting to see.
A “fully working” exploit for the most modern browser is probably possible frankly, but it’s not something that anyone is looking at with seriousness because everyone has mitigation’s enabled anyway. It’s the very definition of high work low reward.
No you didn't. You posted some code that doesn't do anything.
>I’m not sure what you were trying to prove, but you can’t tell people to prove you wrong
I didn't ask anyone to prove me wrong. I asked georgyo to prove his claim that an exploit is possible. You chimed in and posted some nonsense code which obviously does nothing and never did anything even before Meltdown and Specter were mitigated anywhere because even the original native PoC's were much more complicated.
>A “fully working” exploit for the most modern browser is probably possible frankly
Stop talking out of your ass.
IMHO in the cloud a sane mechanism to protect users would be to guarantee exclusive access to CPUs. The primary risk in the cloud is sharing CPUs with other cloud users. Cache flushing sounds like a sane thing to do in non performance critical applications where CPU sharing happens. There's a performance penalty but if you'd care about that you'd be using a different instance type anyway where you are guaranteed to have dedicated CPUs.
For bare metal/desktop use, flushing caches makes less sense.
I remember those days, and actually shipped "dynamic" webpages before JavaScript/AJAX/DOM. I'll take cache flushing.
Imagine if every like/upvote button required a screen refresh... imagine Google maps...
The difficulty with both timing and cache attacks is that a "sandbox" approach is not possible... at least not without special hardware and OS level support like the ability to tell the OS to flush caches on context switches for certain processes.
You already went down the wrong road there. Can you imagine how to add functionality declaratively?
The middle ground is extending user agents to support features that improve the browsing experience and enable new functionality, without turning it into an arbitrary application delivery framework. And yes, this means rejecting and scoping out some things that people do in browsers today.
Just like in the past, you as a developer would get to choose whether to write a website (ubiquitous, accessible, runs on grandma's old potato, relatively easy and cheap to maintain; these are the things that made the web popular and businesses around the world decided that it's ok to to make their product a website even if that meant they couldn't have all the features you could with a desktop application), a native application (more effort, more complex, more expensive, more invasive, more friction, more concerning w.r.t. security), or both.
To illuminate you, let's pick some examples from the false-dichotomy post...
Consider form validation. In fact, this is already done (and could always be extended to support more cases). HTML5 has built-in form validation that works without javascript. And of course it's still perfectly backwards compatible with old browsers that don't do validation; they might send you invalid fields, but you will have to validate them server side anyway because you can't trust the client.
https://developer.mozilla.org/en-US/docs/Learn/Forms/Form_va...
Static maps have already been done. Not a super smooth experience, but you could always improve that by speccing a zoomable & pannable tiled image element that'll send requests to the specified URL when you pan outside of the loaded area or zoom in. Add a set of loadable elements that are embedded into this image element and you get something that starts to resemble SVG. No JS required.
Information about points of interest could already be shown with the hover selector, but there's no reason we couldn't spec an element whose visibility can be toggled with a click, no js needed.
Upvote/downvote buttons just need an attribute that tells the browser to post the request but stay on the current page. (This also degrades trivially with browsers that don't support the attribute) You could even toggle the visibility of the arrows after posting; similar CSS selectors for checked inputs already exist.
In general, there's no reason we can't have post or get requests that display the response in a new element without reloading the entire page. Semantically, not very different from target="_blank" or whatever you use to load something in a new tab / windows, except this time you want the target to be an element.
(At this point I'd also like to note that frames exist and yes they suck but hilariously a lot of the new web does exactly the kind of stateful non-linkable things that framesets were derided for; only worse, because you actually could right-click a frame and link directly to it, but you can't right-click and link the arbitrary DOM that was cooked by your client-side javascript)
Going with the tiled image element theme, there's no reason we can't have more elements that instruct the browser how to load more data on demand. These same elements could let the user agent decide whether to paginate or scroll infinitely, or how many items to display per page.
(There's no reason we can't load images progressively and on demand.. progressive JPEG exists already, but for some reason devs still insist on giving me a blur and nothing more will load unless I enable scripts)
The way we currently do things really sucks for the user (because they have very little control over how the script behaves; the user agent is degraded to a mere dumb client with little meaningful configurability) and it sucks for developers who would rather just focus on the content and let the browser provide whatever UX fits the user & their platform best.
Web devs are in a hurry to paper over the deficiencies of browsers but in doing so (and not fixing browsers), we end up with something worse and every goddamn website becomes a complex application. We're stuck in a worst-of-both-worlds state, where the browser runs applications that lack the power of desktop applications, yet are invasive, heavyweight -- don't run on grandmas old potato, complex & expensive to develop and maintain (every website shipping complex UI logic that should be part of the browser instead), increasingly less accessible and less reliable, less linkable & crawlable, less secure.. it's all I never wanted.
Your reaction will probably be that I'm missing the point, that if we went with a JS-less world there would be solutions for this. But I strongly suspect that the solution would be to not use HTML and instead use some other technology that was capable of general computation on the client.
Nitpicking the details of how it is currently implemented is indeed beside the point. Ideally, the spec is made loose enough to give user agents & users the freedom to configure the behavior to their liking (and if someone can make the case for a particular behavior must be followed in some situations, then an optional attribute is added to "force" that behavior).
In general, I'm very tired of the status quo, which is that every site developer is responsible for providing good UX and people nag at them, when their preferences could be accommodated for by the browser itself. As long as the behavior stems from javascript, there's very little a browser can do to accommodate user preferences without breaking the web at large. You know, maybe I don't like form validation the way you'd implement it in JS.
People are so vested in the status quo that some of them even get angry when you e.g. suggest that they could use the browser's reader mode (instead of nagging at the site's author) to make a site readable for themselves. Bikeshedding about colours and fonts on front page HN postings happens all the time... of course, reader mode is a hack that fails very often, so disagreeing with that suggestion is somewhat justified. But really, we could've built the web around the user agent instead of vice versa, and then your web browser would be your reader mode by default. You could blame at your browser vendor or yourself first of all if the colors and fonts (or input form validator behavior before you've entered anything) don't please you.
So maybe the thin client model? All the code runs on the server and the UI is streamed to the client? But the lag would be higher and the cost for the web server would probably have prevented any internet boom.
Also, this problem is fundamental and not limited to a single core, package or machine. Any time you are making opportunistic optimizations (of which caching is one), and the opportunities available depend on what happened outside of a privilege boundary, you can have this problem.
A similar thing happens at a higher level of abstraction with, for example, btree operation timings in a database.
The problem of course being that CPUs that rival performance of a decently-specced x86 VM are going to be pricy, mooting the point.
Based on these architectural features Intel has had CAT, which essentially turns LLC slices into private caches for certain cores. That's intended for performance, but is now also relevant for security.
Graviton 2, by ditching SMT, has one such mitigation.
This goes against decades of OS design, possibly all serious OS design once you pass over program loaders like MS-DOS and various ROM BASIC iterations. More to the point, there's enough advantages to being able to run untrusted code that a performance trade-off is worth it: If it comes down to being able to run your business at a penalty and having to close up shop, it doesn't take much to figure out what Amazon is going to do with AWS.
[1] https://www.theregister.co.uk/2018/06/20/openbsd_disables_in...
Just manually figure out all(or the chosen ones) thread ids spawned by the chrome process.
Among other things there are now mechanisms for doing page rasterization (PaintWorklet) and sound synthesis (AudioWorklet) in JS, so the overhead involved in making all your JS share a single core becomes more dramatic. There's also lots of stuff out there that uses Shared Workers and Workers to do background computation and that won't be background anymore if you pin them all to your JS core.
This used to only be true on x86--IBM and DEC took security seriously and bitched about this incessantly.
Nobody cared. x86 was cheap.
Eventually everybody just threw up their hands and went to superscalar, deeply predicting, out-of-order microprocessor architectures because that got you better benchmarketing
x86 was always insecure. It's just that nobody cared until The Cloud(tm). Malicious client-device Javascript just made it all worse.
All of the speculative execution mitigations can be turned off with a single flag, I much prefer to have the option to turn them on/off, rather than either not having them at all or not having the ability to disable them.
So from my point of view, we're exactly where one would hope we would be in the face of these hardware flaws.
2. It's for the paranoid. So you don't have to do it
3. It's on context switches. These aren't that common, and reloading it will pull entire cache lines in from the L2 cache which is pretty quick anyway.
But this I agree with:
> If untrusted code is running on the same core/package/what have you, your security has already been breached.
The biggest untrusted sod to be running is usually the browser. Of course if you're worried about that, how about turning off JS thereby blocking the biggest attack vector in it? (and loads of bloody irritating behaviour, as a blissful bonus).
Edit: why the world-is-going-to-shit attitude I keep seeing everywhere? The worst possible interpretation is put upon everything, instead of evaluating the risk/reward rationally then choosing appropriate actions.
Until it's proven not to be so, through another POC.
And then OPs point stands: Lots of Intel's performance gains since the 90s has been through out-of-order execution and branch-prediction.
If those improvements are deemed incompatible with being able to securely run JS in your browser, I would argue Intel is having a very fundamental problem now.
Hopefully AMD does better, but I don't think they are entirely immune to this category of security-issues either.
I also suspect the performance hit will be minimal, however I'm not an expert.
> If those improvements are deemed incompatible with being able to securely run JS in your browser, I would argue Intel is having a very fundamental problem now.
Well it is, yet people are overwhelmingly willing to expose a turing complete language controlled by some 3rd party they know little or nothing about directly to the open internet. The problem there is nothing to do with hardware. It's people.
(agreed about AMD)
A context switch probably happens thousands of times per second on modern systems. What do you mean by "not that common"?
Plus the other overheads of switching are already there - it's not cheap. I don't expect the overhead of reloading from L2 cache to add much (to repeat, I'm not an expert though).
Same as - Volkswagen (dirty cheating cars) - Samsung (exploding phones) - Boeing (falling planes)
Eventually tech companies will join them. The signs are already there, the dirty growth tricks at Google and Amazon. It'll take more time.
The issue is that gains at performance (...environmental performance, form factor, battery density, etc) are not linear anymore - the investments to make further gains becomes increasingly more expensive and time consuming and all the quasi market duopolists and too big too fail national infrastructure companies are not able to grow slower- stock market dynamics would punish them, execs wouldn't get their entitled pay day, politicians would lose jobs, tax income.
And so corners are cut (Boeing, Samsung, Intel) or performance tests are cheated (all of the above) and slowly infrastructure of dependent industries (cloud, transportation) is built on more and shaky ground.
So why? Market concentration, entitlement, too big to fail dynamics, endless growth doctrine.
It would have been better to preserve users (UIDs) as the 'security' boundary for data. Leaving processes (PIDs) to offer the containment for safety (virtual memory etc.), and specifically not security of data.
Then this specific case would only need the cache flushed on context switching between UIDs, but not all processes.
Attempts to retro-fit security of data between processes are resulting in as much a negative as a positive -- eg. performance loss; or usability issues like regular users on Linux gdb'ing their own processes now has to be explicitly enabled by root.
Of course, the new thing is all this untrusted code we're running; understood.
I would suggest to preserve UIDs as the security boundary for data, then direct these issues through a mechanism that makes it as easy for an unprivileged user to 'fork' a UID for a purpose, just as they fork() a process. This gives a clear role to UIDs/PIDs and a sandboxing capability that makes use of all the existing implementation and boundaries.
Separate UIDs quickly get confusing too.
I think such separation would be better done via cgroups and namespaces instead.
cgroups and namespaces are privileged mechanisms that only 'root' can use, and share UIDs and other scopes across them (UID namespaces are complex and a risk in themselves.) It's required to re-implement all the required policy.
Separate UIDs need get no more confusing than PIDs. I agree we really don't want /etc/passwd with a gazillion entries; we should only be managing them only as much as we "manage" temporary PIDs. But imagine a command to view the "UID" tree like the PID one.
I agree it might not be how we'd design it if we designed from scratch. But then we may have chosen cgroups that integrated into the process tree and several other design decisions would change.
The name "user" ID distorts the conversation because it's unintuitive, but I am suggesting that UIDs already embody much of the policy that is needed.
Mobile operating systems are far ahead of desktop operating systems when it comes to making the "application" a first-class entity. macOS comes closest since most ordinary applications can be largely self-contained within the app bundle and the sandbox is steadily getting richer protection/isolation mechanisms. Linux applications have fairly well-controlled install and uninstall processes thanks to distro package managers, but post-install behavior cannot be managed with application granularity in any standard way (though there are some projects seeking to accomplish this, if they can first succeed in replacing existing distro package managers). And Windows is still largely allowing applications to spray files all over the disk and run whatever in the background.
The challenge for the desktop is that we don't really want to switch from a multi-user paradigm (with all applications in the same security domain) to a single-user paradigm with per-application security domains. We need a multi-user OS with per-application security domains, and mobile operating systems aren't quite there. (I've heard that Android can be multi-user, but I've never encountered that functionality in the wild.)
like cgroups?
Maybe I'm misunderstanding, but to me what you suggests sound like it would allow malicious JS running in my browser should be able to snoop data from my secure password manager (running as a separate process), because they both run under the same UID?
Is that correct? Since most systems are pre-dominantly single-user systems, I honestly think the PID-isolation model makes more sense, as I'm not trying to defend against other users trying to spy on me, on my own laptop.
Your browser has the role of bringing in untrusted code, and running it. The browser code would 'fork' a UID to run the untrusted code (and only that), and then we make good use of all the existing UID-based policy in the kernel.
What would be the privilege set of that new UID though?
And it would be a poor UX if that separate UID had absolutely zero access to my (human) UID-secured files because then it wouldn't be able to access my browser cache and history - and I'd rather Chrome didn't decide to require each browser process to have its own non-shared cache. It's bad enough Chrome is now using 3.5GB for 4 windows (16 tabs total) on my desktop.
Yes, for the reason you describe, inter-operability with your own UID is actually good. My previous job taught me just how much this can be a feature not a bug; we made extensive use of applications, plugins and various forms of IPC to allow desktop applications to inter-operate in powerful ways.
There are other mechanism already exist. When a process (or thread) is forked, various resources can be passed over the boundary; file descriptors, shared memory etc.. For example, where un-trusted code needs access to a file, the mechanism to do that is already there in a nice "opt-in" manner.
Since you can put uid and gid in firewall rules it makes for interesting belts and suspenders component separation.
Combined with ZeroMQ, and SPARK and lots of generated code for the inter-thread comm and you can build modular designs.
Just don't look at you ip link / ifconfig :-)
We routinely rely on virtualization and containers for isolation. Even in this very browser you're using, you are supposed to expect untrusted hostile code to run in other tabs. Imagine going to a website and their JS is reading memory from your password mananger, and would you then have a similar reaction?
Why do you need single core performance so badly. I can't even think of an example where single core performance has been an issue for me, and I routinely run tasks on decade old CPUs. And even if it was an issue, You are essentially saying security should be an after-thought,sorry but your reasonig is very dangerous. Would you get in a car where the engineer of the car thinks "if people are ramming your car on the freeway,you have bigger problems,let's focus on making it lightweight,fast and fuel efficient".
And to be frank with you, people that have been doing admin/engineering work since the 90's with that attitude are a bigger security threat to most orgs (and themselves) than any hacker (with the exception of the few orgs/people that receive targeted attacks frequently). The days of treating security as a perimeter issue have been long gone for about a decade now. Whether it is network, system or software security,the entry points and perimeters have been rendered meaningless (lookup zero trust, I think it applied here too).
Because on the cloud, performance is directly correlated to your monthly bills.
Intel should be held accountable. A fine every quarter they continue to put national security at harm.
I think it's much better to lay the blame at the feet of the companies who made these decisions, rather than generically blaming everybody.
Theo De Raadt called out Intel in particular when they started taking all kinds of crazy shortcuts like this, all the way back in 2007.
I am unsure if this is just idle speculation (heh) that there may be issues in this area or there are issues that have been disclosed to vendors but not the public yet?
May seem over cautious, but certainly for most, a default position of over-cautions and for those that know how to play on the edge, well those would be able to compile their own kernel.
Given Linux in so many devices and so many blackbox left alone systems, not a bad default position to be taking - planning for the worst, expect the best.
This is L1d cache which is just 48kB for Ice Lake. We are also talking about context switches which are not happening very frequently. Applications that are generating load don't context switch all the time because they are busy doing work.
Then, when you context switch it is likely the context to which you are switching would like to use that cache for something. By the time we switch to your original thread it is very likely L1d has already been filled with something else.
I am pretty sure you would not notice anything except for very special, rare situations.
100ms is a huge amount of time and 48kB is a tiny, tiny part of what processor does during 100ms. Gigabytes of data can be transferred during that time, 48kB isn't really much.
As I have pointed out, that cache has very little value over context switch anyway. The cost is removing data from cache that would be usable after we have returned to the original context. But it is already very likely the data in the cache is already for a completely different context and hence completely unusable.
Say you have apps A and B and OS.
You are running A which has 48kB of data in L1d. It switches to OS which causes some of L1d to be evicted and puts its own data there. Then it switches to B which is likely another process, this causes very likely entire L1d to be evicted unless this is extremely small process. Then we come to OS and again to A. By the time you are at A, there is no data from the original L1d state.
Cleaning L1d upfront on context switch is likely not hurting anything.
Any noticable perf overhead is going to be from the act of cache flushing taking some super slow path for some reason, or much more frequent context switching than 100ms timeslices.
[1] https://stackoverflow.com/a/4087331
[2] 1000x 32-128B cachelines = 32-128KB, definitely in the ballpark to completely refill a 48KB L1D cache.
[3] https://en.wikipedia.org/wiki/DDR4_SDRAM#Modules
[4] https://www.wolframalpha.com/input/?i=48KB+%2F+12800+MB%2Fs
Consider that a single core on a modern CPU running at 2 GHz can execute over 20k instructions in those 100ms.
Anyway, 100ms is quite a lot in the life of a modern CPU.
Lots of stuff happens during those 5us. The message is read from the network device (directly by the application, no Linux or syscalls anywhere during those 5us). Then it is parsed, deduplicated (multiple multicast channels carry redundant copies of the messages), uncompressed (the payload is compressed with zlib), the uncompressed payload is parsed, interpreted (multiple types of messages). Business logic is executed to update state of the market in memory then to generate signals to listening algorithms. The algorithm is run to figure out whether it wants to execute an order. The order is verified against decision tree (for example to check whether it does not exceed available budget). The market order packet is created and sent over TCP.
Now imagine, all that stuff happens in 1/200th of 1ms. In comparison, transferring 48kB from L2 or L3 to L1 is pretty damn insignificant.
[1]: https://github.com/torvalds/linux/blob/master/kernel/Kconfig...
Presumably more of a problem if all cores are busy, which is more likely if there are few cores. Also dependent on the number of interrupts (e.g. high network traffic of small packets etc). Presumably not a problem if there is an idle core that can run the interrupt code.
Either way, I am sure there are plenty of devices that can cause a lot of interrupts (USB?), not just network IO. Presumably there is a way to monitor the count of interrupts per second in Linux?
The cost would be right if the cache was usable after context switch. Since it is likely stale, the new context will be pulling new data into cache as if nothing really happened.
But FWIW: most HPC computing is, in fact, "shuffling memory around", yeah. Very few architectures are actually interrupt bound, and the ones that are work very hard to address that (because hardware interrupt parallelism is an even harder nut to crack than context switch overhead).
Edit: I wonder why the downvotes. Switches between in and out of kernel have never been called context switches that happen between threads. I know no one who calls them 'context' switch as the context, i.e. registers that point to the thread/cpu core remain the same.
It provides a command-queue/response-queue dual-ringbuffer interface to the kernel, mostly providing benefits in terms of less per-IO-op overhead and offering non-blocking buffered disk IO.
It can work in a zero-syscall steady state after program startup for applications such as (for example) web servers.
The other thread may have been doing work with memory on a GPU. The other thread may already have a hot cache at another layer. It's definitely not an edge case, or else the L1d cache would not have been designed to maintain state between context switches in the first place. There are going to be consequences to this.
Also, context switches can be very frequent in some designs. For example, in micro kernel systems you often have ping-ponging with processes communicating with servers via RPC. Wiping out your whole L1D every time that happens could be pretty unpleasant.
That depends on how your software is written. If, for example, you're running a web server that uses a thread-per-connection, you'll be context switching all over. Hi Apache!
The real reason flushing L1d is not going to be noticed is that even without flushing the cache is unusable after context switch. It is highly unlikely the next thread that gets ownership of the core will require exactly the data present in L1d.
On a busy web server the two most frequent reasons to switch context will be:
1. The thread is waiting on I/O so it yields the rest of its time share back.
2. The thread has finished processing request.
Now, if you imagine a thread that just did a bit of I/O returning its time so that OS is switching context to another thread... it is very unlikely any of the data in L1d has any meaning or worth for the other thread. Anything that the next thread will do will require fresh data at least from L3.
So L1d is practically worthless and blanking it isn't going to do anything noticeable.
(I have intentionally omitted all the interrupts happening in the meantime and OS also using the cache which is the proverbial nail in the coffin when it comes to usability of L1d after context switch)
L2 would be absolute crazy town though.
”Burks, Goldstine, and von Neumann, "Preliminary discussionof the logical design of an electronic computing instrument," 1946.
[1]https://support.microsoft.com/en-us/help/4497165/kb4497165-i...
[2]https://www.windowslatest.com/2020/05/21/windows-10-kb449716...
Interestingly enough: https://www.theregister.co.uk/2020/05/24/linus_torvalds_adop...
Why one or two CPU cores couldn't be dedicated to the OS aspect and locked out of user-space of any form, certainly would be something worth exploring.
Be nice though to have a proper isolated core or two for the OS, after all - that is exactly what is done for enclave based security and management systems. Though some not all a great track record.
Add more fine grained memory partitions to let the memory hierarchy in on what you're doing.
Make Rings 1 & 2 Great Again
I'm not even sure they could have delivered equal performance at the high end if they had included these mitigations earlier. Whether to produce processors which are safer or faster depends on what customers prefer, even now. Not everyone needs ultimate security or wants to pay for it (in money or performance). So unless the law says less-secure processors must never be sold, this situation was and still is inevitable.
Products were recalled for smaller issues than that.
https://www.recallmasters.com/mercedes-recalls-vehicles-defe...
Also, intel did replace CPUs with issues before:
https://en.wikipedia.org/wiki/Pentium_FDIV_bug
> On December 20, 1994, Intel offered to replace all flawed Pentium processors on the basis of request, in response to mounting public pressure.[5] Although it turned out that only a small fraction of Pentium owners bothered to get their chips replaced, the financial impact on the company was significant.[citation needed] On January 17, 1995, Intel announced "a pre-tax charge of $475 million against earnings, ostensibly the total cost associated with replacement of the flawed processors."[1] Some of the defective chips were later turned into key rings by Intel.[6]
Part of the initial stages of a class action is verifying that there does indeed exist a class.
There's value in him talking about his grievance publicly and not just rolling over when he gets screwed by a corporation, just because that can beat him one on one in a legal brawl.
> We accept that our $200 wafer of silicon with 13nm features that can execute billions of mathematical operations is "good enough". Sometimes there are bugs. But we don't know how to make these things ourselves, so we deal with them and aren't really looking for a pound of flesh from Intel because the billions of instructions their CPUs can execute per second is a slightly lower number of billions.
I also can't build a modern car. Or insulin. Or a million other devices in my life. The bar isn't "you can only complain if you can make it yourself better", it's "you can complain if what was sold to you didn't meet it's advertised specifications".
A replacement CPU would make more sense. It should support the same operating systems and motherboards. I wonder if Intel would be asked by a lot of people to replace their CPUs.