Mitigations landing for new class of timing attack
blog.mozilla.org
blog.mozilla.org
I am really happy to see this happening at long last. One can argue that almost every x86 SW side channel attack (Spectre and many prior ones that, for instance, leak parts of cryptographic keys from variable-time SW implementations) are being aided by userspace processes having access to high-frequency/high-res counters for which, in almost every case, they have no legitimate need. Perhaps future silicon should block direct access to the TSC & friends, except for processes that have a privileged flag set (hardly a new idea; IIRC some IBM PowerPC chips from the late 90s support that).
OS-level timing calls such as gettimeofday() could be degraded to a few microseconds resolution by default - poor enough to obscure delays caused by the state of various HW caches, branch predictors, etc.
Ultimately there is little reason to give every website js program access to a nanosecond-level counter by default, and many defense-in-depth reasons for not doing that. So kudos to the firefox team here and hopefully chrom[e|ium] quickly copies this.
You need to address the specific vulnerabilities, not shoot the clock_gettime messenger.
[1] https://www.usenix.org/system/files/conference/usenixsecurit...
No. It's true that with a low-res clock source in some cases an attacker can still get the precise measurement they're after by repeating an operation 1000x+ more times. In some cases though that is not possible, because at that timescale the signal they're after is degraded by noise that they cannot predict or control: operating system interrupts, memory traffic from other threads, etc.
Anyway, even if a lower-res clock source helps only 10% of the time, on defense you should always prefer to make an attack complicated, slow, partially-reliable rather than trivial, fast, highly-reliable.
Developers who need a high-res clock source for profiling, etc, should of course still be able to enable one selectively.
> You need to address the specific vulnerabilities, not shoot the clock_gettime messenger.
You can and should do both, to gain some defense-in-depth against future attacks.
Wrong, those signals will average out over a long enough data collection time.
But if you have a signal overlaid with random noise [1], and you know what your signal's looking like and when it's happening, you can correlate. For example, a small delay that occurs at certain known points (or not), will introduce a bias into a timer measuring it, no matter how noisy or coarsely quantized that timer is.
Similar techniques have been used in other fields for decades to pull useful signals from far below the noise floor (e.g. a lock in amplifier can go dozens of dB below the noise floor, because it, essentially, correlates frequency and phase and thereby eliminates all but a tiny sliver of noise. E.g. GPS signals are typically 20 dB below the noise floor.
[1] It doesn't have to be random.
——
So these mitigations just make the attacks harder, hopefully hard enough that they become not feasible to be exploited widely.
Exactly. The same pixel isn't imaging the same location on the distant object. If it were, then what you say might be possible.
Contrary to intuition, the presence of random noise can actually make detectable a signal that is otherwise below the minimum detection threshold. See, e.g. stochastic resonance. (Essentially, the noise occasionally interefers constructively to 'bump' the signal up beyond the threshold to make it detectable.) If you are able to introduce and control your own noise, you may also be able to take advantage of coherence resonance.
Randomness itself can be a very useful tool in many signal detection and processing systems, e.g. sparse random sampling in Compressive Sensing techniques can reconstruct some kinds of signals at frequencies beyond the Shannon-Nyquist limit of a much higher-density fixed-frequency sampling -- something thought impossible until relatively recently.
I would not be at all confident that such 'system' noise could not be filtered out statistically; it might even be used to an attacker's advantage.
Always? I would think it would depend on what the trade-offs are, what costs you are paying for doing that (in inconvenience or damage or cost to the non-attacking users and use cases; in opportunity cost to other things you could have been focusing on, etc) compared to how much you lessen threats. Security is always an evaluation.
Timing attacks are the worst though. I think this may only be the beginning of serious damage to the security of our infrastructure via difficult to ameliorate timing attacks.
Also note that if you reduce precision you might still be able to tease out the data simply by gathering more samples. An exploit that gets you 1000bytes per minute might take an hour instead, but that could be enough to find cryptographic keys.
I agree that for a lot of cases, it's plenty. But if you're trying to account for phase differences in multiple data sources, this puts a major blindfold on.
Let's think about it in terms of lower frequencies. Imagine that you're sampling from a pool of events that occur roughly 100 times/hour. Let's say you want to sample 2 event per hour.
Well you might say "since you're only sampling at a rate of 2/hr, why would you need any better granularity than 30 minutes?"
Here's the catch: you need to get the event that's closest to the 15 minute mark of each hour and the one that's closest to the 30 minute mark; you will then put those two together using your special formula.
Imagine that all your events came into a mailbox on an half-hourly basis. Imagine a stack of ~100 envelopes, each labeled "8 - 8:30 A.M." or "8:30 - 9 A.M.". You're tasked tasked with the job of trying to pick the two closest to 8:15 A.M. and 8:30 A.M. Good luck! It's gonna be hard, if not impossible, to do this consistently with any accuracy.
(Even if you had 15 minute granularity, you'd still have to pick from each one from either 25 or 50 envelopes, depending on where the windows fall relative to the 15 minute marks.)
Now, to make matters even more complicated, imagine a situation where you're integrating each hour's result with the previous hours' result, e.g. taking a rolling sum or product, or computing some kind of feedback/delay/reverb filter. In cases like that, you're done for.
Full disclosure, I'm coming at this from the perspective of an experimental digital artist who some day would like to do interesting creative things with high frequency sources, like the sounds of bugs, dolphins, and higher frequency ambient sounds. I'm also interested in building applications that crowd-source and integrate audio from multiple smart-phones in a room. The web seemed like a perfect platform for this.
Depends on the system IMO. I certainly want the TSC when I'm profiling something highly-performance-sensitive on my workstation. Yet I am hard-pressed to see why any userspace app on my Mom's chromebook requires a ~1ns counter.. some stuff there may be currently using it, but that is different from "actually requires it"/"should continue to have it from an overall cost/benefit point of view".
If Alice performs an operation only once, and it leaks sensitive data through a timing variation of 100-200ns, Mallory is in great shape with a 1ns clock source and in pretty awful shape with a 10us one.
This problem has been studied a lot. The venerable TCSEC Rainbow Series dedicates an entire volume to covert channels (the light pink one iirc).
It is a statistical problem. Even if you reduce the timing precision or randomize it and effectively raise the noise floor, it just takes a little bit longer for the attacker to get his data.
A good analogy would be Differential Power Analysis (DPA). Measurements are collected over a period of time to enhance the signal.
It just takes more samples and in the end it solves nothing.
>>non-attacker-controlled threads running on the same core.
Again it might take more samples. However, worse: the system is like unresponsive at the time. Also the timeslices are long enough to carry the task, unless there are way, way too many and unpredictable context switches (which would be bad for performance). ---- Back in the days of old there were no built-in timers and people used to count cpu cycles to accommodate for external io.
> In the longer term, we have started experimenting with techniques to remove the information leak closer to the source, instead of just hiding the leak by disabling timers. This project requires time to understand, implement and test, but might allow us to consider reenabling SharedArrayBuffer and the other high-resolution timers as these features provide important capabilities to the Web platform.
Glad to see people finally waking up to the fact that high accuracy timers are security vulnerabilities on most synchronous systems.
The cache attacks never received as much publicity for some reason (the papers being harder to read may be one). I do wonder what would have happened if this hadn't come out at the same time as meltdown. It's quite possible it would have been brushed under the carpet yet again.
There were token 'fixes' to these things (coarsening timers) in the past, but they never worked and everybody involved knew they wouldn't work. The introduction of features like SharedArrayBuffer revealed how (un)seriously they really took the problem. They knew it could be used to implement high precision timers but it got added to browsers anyway because it was central to the project of making the web an application platform.
They perceive a need to allow high precision timers (or features that can be used to implement them) because without that the web won't be able to do a lot of things that are possible in native applications.
I'd like to think that this is the moment that browser vendors come back to their senses and rethink what they are doing but I doubt it. Google is a multi billion dollar company based on the web as a platform, and running untrusted javascript on other peoples computers. Dropping the idea of the web as the platform to end all platforms would be an existential crisis for Mozilla. They are locked into this madness with no way to stop that wouldn't effectively be corporate suicide. Expect years of half hearted 'fixes' which don't fix the problem.
[1] https://news.ycombinator.com/item?id=10455735
[2] https://www.nds.rub.de/media/nds/veroeffentlichungen/2014/07...
[3] https://contextis.com/resources/white-papers/pixel-perfect-t...
[4] https://www.mozilla.org/en-US/security/advisories/mfsa2013-5... and https://bugzilla.mozilla.org/show_bug.cgi?id=711043
This is very insightful. A lot of web standards are attempts to add what traditional native-app capabilities to the browser (e.g. 2D Graphics => Canvas).
What's very clear now is that native-like web capabilities imply native-like vulnerabilities, but delivered over the network.
Browser vendors rethinking what causes a user to launch complex Javascript on tabs they visit (or worse, in invisible iframes) would be a great start. It was one thing when all Javascript could do is style and manipulate the DOM. We can now compile vim into Javascript and that demands a completely different response.
This is not to stop progress on web standards, but if the web community takes this opportunity to level-up their security practices, it'll help them (and web application developers and users) in the long run.
>The idea behind the TTT measurement, as shown in Figure 4.4, is quite simple. Instead of measuring how long a memory reference takes with the timer (which is no longer possible), we count how long it takes for the timer to tick after the memory reference takes place. More precisely, we first wait for performance.now() to tick, we then execute the memory reference, and then count by executing performance.now() in a loop until it ticks. If memory reference is a fast cache access, we have time to count more until the next tick in comparison to a memory reference that needs to be satisfied through main memory.
>TTT performs well in situations where performance.now() does not have jitter and ticks at regular intervals such as in Firefox. We, however, believe that TTT can also be used in performance.now() with jitter as long as it does not drift, but it will require a higher number of measurements to combat jitter.
So, what stops this method from working, even with 20µs resolution performance.now()?
The issue cannot be fixed without disabling speculation entirely on current hardware.
That will reduce the data rate of this particular covert channel, not prevent the attack altogether. Even adding random noise would not rule out the attack.
Limiting JS execution resources, and in particular CPU cycles, will actually stop instead a whole swath of timing and resource-dependent attacks.
Please put your thought and chime into this thread:
https://bugzilla.mozilla.org/show_bug.cgi?id=1414675
Allowing infinite resources for remote programs is something we don't even do for local programs. Giving a ceiling to the JS runtime is a sound reasoning.
Please comment on the bug tracker!
Is chrome or edge affected?
Should the SharedArrayBuffer mdn page be updated; or maybe moz://a will fix it quickly enough to allow it on by default again.
By the looks of it, this will affect a lot of webgl code.
[0]: https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
> In line with other browsers, Chrome 64 will disable SharedArrayBuffer and modify the behaviour of other APIs such as performance.now, to help reduce the efficacy of speculative side-channel attacks. This is a temporary measure until other mitigations are in place.
https://sites.google.com/a/chromium.org/dev/Home/chromium-se...
Of course, that will likely be a browser-wide setting, so telling others to enable it will put them more at risk of these attacks.
(As mentioned in another comment, the more robust solution is to use window.postMessage as a fallback, although it depends on exactly what you're doing.)
I can't think of something that needs microsecond level timing though.
I realize there's a whole group of people who want to make the web browser as powerful as native Apps but I for one do not.
The option to blacklist an app from such requests should also be present.
If the app itself absolutely cannot perform some function w/o high-precision timing, that becomes its problem to communicate to the user.
I've made a practice of denying application permissions, and deleting apps that make such requests. I'm moving to deleting Android entirely, which has proved overall to be a poorly-performing and functioning virus and attack vector, at cognitive, social, economic, political, and other levels. Also, frequently, software.
I hope they find a better solution that keeps shared memory working.
Certainly shows why it’s useful though.
I see over 100 pages of commits on github referencing it, which tells me it's not exactly rare?
So if you want the next Photoshop or Final Cut Pro to be browser based we’ll need to make it work securely.
That seems feasible without significant performance loss in most cases.
There is still a fingerprinting issue though that might not be possible to fully remove without huge performance cost.
I personally used it for a experimental hackathon project in April (in Chrome Canary, behind a flag), and I believe it was necessary for one of the aspects of the project, so it's a shame to see it disabled, but my impression is it's more "future web tech" than "the way the web works today".
[1]: https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
After a bunch of google searches, I have not found anything suggesting how accurate clocks are for clients besides Charlies algorithm [0] (I want to know UTC time/epoch... whatever from each client so i can compare them) which I am concerned about adopting due to the wiki page assuming a high quality network.
> Specifically, in all release channels, starting with 57:
Firefox 57 was released nov 14, did it already include those changes?
Though cpu cleans registry, it doesn’t clean up cache. So later through timing you can gain knowledge about something you shouldn’t know.
That is why, reducing timers precision can help
If there ever was NSA hardware bug this is how it would look like. I doubt we’ll ever know.
Hypervisor Guest OS
Is there double the performance hit of all these attack fixes?
I can't invoke SharedArrayBuffer for sure but what about the resolution of performance.now()?
E.g.,
(function() {
var global = (function() { return this; }).call(),
_Performance = global.Performance,
_performance = global.performance;
global.Performance = /* ... snip ... */;
global.performance = /* ... snip ... */;
})();
[Edit]: This could be injected by a browser-plugin on the user's side, as well.