Nvidia GPUs can break Chrome's incognito mode
charliehorse55.wordpress.com
charliehorse55.wordpress.com
Basically, this issue is not restricted to NVidia GPUs or specific operating systems - This can be reproduced on Windows, Linux and OSX. Basically the concept of memory safety does not exist in the gpu space - which is the reason why the webgl standard is so strict about always zeroing buffers. The issue of breaking privacy and privilege boundaries on a multiuser system is very real, and there is no workable solution. This seems to be one of those problems where a lot of people are aware, but no one is sure how to fix it and so it just stays how it is.
The fix is pretty simple, the GPU manufacturer just needs to update their driver to zero the VRAM like an OS would with RAM.
Seems like an obvious way to increase the take, in any event: "Chip X has a security flaw. Chip Y does not. Buy chip Y now or Evil People will own-zor all your cash-zors."
Yep, all the time. http://i.imgur.com/8k3ffLa.png
Version: 359.00 - Release Date: Thu Nov 19, 2015
Version: 358.91 - Release Date: Mon Nov 09, 2015
Version: 358.87 - Release Date: Wed Nov 04, 2015
Version: 358.50 - Release Date: Wed Oct 07, 2015
[1] http://www.nvidia.com/drivers/betaIn a sense, BSODs aren't anything special -- all a BSOD means is that some code running in kernel mode has crashed or raised some exception that went unhandled. The same thing, when it happens in a user-mode program, gets you the error dialog box 'Program has stopped working'.
So the causes of BSODs and user application crashes are the same. The reason Windows has BSODs is that it's dangerous to keep the system going when something in kernel mode crashes. Things running in kernel mode have access to everything (think - all memory) and are deemed important enough to the operation of the whole system that a crash in one of those is a significant event that's worthy of special logging and rebooting. You can't guarantee, for example, that a display driver crash hasn't corrupted other parts of memory, cuasing potential for data loss if the system were to continue operating.
So, back to the original point. Device-driver BSODs from the big vendors are probably rare enough in general that you should suspect a hardware problem or glitch if you suddenly see one out of the blue. Graphics drivers, given their complexity, are a bit more prone to crashing though. Also, things running on the system can interact and cause the driver to crash.
Windows has lots of infrastructure in place for making sure device drivers behave safely. There's also good facilities for figuring out exactly what caused a BSOD beyond the usually cryptic-looking error code you see on the screen.
Resplendence WhoCrashed is handy: http://www.resplendence.com/whocrashed
Though if you really want to dig deep, the tools with the Windows SDK (particularly WinDbg) can let you achieve the same thing; they are developer tools though, so targeted more to that audience.
EDIT: Just to add in answer to your original comment, big-vendor graphics drivers are VERY often updated. I'd bet they're the most often updated drivers on a system. There are myriad reasons for this, both technical and competitive. That doesn't mean that long-standing problems are necessarily fixed, but both AMD and Nvidia have very regular releases with fixes and performance improvements.
Two examples that demonstrate this point well:
- There are various tools out there that you can use to perform a live memory capture on a Windows system; not just doing a memory dump of a single process, but doing a live memory dump of the whole system without having to halt or reboot. I've used one of these and it works by loading a 'driver' component when it is run that does the memory capture from kernel-mode (it requires Admin elevation to run, obviously).
(For examples, see: http://www.forensicswiki.org/wiki/Tools:Memory_Imaging
I don't remember if it was one off this list that I tried though).
Another example: A friend of mine had a system that would inexplicably BSOD if he left it running for a long while, unattended (especially overnight). We initially suspected perhaps a heating issue (it was a small Intel NUC). After setting up for full memory dumps and then analyzing them after a BSOD occurred using WinDbg, we actually found out that the BSOD was being caused by a kernel-mode component of the anti-virus suite that he had installed -- I think at the time it was BitDefender, but not sure. When he consulted the AV vendor support website, I believe it turned out to be a known issue with a fix.
On my own systems, by far the largest cause of BSODs (of the few that I've seen over the last couple of years) has been RAM going bad. These typically manifest as BSODs out of the blue that seem to come from different modules each time they happen, or they come from a module deep in the system that 'shouldn't' have crashes. My personal rule is, if I see one, be vigilant. If I see another one, reboot and run MemTest86.
Technically this should be a OS responsibility, but practically the vendors have made that all but impossible.
You seem to be using OSX (judging by the screenshots).
You should be aware that OSX's GPU drivers are written by Apple (or at least they act as the gatekeeper). You need to send the bug report to them. And perhaps update the title of your post to "Apple breaks..."
I've seen this exact same behavior on OSX with an Intel GPU.
As mentioned elsewhere in this post, e.g. Windows WDDM drivers require memory to be zeroed out.
Well, MMUs on GPUs have been standard for a while. They just need to use them properly and at least have an opt-in mechanism to enforce zeroing of newly allocated pages.
Your porn browser declares it wants to share its window buffer with the WM, the kernel maps the buffer into WM's GPU address space and now WM's shaders can read from this buffer.
A lot of consoles had no concept of the CPU being the final arbiter of the system being running or halted—things like the PPU and SPU and so forth would just continue to merrily loop on even if the CPU halted, or was executing a CPU-reboot instruction.
You can see this in many NES and SNES games, where the game will "soft-lock": the CPU crashes, but the music (being a program running on the SPU) keeps playing, and the animations on the screen (being programs running on a PPU, or dedicated mappers feeding it) keep animating.
But this isolation can also be used deliberately, especially where the framebuffer is concerned. Since systems up until the 1990s-or-so had extremely small address-spaces ("8-bit" and "16-bit" are scary terms when you're trying to write a complex program), and since console games were effectively monolithic unikernels (even the ones "running on" OSes like DOS: DOS would basically kexec() the game), frequently a game's ROM or size-on-disk would exceed the capacity of the address space to represent it.
The solution to this was frequently to actually have several switchable ROM banks or several on-disk binaries, and to effectively transparently restart the system to switch between them. This isn't equivalent to anything as "soft" as kexec(); you wanted the CPU's state reset and main memory cleared, so your newly-loaded module could immediately begin to use it. Any state you wanted to preserve between these restarts would be stored on disk, or in battery-backed-RAM on a cartridge.
This is how C64 games managed to fit a rich-looking splash screen into their game: the splash screen was one program, and the game was another, and the splash screen would stay on the framebuffer while the C64 was rebooting into the game.
This is also the architecture of games like Final Fantasy 6 and 7: when the credits list developers' roles as something like "menu program", "battle program", or "overworld program", those aren't mistranslations of "programming"—those were literally separate programs that the console rebooted between, hopefully finishing in the time it took for the console to finish executing a fade-out on the PPU. When a battle starts in a Final Fantasy game, the CPU has been reset and main memory has been entirely cleared; everything the game knows about to run the battle is coming from SRAM. (And the reason the Chrono Trigger PSX port feels so laggy is that CT has this architecture too, but reading a binary from a CD takes a lot longer than switching a ROM bank. Games designed in the PSX era took that into consideration, but ports generally didn't.)
I've always thought it'd be cool to re-introduce this idea to game development, through a kind of abstract machine in the same sense as Löve or Haxe. You'd have a thin "graphical terminal emulator" that would contain PPU and SPU units and the SRAM, and would be controlled through an exposed pipe/socket (sort of like a richer kind of X server); and then you'd write a series of small programs that interact with that socket, none of them keeping any persistent state (only what they can read out of the viewer's SRAM), all of them passing control using exec().
(There's another thing you'd get from that, too: the ability to write strictly-single-threaded, "blocking" programs that nevertheless seemed not to block anything important like frame-rendering or music playing. You know how "Pause" screens worked in most games? They just threw up an overlay onto the PPU and then stuck the CPU into a busy-wait loop looking for an 'unpause' input. The game's logic wouldn't continue, but the game would still "be there", which was just perfect. This also allowed for "synchronous" animations—like Kirby's transformations in Kirby Super Star, or summoning spells in the FF games, or finishers in fighting games—to just run as a bit of blocking code on top of whatever was currently on the screen, without worrying that something would change state out from under them.)
Or was there secondary memory besides the RAM and disk that allowed for data to be passed between resets?
In later consoles that had their own MMUs, like the PSX, this wasn't a full hardware feature anymore, but rather a simulated convention. You'd "reboot" by dropping all your virtual-memory mappings except one, then asking the disk to async-fill some buffers from the binary you wanted to launch and then mapping those pages and jumping into them on completion. (Basically like unloading one DLL and then loading a different one, except you're also forcefully dropping all the heap allocations the old DLL made when you unload it.)
In both the hard and soft implementations, the "volatile SRAM" page could be thought of as basically a writeback cache for the state in the actual battery-backed SRAM. You wouldn't want to do individual byte-level writes to SRAM (writes to SRAM were slowwww), so when the game booted, you'd mirror SRAM to your state-page, and then update the state-page whenever you had something you wanted to persist—finally dumping it out to battery-backed SRAM when the player hit "Save". Basically, most games were "auto-saving" from the beginning—but they were auto-saving to volatile memory.
But even games that had true "auto-saving", like Yoshi's Island, still kept an "SRAM buffer" like this; the write-to-SRAM event was just managed as a sequence of smaller bursts of memory-mapped IO done by modules that sat there playing music+animations and not doing any logic, like YI's "field/title card" submodule when re-entered from a loss-of-life event, or its "field/score card" submodule entered by completing a stage. If there is ever what seems to be a "pointlessly long" animation in an auto-saving game of that era near a state-transition, it's probably by design, to cover for an SRAM cache-flush. (The fact that flashy "rewarding" animations turned out to also be good game design, favored by slot-machine and casual-game designers the world over, is mostly coincidence.)
---
ETA: when you had neither kind of SRAM, but still wanted to preserve some state across a reboot, what could you do? Well, write it to video memory, of course!
On a system with a PPU, the PPU owned the VRAM; it wasn't the CPU's job to reset it, but rather the PPU's. On a system with only a framebuffer, nothing owned the framebuffer (or rather, the framebuffer was an inherited abstraction from the character buffers of teletypes: the "client" owned the framebuffer, so it was up to the "client" to erase it. Restarting a mainframe shouldn't forcibly clear all its connected teletypes; disconnecting from an SSH session shouldn't forcibly clear your terminal, but rather optionally clear your terminal as a way your TTY driver is set to respond to the signal; etc.)
Either way, if you could get your data into VRAM generally or the framebuffer specifically, you could very likely read it back after reboot.
VRAM is also where many "online" development suites—the BASICs and Pascals of the time—expected you to write exception traces. Rather than trying to "break into" a debugger (i.e. cram a debugger into the same address space as your software), you'd simply have your trap-handler persist the stack-trace to VRAM, switch your tape drive over to the devtools disk, and reboot. The "monitor" would load, notice that there's a stack-trace in VRAM, parse it, read the pages it mentions from your program (now on the slave tape) and display them.
Another effect of the screen RAM being regular RAM was that you actually could run programs in it while it was being displayed. You could watch a program run in the literal sense. This was often used by unpackers. The unpacker run in the screen RAM, filling the rest of the RAM. After its job was done the game started and filled the screen RAM with graphics destroying the unpacker.
EDIT: Found a video showing activity in the screen RAM of a C64. I'm not sure if this is the unpacker or this is really code execution, but it looks similar how I remember it.
https://www.youtube.com/watch?v=5nDzFsCEZT8&feature=youtu.be...
If it can persist across restarts... does a shutdown do anything differently? I'm not too familiar with GPU programming, but there has to be something we can do.
The truth is a GPU is an entire second computer attached via PCIE bus. As far as security is concerned this will continue to be a shit-show until we accept that fact and act accordingly.
OP's problem stems from the fact that some video buffer used by his browser hasn't been cleared after deallocation. At some later time this buffer either has been erroneously displayed instead of the game's buffer or has been allocated to the game which erroneously displayed it without filling with new content.
Forgive my ignorance (not a graphics programmer), but why can't the drivers simply clear the buffer before handing it off to another application?
Not zeroing a buffer cuts a big constant out of overhead. If you know which of the benchmarks will fail if you don't zero the buffer, you code in an "exception" so the benchmark doesn't fail and other applications act wonky. This isn't the "first time" nVidia has been caught doing this, see:
http://www.geek.com/games/is-nvidia-cheating-on-benchmarks-5...
http://www.cdrinfo.com/Sections/News/Details.aspx?NewsId=288...
To give an example, consider the difference between memcpy() and memmove(). On most systems memcpy() is as same as memmove() in the sense it works even when the source and destination overlap. Then you decide to optimize memcpy and to prevent bugs like this https://bugzilla.redhat.com/show_bug.cgi?id=638477 you will need to set a flag USE_MEMMOVE_INSTEAD_MEMCPY for every app that you know to memcpy between overlapped regions. You could call this "cheating" or could be a reasonable person and say something like this https://bugzilla.redhat.com/show_bug.cgi?id=638477#c129 instead.
As for the original question. I am not an expert on the windows driver model but have written some GPU drivers and can tell that a) memory release is asynchronous i.e. you cannot reuse the memory until the GPU finishes using it and b) clearing graphics memory from CPU over the PCIe is slow and drivers, in general, do not program GPU on their own. Taking these into account, it seems the driver is not well positioned to do this and this is a task for the OS instead.
Probably even Intel didn't anticipate protected mode with its 24 bit address bus when designing the 8086. 1MB was enough for everyone at this time.
See this fascinating post:
http://www.gamedev.net/topic/666419-what-are-your-opinions-o...
my point was that an optional post delete cleanup feature was added to the protocol ready to be used, which is a perfect example on how to evolve long term features. then I said the post cleanup feature for the GPU should sit on the driver, since the GPU driver is the one knowing how to talk to the hardware, as there is not a shared protocol between boards (except vga modes etc but those contexts are memory mapped and os managed) and knows when a clear is performed, since all operations go trough it.
Games Studios, IMO, should be made to fix their bugs themselves. They all have patching mechanisms these days, so it's not like it isn't impossible, or even unfeasible.
Currently the engine developer in graphics programming writes something and in reality he has no way of knowing what actually happens on the hardware (the API is just too high level to able to really know much). From there it is the hardware providers job to take out their own debugging tools and make sure correct things happen by having a custom code path in the driver.
Lower level APIs like DX12 and Vulkan remove the competitive advantage vendor dependent performance creates, so well-coded games can perform consistently with lower overhead across ranges of hardware without having to rely on vendors to patch in the shortcuts through their drivers.
Currently, it's like filming a movie with IMAX specifications, then finding out that at different cinema chains it played with quality aberrations because their projectors didn't truly follow IMAX spec. The chains can fix it, but you're already getting blamed for the movie's issues. However, for a little money, on your next film they offer to work closely with you to ensure it shows the way you intended in their theaters. And no, they can't just tell you how to fix it-- their projection technology is a trade secret.
The driver simply has to zero the buffer when the new OpenGL / graphics context is established. It's once per application establishing a context, not per-frame (the application is responsible for per-frame buffer clearing and the associated costs). At worst this would lengthen the amount of time a GPU-using application takes to start up and open new viewports, but that hardly seems like it would matter or even register on any benchmarks.
Why not just fix this in the browser? The real issue here is that this data isn't just being shared across processes but potentially with websites through malicious webgl.
You can't opt out of these security features on upstream/open source linux drivers either.
Now of course this won't insulate different tabs in chrome since chrome uses just one process for all 3d rendering. But GL_ARB_robusteness guarantees plus webgl requiring that you clear textures before handing them to webpages means that should work too. On top of that webgl uses gl contexts (if available), and on most hw/driver combos that support gpu MMUs even different gl contexts from the same process are isolated.
This really is a big problem with binary drivers, and has been known for years.
I was watching some adult material using Quicktime on Windows 98. A few hours later, I wanted to show my mom something on my computer. As it loaded the new video in Quicktime, the last frame of the porno sat there in inverted colours until the new video began to play.
I had closed Quicktime hours ago... what was that still doing there in memory?
Needless to say it was very awkward.
X crashed a lot back then, so everyone learnt pretty quickly.
http://webcache.googleusercontent.com/search?q=cache:_JGpv1r...
I had a similar problem on iOS. When I load Safari, there's usually a flash of the previous screen (probably cached as a PNG), then the page loads. I think it looks junky; I'd prefer a "loading" screen. It would flash the previous screen whether I was in private mode or not. So porn would flash on my screen. I didn't file a bug report or mention it on my twitter because I'm a little afraid of the reception. So, again, thank you charliehorse55.
edit: i said "cached as a PNG" but that's just what I thought prior to reading this article. it could be many things, including this bug.
Anti-porn crusaders aren't necessarily hypocrites. They also don't need to be hypocrites to be wrong.
Nothing makes people quite so alien as differing moral codes. I'm "showing my hand" as regards something that simply does not matter to the majority of people on this website. Trying to make a big deal out of it simply makes you look strange.
Apps like 1Password make their screen blank when they go into background to prevent this… Firefox for iOS does this too!
My e-banking does the same thing - as it requires password when switching back into tab, they show just a big bank logo to hide your account history from multitasker.
Edit: Also, "google chrome incognito mode is apparently not designed to protect you against other users on the same computer".. what? Isn't that the only thing it can and should protect against? It's not like it can protect against non-local users (i.e. HTTP network interceptions)
You don't need webgl for this kind of infoleak either, regular good old 2d canvas also supports allocating memory. It also supports reading the current state of all of the pixels in the buffer through Javascript, so if you have an exploit that gets you an uninitialized canvas you can easily send whatever memory contents you got back to your server for later analysis.
If the DOM element you draw has the same origin as your canvas it seems like (from my reading of the spec) you should be allowed to do what you describe.
WebGL being enabled by default is insanity in my opinion.
Imagine being able to open an incognito terminal to type commands that won't get saved to history or pollute what you already have.
unset $HISTFILEPeople considered it a profile switching for dummies, now no one remembers that browser can have multiple profiles.
As I understand it, this is not how frame buffers work. All they contain is the rastered data to drive the a given frame.
That's nonsense. For most users, that's exactly what it's used for.
I really think Google is dropping the ball here. I know it's not their bug, and they shouldn't have to work around it in an ideal world, but this is a pretty clear leak of data outside of private mode. It wouldn't impact performance in any noticeable way (you're closing the window anyway at this point), and would just be an extra safeguard.
Very short-sighted of them to ignore this bug. Perhaps we could ask distro maintainers to add patches for this to their builds of Chromium.
Modern operating system zeroes memory pages all the time. It is a security measure, and ensuring security is by no means a waste of time.
If this is the case there might be a compliance issue on nVidia side which makes me wonder if webgl is vulnerable also.
WebGL was amended to request a zero when provisioning or disposing of a buffer but it relies on the API which is handled by the driver if nVidia is taking some shortcuts to save time it might be possible to leech stale memory this way.
Which Windows version introduced this WDDM version? Could it be that OP is running an older version?
> WebGL was amended to request a zero when provisioning or disposing of a buffer but it relies on the API
This is indeed a tricky situation. All modern GPUs do "zero bandwidth clears" which means that upon clearing, nothing gets written to the actual framebuffer, the memory is just marked "cleared" (by writing some special bits to the L2 cache, for example). This makes it difficult to reason whether there's any sensitive content left in the framebuffer.
edit: nevermind, the OP seems to be using OSX, so it's not WDDM. Additionally, the OSX GPU drivers are written by Apple.
As far as WDDM goes 2.0 requires that for sure I'm pretty sure this was part of the original WDDM GPUMMU spec also but I can't really find those details anymore on MSDN since most of the pages refer to 2.0 atm.
It's basically heartbleed in your ethernet driver in 2003.
What's the purpose of incognito mode then? It doesn't protect you from your ISP, websites, or users on the same computer. I'm not sure what other use case there is.
I've seen the same behaviour on OS X with an Intel GPU: https://i.imgur.com/3fagsYx.jpg (screenshot of the contents of a browser tab - pretty sure it was Chrome. The Rooster Teeth page you can see parts of had been closed hours prior)
http://arstechnica.com/security/2014/08/stealing-encryption-...
In particular, they refute via counterexample the arguments that VMs or secure deallocation alone are sufficient.
There's a reason Linus Torvalds flipped the bird to NVidia. Here is a perfect example of the reason why closed source drivers suck.
If you were them would you take the performance hit?
If I (with training in how to design an OS and the risks of handing nonzeroed pages to another process) were them? It'd be part of my standard process for designing a memory repurposing library. But I can 100% understand how this mistake gets made; I wouldn't be surprised if it wasn't an explicit performance decision.
For example, Nvidia claims that the GTX 980 has a memory bandwidth of 223 GB/s. (1920 * 1080 * 3)/223e9 = 27us. Clearing all 4GB of VRAM would take 4/223 = 18ms. This would have a negligible impact on user experience in most cases.
I guess the driver could also erase memory in the background as soon as it is deallocated, with zero user impact.
And even if it's <0.1%, there is strong "optimization mentality" in those companies (because perf matters) so it's unlikely to happen in the current climate.
Also note that memory bandwidth is typically the bottleneck in modern games.
The best place to do this would be in the browser, clearing out any textures and buffers before deallocating them if the contents are deemed private.
If you're doing write-only operations, the CPU can queue them behind the clear. If you do a read, then the CPU has to wait whether you clear or not.
Latency doesn't matter. Clearing can be slotted in with other operations, such as first use.
GPUs generally have less memory, many fewer (but larger) allocations, and way higher memory bandwidth than CPUs, so it shouldn't be a problem for them to do this.
I am not familiar with GPU internals enough, but my understanding is that the GPU should be smart enough to know that a given texture or framebuffer will occupy n full pages, and so when either is written in its entirety, the zeroing only has to occur at the edges. (I would assume that the write would start on a page, but I don't know anything about GPU internals.)
Caveat emptor: I will reiterate I know very little about memory internals. It seems like a bigger issue is that GPU memory is not virtualized and all users get access to the same memory. It's as if three decades of understanding the utility of virtual memory were forgotten.
I think you're forgetting the most important thing here - GPUs are meant to be fast. Virtualization will add like what, an order of magnitude to the access times?
Correct. When the GPU page faults, it causes a CPU interrupt and the driver will handle the interrupt. It's not possible to resume execution on a GPU in a timely manner so the only option is to terminate the process that caused the page fault.
> so you can't really use the MMU for clever things like demand paging
Recent GPU generations support "sparse" or "tiled" memory where the GPU can detect if a load or a store would access non-resident memory and then act accordingly. This requires a specialized shader and some CPU-side logic to actually stream in the memory. This can be used to on-demand paging for textures and buffers as well as implement workarounds to reduce visual artifacts from streaming.
On a related note: Doom apparently does contribute to security proof of concepts though. http://www.techtimes.com/articles/15606/20140916/security-ex...
Which makes me wonder if the non-clearing memory issue exists for the printer's video driver and whether that could be used to retrieve something like a saved password or ssh key.
Maybe the system doesn't pass enough information through to the driver to let it determine this, though...
Even further this has very little to do with chrome. The only way chrome could actually fix this issue would be if it nuked the frame buffer when it released it. This is a fine idea, but if I was a dev in that context I would assume the OS would make stronger guarantees than that??
If anything, this is an edge case Chrome devs (and other developers) could protect themselves against if they were so inclined, but I'm not surprised they didn't assume they needed to protect against this.