How I reverse-engineered Google Docs to play back any document's keystrokes
features.jsomers.net
features.jsomers.net
As somebody who has worked full-time, over-time, and essentially in my sleep with Word files, PowerPoints, and highly sensitive bid documentation...yeah, I have to agree working in the cloud for this kind of stuff strikes me as career suicide. Maybe not in your turf, re: software development, but with respect to management, operations and marketing people, there should only be one person with the key to the kingdom. I'm not kidding about this, even if just talking internal development.
Also, this is why everything went out as a locked down PDF, unless explicitly mandated otherwise by RFP/etc specifications...and even then, Track Changes > Accept All Changes is gospel. Anybody in my line of work saw what the .GOV did with converting PDFs and simply redacting with a graphic over the text...yeah, that's why I'm a first-class proposal developer, because I've seen carnage yo.
I mean, I know how to really whip the donkey spit out of "Compare" and the available tools, but until you've seen just how sideways this stuff can go, there's always room for a professional pest...an expert Proposal Writer...there's so much learned in the trenches that we all have a certain cynicism built up, and that's the fuel that gets the junk in by deadline time.
(Of course, the big downside is that all your collaborators have to know LaTeX, and not use funny macros that the others don't understand)
Document history enables thoughtcrime detection, especially as it becomes more and more atomic towards actual keystrokes. I wonder how long it will be until typed and deleted text is used as evidence in a court of law.
Is this really a privacy breach? It's been obvious that Google stores revision history since it launched—you've always been able to access a thorough revision history in the UI itself...
Google Docs has always had a revision history tab.
I am curious to see who is (brave enough?) to show their writing process in all its glory.
Finished version: http://www.scribd.com/doc/156040925/Pink-Paint-Rain-Vernon-W...
I have about 7-10 print outs of the story with mark ups. Before switching to English & Creative Writing I was in Computer Science and pretty skilled with C++ at the time, so I can come up with a decent correlative:
Every version I printed was like running it through a compiler - this is to evidence that even if a piece of code or a line of text is functional, is it efficient, and, if at all possible, an excellent construction? These are subjective concepts that are innate in language. There may be several writers who can put words to paper directly from their head with no revision process - hell, it's what I do when I'm "practicing" on my IBM Selectric III to commit to writing in permanence (think before I type) - but for the greats, it has always been an iterative process.
To keep this from being all about me, allow me to provide a link to something that may be to your liking:
http://www.amazon.com/James-Micheners-Writers-Handbook-Explo...
In programming, we have editors that have strong support for writing because we know exactly what the semantics of code is, and what good operations for editing/refactoring are. With writing prose, the best we currently have is edit histories from wikipedia articles, which are on a much larger timescale and full of things that should not be part of the editing process (vandalism, NPOV wars, etc.)
I know I've accidentally done the "paste password" into those places accidentally at times.
I've also chucked a few passwords into IRC in times past. Fortunately non-essential stuff but really motivated me to sort out some better solutions (SSH keypairs, etc).
At least typical IRC clients don't transmit until hitting enter. Browser omnibars and Javascript can send away every keystroke as it happens. Now I want to search all my Google Docs for my passwords -- let alone other stupidity I don't care to share.
This password security measure is a Chrome extension that's required by company policy to be installed on all corporate machines. It watches all input (to browser forms) and if it detects your password being typed anywhere other than an actual sign-on page, then the next time you sign on successfully you're required to change your password. I believe there's also an e-mail notification, but it's delayed.
This is actually a pretty good password security technique, specifically because people often inadvertently type their password into the wrong forms due to focus errors, lack of caffeine, etc.
Because I can't think of an efficient way to do that that doesn't involve having the extension have access to the password.
I mean, you could store the password hash + length, but then you're securely hashing every single overlapping substring of what you enter, which is not exactly fast. Especially as KDFs are designed to be slow.
And if you store the password hash then you enable an offline attack.
Even assuming that the connection is secure (never a good assumption), that still means that there is a single point of failure. And one with drastic consequences.
I do agree about the single point of attack though. Perhaps you could do an asynchronous substring check locally when the CPU is idle.
Well, if you paste something in the location bar, presumably it'll already be on it's way to google (or whichever service handle autocompletion/suggestion)...
I'm currently using ClipMenu on OSX, which hits all the right needs for me. Anyone have a suggestion for Windows?
Thanks for Ditto!
(make me want to go back and see whether I can see other peoples comments .... hmmm)
Is this a typical way to write a magazine article? I wouldn't have expected so much time revising the opening sentences before getting the rest of the article in place. (But there's probably a lot of variation between writers.)
Ahh I'm now hyper aware of it and realise I've chopped and changed that ^^ first paragraph a ton of times.
For me, I usually spend a day or two just thinking about it in my head. Going over what would be a logical thought flow and things like that. When I sit down to actually write, I tend to have very few revisions. My first draft is much closer to the typical persons 5th draft (I think), but that's because I've been revising and editing in my head first.
I found that when I first started writing regularly I would spend a lot of time doing constant editing similar to the example shown in the linked article. This would become distracting and time consuming, then I'd forget other things I had wanted to say. So I found it better, for me, to just kind of write things in my head first and then sit down and write -almost more transcribing vs. "writing".
[edit: unintentionally demonstrating the need for more than one revision...]
I dislike revision-able software for a number of reasons. Privacy is the foremost reason. Yes, yes, "if you've nothing to hide, you've nothing to fear..." That old chestnut gets trotted out every time someone worries about security or privacy.
Since about 2000, I keep my documents in plain text only on an encrypted drive backed up several times over -- none of the backups are online, but I'm still good if my house burns down, my machines get stolen, you name it.
No, just no.
I really do like old school and I stick with it. I write most of my docs as text in either vi or nano. I neither want nor need formatting beyond the basics. My CV is the only document I own that has formatting, and I used LibreOffice for it. It's a single page.
I had a friend tell me my CV should be in PDF and locked down to prevent recruiters and others from changing it to something other than the original. I know recruiters are fond of changing things up without informing you.
Can you elaborate? Why would they want to change your CV?
Keeping everything in proper version control (possibly unzipped, to give usable diffs even for office document formats -- or in something like markdown) -- would at least rise the bar a bit -- there'd be different process for sending a single version of a file, and sending all versions of (all) [a] file(s).
I suppose if you're already running an internal mail server, you could just do filtering there, making sure no version/history-rich documents pass out that way...
[1] http://wave-protocol.googlecode.com/hg/spec/federation/waves...
[2] http://code.google.com/p/wave-protocol/ - wave protocol project (initiated by google, now maitained by apache) is the root from where gdocs adopted OT.
The author mentions that his system doesn't handle rich text, which is fine, but I'd just like to comment on how difficult of a problem handling rich text is. If anyone is interested in having a personal text-only replay editor, check out http://sharejs.org/ by an ex-Google Wave engineer.
As far as handling rich text, I've talked to the original co-founders of Writely (which became Google Docs), and I've also spent a good 8+ months on it as well. There are lots of tradeoffs involved, that diff-patch-match (as mentioned in the article) won't work on. Doc's ultimately expresses styles as applied ranges, rather than actual markup.
Point being, Google keeping every keystroke you've made is absolutely necessary for realtime collaborative writing.
ShareJS and Quill have actually been making some great progress on this though: https://github.com/share/ShareJS/issues/1
So will document retention specialists trying to foil laywers doing discovery.
So will hackers looking for sensitive information, and security specialists looking to avoid sharing sensitive information.
There probably really ought to be an "erase history" function.
I don't use Google Docs (and probably never will), but if I did, all those requests - "these /save calls every time I typed something" - would be enough for me to investigate why it's generating so much traffic. I'm using an OS that still has a useful network activity indicator icon, so I easily know when there's data being transmitted/received when there shouldn't be.
There's a line of thought that says those sorts of indicators are unnecessary and a distraction, and that maybe valid justification for removing them, but I can't help feeling like their removal is making users more unaware of what their machines are doing - and thus easier for companies to do things like this to them.
When most people hear "revision history", they think of the versions of the document that exist between explicit saves or periodic autosaves, and not extremely fine-grained per-keystroke activity logging.
This is not so different from source control except that merge conflicts are handled differently.
Differential synchronization [2] might be easier to implement, though.
They would then be able to implement sharing without allowing access to the history.
Alternatively, and more easily, they could record the state when someone is given access to the document, and not allow access to operations received by the server before then unless they are granted a separate permission.
I had no idea they stored the edit history.
A suggestion for further evolution would be an option to color text by the author that wrote it (for collaborated documents).
But I'm curious: Can one delete these kind of revisions displayed here? Those visible in the GDocs UI are only a few, mayor revisions (which may be troublesome in itself for people not knowing about it and sharing a document).
Basically, if you have all keystrokes with timing info, you've got all the keystroke dynamics required to establish an individual's biometric keystroke fingerprint. And that seems freaking scary to me, for a few reasons.
1) Impersonation
Anyone can grab that fingerprint from any shared file on Google Docs and then feed it to a program so that you can impersonate the author for various purposes... Be it typing a blog comment (harmless), or something more insidious like logging into a secure system using that type of auth system.
Another trivial example could be impersonating someone on a Coursera course, where they use such fingerprints for identification on paid / "signature track" courses, which allow you to get a verified certificate from a known university. They use a photo, but that can also be fed by a tweaked webcam driver. So there you have it, you can hire someone to take lessons and pass exams for you. Or fail you.
2) Anonymity
Anyone with access to a shared google doc can now get your fingerprint, and if they implement a similar record on another website, they can identify who you are. Maybe such fingerprints are not 100% unique, but they surely can be accurate enough to pick you from a crowd of anonymous commenters on a website, for instance.
You could also imagine that a software you already have installed could identify you using a similar approach.
In that case, surfing using a Tails/Tor VM and in the incognito mode of your browser won't help you that much.
I'm sure in a perfect world this could sound awesome: no logins required anywhere, just type in stuff and get automatically ID-ed and credited for what you type and say. In our world, that could be bad for some people.
Plus I can only imagine how bad that would be if companies started to include in their web and desktop apps EULA that you agree to share keystroke dynamics with them and that you auhtorize them to redistribute it to partners. BAM, a global commercial database of uniquely identified users, no matter what account or throwaway email they use. Forget cookies and stuff, that won't need that anymore.
Bit far-fetched of course, as that would require some effort. But it's not that much effort that it wouldn't be interesting enough for someone to do it...
Still, a bit worrysome, because it could easily be modified to track it. And for all we know, some sites could be doing that. Facebook was (at least at some point) listening to what you were typing in timeline posts even if you didn't actually decide to send them, so it wouldn't be surprising if some sites did that sort of stuff.
Interesting project idea...
Actually, Firepad does replay the history to display the current version on load (though it also has some snapshotting system to restart faster, but snapshots do not erase the history, they are kept in another location).