I reverse engineered Google Docs to play back any document's keystrokes (2014)
features.jsomers.net
features.jsomers.net
Slides and Doc are extremely closely related projects, sharing lots of code (and even sitting next to eachother, at the time at least).
Our entire team was impressed with the details of the operational transform protocols that the author was able to reverse engineer, and for the most part get correct. IIRC our engineering director ended up calling the guy up for a lunch.
Just curious, how does Protobuf make it easier? Easier than Json? Protobuf doesn’t even include field names or readable descriptors - unless of course you are talking about reading client side JS
Messages and submessages and arrays make figuring out what data is what very easy for the most part.
The eng director may get some cool ideas from talking to him. The guy might get cool ideas from talking to the eng director. It’s also a free lunch.
This guy may want a job at Google in 2 years. He can then email the eng director for a referral.
I've worked for startups and even mid-sized publicly traded companies where managers with sufficient seniority were allowed to extend offers on a whim and it always resulted in friends of said employees being hired. Inevitably, entire teams and even business orgs would be filled with friends or fraternity brothers and their work performance would vary greatly.
No one likes interviews but it's crazy that I keep seeing HN users considering this an acceptable practice.
I hire developers and researchers for my company and we are very selective in who we choose to interview (small company in a bleeding edge field). We only consider engineers with relevant, impressive expertise and academics with solid published papers. Still, many of those interviewed don't receive offers and it's not because we ask them to recite some esoteric solution to a riddle. We have called candidates in just because they have promising github projects related to our work, only for these individuals to completely flop when questioned about basic fundamentals. Our interviews were extremely lenient and informal in these cases and we gave multiple chances to some.
https://observablehq.com/@jsomers/we-need-more-tiny-knowledg...
https://en.wikipedia.org/wiki/Operational_transformation
The new alternative to this approach for collaborative editing (and large scale distributed database concurrency without locks) are CRDTs: https://en.wikipedia.org/wiki/Conflict-free_replicated_data_...
Edit to answer my own question, looks like the original G Suite was in August 2006, but at the time it was just GMail, Calendar, and Sites. Docs was added in early 2007. Wave was a thing under Google from 2009-2012, but was never part of G Suite.
Also have a prototype for what at the time was an improved Outlook to Google Calendar sync (e.g. it maintained coloring by converting color categories in Outlook to Google subcalendars of matching colors) but was hesitant to productize it one reason being I wasn't sure if Google Calendar would live long enough to make it worthwhile.
I wonder if a reputation for prematurely killing their products is harming Google's ecosystem-building efforts.
I didn’t use google calendar for repeating yearly appointments for almost a decade because I thought the same thing. Glad to know I wasn’t alone!
"Yes, we know what the final contract says, but based on your edit on January 7, 2003, at 5:21 pm and 0.210403 microseconds, your real intent was ...". "And then on June 12, 2006, you took out some words that have a bearing on this lawsuit -- obviously you were trying to cover things up."
I was quite surprised to learn that during patent litigation for example, the history and changes made to your application do matter. I assumed that the final patent as issued is what matters and everything else is irrelevant. Well, lawyers have made it matter, and now things will be 10,000 times worse when they want to find something to argue.
So great, now we have to worry about how someone might interpret every single edit rather than the final document.
Sure, until that feature is built into almost everything and it becomes mandatory to use it. Filing an income tax return? You must use this online system that captures every edit and every keystroke. There is no other way to file.
When people complain about lack of privacy when using credit cards and bank cards, would you dismiss it by saying, use something else -- just use cash? I could list a hundred things that you can no longer pay with cash.
Just more of a reason to write your document in notepad or on a collaborative system you control. You don’t control the hardware, then someone else can walk off with the software (or output of the software)
Almost every legal (or otherwise collaborative) document has “track changes” enabled on its source document for the last 25+ years, which captures every keystroke in Word or WordPerfect.
I cannot recall there’s ever been a subpoena for the track changes source. It’s generally immaterial. Negotiation proceedings are not the contract, the contract is what was signed and both parties have a copy as paper, image, or PDF.
Computer forensics are often about determining whether a final image was altered after mutual acceptance. This is also why important documents get notarized.
Obviously no lawyer worth their salt is going to clarify ambiguity via email instead of redlining the contract with a change, but post-contract ambiguity can often be discovered.
While not your lawyer example, a very common example developers would follow while infringing on code. I can think the same for someone writing a manuscript or other large work of writing
How do you know you're doing something really funky? Even your lawyer doesn't want to put it in writing.
Approvals ends up solving some of this by freezing changes to approved documents.
Using this data alongside ML, you could do things such as:
- Identify people based upon their typing style easily, given so much data about them
- Find a lot of secrets that were accidentally pasted into document including passwords, keys, etc
- Try to model the emotions people were feeling as they typed things
- Much, much more, given some data and time
[0] - https://dl.acm.org/doi/abs/10.1016/S0167-4048(03)00010-5
https://github.com/mturnshek/keystroke-timing-identifier
I did not test it on a large group, but it worked pretty well for a group of ~8 styles and minimal data. I believe a deeper model with enough data could identify on a large scale through timing alone.
Have you noticed how, for example, you can identify a guitar player easily, no matter what guitar they're playing on?
When guitar players are first learning or are transcribing, they often slow the music down. If you slow it down enough, you can hear a pattern of the player slowing down and speeding up (behind and ahead of the beat). I believe that has a lot to do with that instant-identifiability, even though it is so hard to detect at speed.
https://webapps.stackexchange.com/questions/73080/is-there-a...
Though, I suspect that's intended as a generalization of the various VCS options they support.
---
I do think it'd be good to move away from the negative connotation associated with blame. I want to know who to ask about some code, not who to blame for it breaking.
discussed at the time: https://news.ycombinator.com/item?id=8562483
No shit. I'm stopping to share anything on Google docs now. Who knows what secrets are here somewhere ?
Usually the expectation is that as long as you don't hit "send" it's fine, but this is clearly not the case here.
... But seriously, what? What is a virtual on-site meant to be?
Maybe the name "virtual on-site" is supposed to convey to both interviewer and interviewee that "this is for serious".
This may vary by hiring committee, but on mine we don't look at feedback from the phone screen interviews at all.
As an interviewer, I will give a slight benefit to the doubt for “screens” because I understand some candidates don’t take it as seriously as the main interview.
I would have hoped that it cuts off diffs/undos to a certain point in history.
It's interesting to think about what is it about keystrokes that tend to make people worry about privacy. I don't think it's as simple as "they are storing a lot more data than I assumed". It can't just be a matter of degree. The reason must be broadly along the lines given in the Edgeworth quote.
Storing and indexing keystrokes is unique, and it might even be useful, yes. However, this does not at all seem safe to store without user knowledge for every document created. What could possibly go wrong....
File -> version history
That is just for the file menu. There are 7 other menus each with their own long list of options.
It's probably one of those situations where you'd write a "Diversion Report" for the security team where you talk about limited assets at rist and a huge investment in engineering resources.
FYI - this is a kind of 'biometric' leakage eh.
Your typing style is probably as unique as your finger print.
It's not shockingly bad, but probably something, under thoughtful review, they should take care of.
Edit: by 'no use case' I mean, no use case for long term history. Obviously there's a good use case for 'very short term' individual keystrokes.
Also Edit: the answer to this issue is not 'Don't worry about it, it's a low risk problem'. This is the 'slow boil' SV way of thinking that is seemingly benign, but not acceptable in the long run. The answer is: don't store personal data that is unnecessary. It's that simple. It doesn't diminish the product a single bit.
Storing the change vector is essential for collaborative editing in the moment. But the UI should highlight a Flatten feature as part of explicit Saving or Sharing actions.
For example, as far as I am aware this isn't information that is included in Google Takeout, nor is it included in any SAR that you may ask for. There is no ability to have a list of comments you've made and a choice to delete them. Even if you delete your account I am not sure these pointers are actually removed from the database.
Google collects so much data across so many vast areas I am not sure how they can ever be truly GDPR compliant and in accordance with the law as it's written.
For example, how do I delete my Google Fit data? It is seemingly impossible, of course this is where you have to ask them to. How about deleting every single edit history across all of my documents, or a specific edit in a specific document? Will they be able to remove that unique entry in the edit history?
I do know they store edit history for an indefinite period of time, I've accessed documents 6 years old (never edited again, unopened for many years) and the edit history is still perfectly intact. I wonder how many PB's of data Google has of pure and utter data-rot, and why they don't have systems in place to clean that up (especially old, inactive data like document comments or revision history).
I imagine every single record Google creates and logs must be liable to the GDPR, so of the trillions of data points they have on their systems it must be a pain to try and ensure any of them can be modified if required to do so.
However, Google has demonstrated their bad faith and unwillingness to comply with the GDPR with their non-compliant consent popup so I would be very surprised if they suddenly decide to spend resources on modifying this feature to make it compliant.
Look at the image from the article. Ughh. Maybe 3rd of the monitor is actually used to display text
https://cloud.githubusercontent.com/assets/21294/4930884/70f...
CSS permitting
Safari on mobile devices uses a visual zoom which is like scaling up an image, while desktop browsers usually scale up CSS units and reflow the layout with the bigger dimensions.
it is not true that all browsers can zoom in. there are some cases, e.g. mobile platform with stock browser, where CSS dictates no zoom and the browser honours this and therefore cannot zoom
that might be due to browser implementation but ultimately CSS dictates it
I have no interest in mobile Safari but its non-vector based zoom sounds useless, and desktop zooming is unreliable
Also, just realized that screencap isn't even Google Docs.