Show HN: Keytap2 – acoustic keyboard eavesdropping based on n-gram frequencies
github.com
github.com
In this post, I present my new approach which I think is quite interesting - it matches detected keystrokes with each other based on their sound similarity and converts the problem of recovering the unknown text into a problem of breaking a substitution cipher! The great advantage of this method is that it does not require prior training - it exploits the n-gram frequencies in the English language.
Would love to hear your thoughts and suggestions about this.
For waveform similarity, you may achieve better results by comparing the spectrogram (computed by taking the FFT of each "clack") instead of comparing the time domain signal.
Next, the best performing sound classification algorithms (at least in music note classification, which is a very similar problem to yours) utilise non-negative matrix factorisation[0] to do the comparison against a bank of note templates.
I recommend reading "Real-time Polyphonic Music Transcription with Non-negative Matrix Factorization and Beta-divergence" (Dessein et. al.) for a more comprehensive description of the algorithm[1].
I actually tried applying my algorithm to detect taps on different parts of my laptop for fun, so that I could send keyboard events when e.g. I tap the bottom right corner of my laptop, and it was quite effective. The approach is incredibly accurate for classification of short spectrograms.
Feel free to reach out to me if you want to discuss a bit more.
[0] https://en.wikipedia.org/wiki/Non-negative_matrix_factorizat...
[1] https://www.researchgate.net/publication/220723421_Real-time...
Edit: Also, this is great work. I don't mean to denigrate it at all.
I will definitely check the resources that you provided and see if I can utilize this approach in Keytap.
A quick side note, 16kHz is quite a low sample rate for your use case, since you won't be able to capture frequencies over 8kHz (the Nyquist frequency), and key clicks at least intuitively I'd imagine to contain a lot of detail at higher frequencies. I have no data to back this up though!
I switched down to 16kHz from 48kHz, because I found that with my approach there wasn't much of a difference in the key-matching performance. At the same time, I gained computational performance.
At some point, I was looking at the spectrum of the key sounds and I think that almost the entire signal is in the lower few kHz. It might be just my keyboard - not sure.
I still have the tool, so I can easily check if there is some high-frequency portion that I am missing.
Edit: I didn't mean for this to sound so closed minded. I don't know that much about more recent audio processing practices and I think you probably know things I don't. If you have any reading about wavelets vs DFT I'd be very thankful and interested.
A Discrete Fourier Transform (DFT/FFT) balances frequency vs time resolution based upon its parameterization. In this application, I think that frequency differentiation is less important than time differentiation for events. To get an advantage from utilizing spectral analysis, I think that the STFT can be used with very small overlapping analysis windows. (https://en.wikipedia.org/wiki/Short-time_Fourier_transform) However, due to the importance of transients I would consider use of wavelet transforms instead (https://en.wikipedia.org/wiki/Wavelet_transform) since they can resolve temporal characteristics better than a DFT.
Like the other comment suggested, a frequency analysis might help.
Here is a good explanation you could quickly use in C++:
Otherwise popular libraries like openCV have tools to convert signals, but most are focused on imaging (2d signals).
Even if the keystroke sounds were identical at source, the acoustic environment between the key and mics (key position, shape of the surrounding keyboard area, position of hands, screen etc.) adds its own unique reverb signature too.
Speaking of mics - a lot of laptops have stereo mics. That's enough to start locating the acoustic source spatially by phase, in addition to the other characteristics.
See also: recovering speech by filming the vibrations of a plastic bag on the other side of soundproof glass.
https://news.mit.edu/2014/algorithm-recovers-speech-from-vib...
I got a ReSpeaker Mic Array which has 4 mics and have been planning to do some experiments with it.
Presumably the trick would be finding a non-annoying way for the sounds to differ (perhaps if it was beyond human hearing). Of course it would be totally insecure, so not safe for Twitch videos without redacting the audio (but a tool could know what to remove that detail too)
https://en.wikipedia.org/wiki/Remote_control#Television_remo...
My grandma's TV had one of these - you'd press a button and it would basically mechanically impact a tuning fork to make the needed frequency.
"Cracking Passwords using Keyboard Acoustics and Language Modeling"
http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.186...
The same things that drive security researchers to explore different security exploits. I don't condone anyone using this for malicious purposes.
That said, it significantly lowers the barrier for entry, much like script kiddie hacking.
More recent one that works well over VoIP: https://dl.acm.org/doi/10.1145/3365366
I hope you realize what you've created (and made available to every person in the world).
If your coworkers complain, tell them it's a security feature.