No need to go through the speakers, you could detect keystrokes, map them to transients in the recorded audio stream and use a noise gate to reduce their volume (at the cost of a little added latency).
Bonus point: this would allow you to speak and type at the same time with minimal reduction of speech comprehension, especially if using a multi-band noise gate (that acts on various frequency bands independently). It's a technique we used for dynamic de-essing (removing plosive "S" sounds in post-prod recordings) in a previous company.