And this is good because it removes the illusion that they can speak freely and saves them from repercussions coming from eventual de-anonymization.
And this is good because it removes the illusion that they can speak freely and saves them from repercussions coming from eventual de-anonymization.
What next? A publicly available ML web crawler that analyzes speech and belief patterns, triangulates them with metadata, and returns someone’s identity with 99% confidence? I’m not naive enough to believe the government won’t build that, but for that to be freely available is just a recipe for chaos. It should be illegal, and it should be illegal to distribute many intermediate tools as well.
Close but no cigar.
But seriously, correct me If I am wrong but if I were to propose a solution to defeat this `ML web crawler' confidence in it's results, then the first thing I would do is to feed the internet with invalid data. How does ML-based analyzers deal with this kind of data set?
I apologies but that's disappointingly vague..
The way I see such a thing working is that you train it to identify what makes your writing unique . Everyone has a highly esoteric writing style, akin to a thumbprint.
An ML optimizing for uniqueness can identify:
- Relative frequency of certain words
- Diction
- Written tics
- Distance between certain words that the author tends to cluster together
- Mean clauses per sentence, clause variance
- Symbol usage
- Affect
- Interests
and abstract patterns that we haven’t even recognized yet. You can limit the search space at first by pointing the algorithm at certain websites and sub-sites that you’re fairly certain the person uses, but eventually I think even that will not be necessary.
Aha. Indeed. I would prefer to have someone actively working in ML to weigh in as well.
Would you be interested in PoC'ing it out for such a trivial project?
I don't think you can, not without special means.