With 85% accuracy, 85 of liars are flagged as lying (correctly), and 15% x 9900 = 1485 of non-liars are flagged as lying (incorrectly).
Thus, a bit more than 5% of people flagged as lying are actually lying, while nearly 95% of people flagged are innocent. This is not even taking into account the possibility that hardened criminals might be less nervous than somewhat anxious normal people.
Enjoy your border crossings, everyone.
EDIT: fix italics
EDIT to add: And that's after they get the accuracy up to 85%. And unless accuracy is defined somewhat differently.
This is not what accuracy means.
85% accuracy just means that 85% of all the decisions the system makes are correct. A system in such a setting, where a single false negative matters a lot more than a single false positive (which would simply be handed over to a human for further investigation) would necessarily be tuned for extremely high recall at the cost of precision. In other words, it would often flag innocent people for further investigation (as you've said), but it would almost never clear people that should've been flagged.
Let's make the spherical cow approximation that "a lie" is a fully defined concept, then we have 4 conditional (bayesian) probabilities:
P( "sincere" | sincere) The probability a sincere person is reported as "sincere".
P( "lying" | sincere) The probability a sincere person is reported as "lying".
P( "sincere" | lying) The probability a lying person is reported as "sincere".
P( "lying" | lying) The probability a lying person is reported as "lying".
The first 2 probabilities should sum to 1, and the latter 2 possibilities too, so we have 4-2 = 2 degrees of freedom. A reported "accuracy" tells us nothing without knowing the distribution of liars and sincere people in the test group..
https://en.wikipedia.org/wiki/Accuracy_and_precision#In_bina...
Unless I'm mistaken (and that's possible, I've changed my opinion twice now), my example outlined above is
- conceivable, and
- has 85% accuracy (85 people correctly identified as liars, 85% x 9900 = 8415 correctly identified as non-liars, thus a total of 85+8415=8500 of 10k total "accurately" identified), and
- still only 5% or 6% of flagged liars are actual liars.
EDIT to add:
And if the system is tweaked as you suggest, to very rarely fail to flag a liar:
- suppose it correctly flags all 100 liars as liars
- suppose accuracy is still 85%, thus 8500 people in total classified correctly
- thus 8400 non-liars flagged correctly, and the remaining 1500 non-liars flagged incorrectly
Now still only 6.25% (100 of 1600) of people flagged as liars are actually liars. Thus, even with the tuning you suggest, this remains.
(Note to self: 1. think 2. write)
You really have to compare precision and recall values to know if the accuracy statement holds true. You could have have 100% precision and low recall and still have 85% accuracy (meaning you could never flag someone as lying and be wrong while missing a bunch of liars and still have 85% accuracy).
but if everything is totally evenly distributed, then 85% accuracy means 85% accuracy and your first statement is correct.
The real issue is that accuracy is only one piece of the puzzle.
This is going to be awful - imagine all the people who are anxious anyway, perhaps dont speak the langauge very well, get confused by the questions etc. There are going to be a lot of people who will get "enhanced screening" and generally treated like a criminal for no other reason than "the computer said you are a liar".
Awful.
So it maybe that the bigger problem is the false negatives.
Even if they bring it up to 85% accuracy with 50% base rate, by the time you are dealing with base rates that are more realistic, you're going to run into way more problems than just 1 out of every 7 people.