1 - (1-3/200)^19800
Which of course is basically 1. The lower bound is 0, the upper bound is almost 1, so the confidence interval is [0, 1). Basically, no information at all. So its much to early to conclude that you're not going to find any typos in the entire book.The best way to state your conclusions is to say "after sampling 200 pages at random, we are 95% confident that the true typo-per-page rate is in the interval 0% - 1.5% or equivalently the total number of typos in the 20,000 page book is between 0 and 300."
"... or equivalently the total number of [pages with] typos in the 20,000 page book is between 0 and 300."
We're not measuring the number of typos on a page, but only whether the page contains a typo or not. So we can't speak about the number of typos.
Just as the article itself states regarding the approximation "Since log(1-p) is approximately –p for small values of p" which relies on discarding p-squared terms as insignificantly small.
And I won't even bother with the argument that presence of one typo on a page might actually increase the probability of another typo on the same page or nearby pages. The idea that this principle is based on uniform distribution is discussed enough in other threads.
So after reading 20 pages with no errors and having no other information the odds of finding an error on the next page are no more than 1/60. If you have 19,800 more pages left you'd expect to find no more than about 330 errors.
Of course, the comments here suggest that it just makes the whole concept confusing to an audience who should either already know about confidence intervals or be well prepared to learn about them.
If you’ve sampled all the pages, then we are talking about certainty which this wouldn’t apply.
So in your example, according to the rule, the probability of errors is less than 3/200. Which it is, either 0/201 or 1/201.