So, your 99.8% confidence isn't 99.8%.
There are a few ways to compensate for this. The easiest is to fix your sample size and not reach any conclusions until you have tested that many users.
You can determine your sample size by, say, figuring out the minimum number of people needed before there's a 95% chance you observe a 10% effect size.
This is a problem in clinical trials where ethical questions arise. If the test appears harmful, should we stop? If it appears beneficial, isn't it unethical to deprive the control group of the treatment?
Anyhow, if you want to observe the results continuously, you need to use a technique like alpha spending, sequential experimental design, or Bayesean experimental design.
TL;DR: If you're periodically looking at your A/B testing results and deciding whether to continue or stop, you're doing it wrong and your significance level is lower than you think it is.