However, teachers naturally want to motivate their best performers (what teacher would grade so that the best student in their class would get a B- ?) and school administrators have a practical incentive to inflate the grades. This means that subjective evaluation of speeches and projects (especially if you're comparing them within a single institution, not to what projects and speeches are considered good and bad in other institutions) and continuous real-time questioning won't lead to an evaluation that's useful to determine where a student stands in comparison to other students in other institutions, and furthermore, they are going to be optimistic. A relevant headline is "40% of A-level results were downgraded from teachers' predictions." - however, that seems normal and expected to me; do we have the equivalent numbers from last year, comparing the teacher's prodictions with the actual exam results?
Technically you might consider taking the "local evaluation" and make an adjustment for the institution, but this is exactly what they attempted to do this year in UK, and the whole original article is about the limitations of this approach.
Also, fraud and cheating is a problem. There are many approaches that have proven to result in somewhat effective remote teaching and learning, however, we don't have good solutions for effective remote evaluation (this is a problem that I personally experienced this spring in my work). For the pre-COVID remote education, the only thing that IMHO worked was a network of in-person proctoring centers supervising the testing process and identity of the people who are taking the test (otherwise people will get others to take the test for them, it's not a hypothetical issue) or, in certain cases, a privacy invasive and labor intensive online proctoring.
IMHO it makes all sense to separate instruction and teaching from final evaluation and certification, due to the unavoidable conflicts of interest and misleading results that happen otherwise. This is also what UK has chosen as its policy - but as they found out, this means that they can't have a meaningful evaluation if they skip the actual evaluation part this year.