This is the sharpest objection in the thread, and I think it dissolves if you redefine "verification" as provenance rather than truth-arbitration. The problem with "the model checks its own facts" is real. But the useful version doesn't ask the model to be a judge — it asks the system to record, per claim: which tool call produced it, what query was run, what source was cited, when. Then a human (or a second, dumber check) can re-run the chain. You don't need the AI to know what's true; you need every claim to be re-derivable on demand.
In practice this is what separates automations I trust from ones I don't. The ones I trust emit an audit trail as a side effect — every output links back to its inputs. The ones I don't just hand me polished text. The checkbox UI the author proposes is one surface for this; the deeper point is that verification-by-construction scales where human re-checking doesn't.