What about quality of results? Are you measuring that too? Did you do so for the traditional reference practice? Using what sort of methodology? How did it your technique compare in quality? What kind of errors was it most likely to make? What techniques have you devised for spotting those errors? Are they the same kind of errors that users would experience when outsourcing? Are the errors easier or harder to spot for one than the other? Are they faster to remediate with one?
I see a clever concept but given the state of LLM's and the nature of how they work, I don't know that nominal cost and speed differences are really enough to sell on. Not for something "crucial to big business decisions." I'd want to know that my failure/miss rate is no worse than when outsourcing and that my net cost and time (including error identification and recovery) still end up ahead. I don't see either of those vital issues touched upon here.