Aren't you worried about the agent missing or hallucinating policy details?
Aren't you worried about the agent missing or hallucinating policy details?
> It is more accurate and consistent than our humans.
So, errors can clearly happen, but they happen less often than they used to.
> It will draft a reply or an email
"draft" clearly implies a human will will double-check.
If you take the comment at face value. I'm sorry but I've been around this industry long enough to be sceptical of self serving statements like these.
>"draft" clearly implies a human will will double-check.
I'm even more sceptical of that working in practice.
The wording does imply this, but since the whole point was to free the human from reading all the details and relevant context about the case, how would this double-checking actually happen in reality?
On the off chance it’s not for that reason, productivity requirements will be increased until you must half-ass it.
That's your assumption.
My read of that comment is that it's much easier to verify and approve (or modify) the message than it is to write it from scratch. The second sentence does confirm a person then modifies it in half the cases, so there is some manual work remaining.
It doesn't need to be all or nothing.
When the AI gets "good enough", and the review becomes largely rubber stamping, and 50% is pretty close to that, then you run the risk that a good percentage of the reviews are approved without real checks.
This is why nuclear operators and security scanning operators have regular "awareness checks". Is something like this also being done, and if so what is the failure rate of these checks?