> not to mention the egregious target leakage
I was curious about this so I skimmed the paper [0]:
> SalesRLAgent achieved 96.7% accuracy, outperforming the best commercial alternative by 23.7 percentage points and the best LLM approach by 34.7 percentage points.
For a fuzzy natural language task like this, this magnitude of improvement should already set off alarm bells (Though i admit I'm not even sure what accuracy is even measured here, and the paper doesn't help either). Also, "best LLM" here refers to GPT-4 (at the time of upload, the public already had access to GPT-o3 and).
I would have loved to contextualize the performance by looking at model size, but the paper is frustratingly devoid of detail in that regard:
> The core of SalesRLAgent is a reinforcement learning
architecture consisting of:
• A state encoder network that processes Azure OpenAI
embeddings and features
• A policy network that estimates conversion probability
based on the current state
• A value network that estimates the expected cumulative
reward
• A meta-learning module that assesses prediction confi
dence
Also:
> Beyond technical metrics, we evaluated SalesRLAgent in
real-world sales environments through A/B testing. [...] After 90 days across 217 representatives and 12,433 con
versations, we observed:
• 43.2% increase in conversion rate for the test group
This would be a pretty huge result but the fact that this is just shoved into a single paragrpah with no further discussion on methodology, baselines and setup makes me very suspicious.
[0] https://arxiv.org/abs/2503.23303