Strong agree that static evals saturate — the decay property of markets is the genuinely useful part: historically, quant alpha decays on the order of 30-50% per year as capital crowds in, so a live market eval is self-difficultating, exactly what model comparison needs once benchmarks plateau. The hard part I'd flag is comparability: market paths are stochastic, so two runs of the same model can land on wildly different difficulty depending on the realized path — without controlling for that, the eval measures luck more than the model. Options that work: fixed-seed regime paths with resampled baselines, or bootstrapped difficulty metrics (percentile of PnL against a distribution of random strategies) instead of raw return. The second hard part is reward shaping for RL: sparse PnL rewards over multi-day horizons give terrible exploration, so you'll likely need shaped intermediate rewards (execution quality, information state) to get gradients flowing at all. Also worth publishing: variance across seeds for each model — that's the number that tells users whether a 2% delta is signal or noise. Are you planning to ship fixed seed regimes, or is path stochasticity part of the point?