Evaluating on the training distribution is not evaluation
Policies trained on a fixed asset set evaluated on the same fixed asset set overstate real-world performance. Honest evaluation needs a held-out catalog the policy never saw at training, which is exactly what most teams cannot afford to build by hand.