Cross-Entropy Games and Frost Training
Preprint, 2026Arthur Renard, Franck Gabriel, Valentin Hartmann, Clément Hongler
Xent Labs · Université Lyon 1
Frost Training improves Monte Carlo policy optimization for Cross-Entropy Games by exploiting the gradient of the reward in embedding space. The signal behind the GCG jailbreaking technique is repurposed to boost training: applied to GRPO for maximum-likelihood infilling, it reaches higher best-of-k scores faster.




