We present our paper "Mitigating Reward Hacking via Adversarial Robustness" at EIML
@ICML2026!
We conjecture that reward hacking is often caused by flipped advantage-sign estimations, and propose SignCert-PO, a new algorithm built on the theory of randomized smoothing! 🧵