OdysSim: Building Foundation Models for Human Behavior Simulation
It is shown that LLM-as-judge RL induces reward-hacking patterns, and that LLM-as-judge RL detectors can mitigate them during post-training, suggesting that behavioral foundation models require rethinking the LLM training paradigm.