Skip to content

Author

Haojia Sun

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

OdysSim: Building Foundation Models for Human Behavior Simulation

It is shown that LLM-as-judge RL induces reward-hacking patterns, and that LLM-as-judge RL detectors can mitigate them during post-training, suggesting that behavioral foundation models require rethinking the LLM training paradigm.

Xuhui Zhou, Weiwei Sun, Weihua Du et al. · 8 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.