Skip to content

Author

Nico Daheim

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Parameter Exploration for RLVR via Variational Learning

Evidence that parameter-space exploration can improve reinforcement learning for LLMs is presented, and a family of methods called Perturbed Parameter Policy Optimization (3PO) is introduced which use different sampling strategies and different rollout grouping for reward estimation.

Vatsal Venkatkrishna, Nico Daheim, Iryna Gurevych · 0 citations