Skip to content

Author

Mihir Dhanakshirur

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

Entropy Regularization: A Free Correction to Cross-Entropy for Verified Demonstrations

Large language models are often post-trained on expert demonstrations using cross-entropy (CE), even when the downstream objective is not to imitate the demonstrated solution but to produce any output accepted by a verifier. This mismatch is seen in verifiable domains with multiple correct solutions, such as mathematic...

Mihir Dhanakshirur, Adam Ousherovitch, A. Tewari · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.