Skip to content

Author

Sumit Gulwani

4 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills

Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic tasks, but pay a 3-6x premium in output tokens on every episode -- much of it spent re-deriving procedures that are shared across episodes of the same domain. We show this recurring cost can be amortized: a coding agent analyses a small corpus of existing trajectories from a training split and compiles a compact natural-language skill that is injected into the non-reasoning model's system prompt. Across four agentic benchmarks (ALFWorld, tau$^2$-bench telecom and retail, and SpreadsheetBench-Verified), skills recover 55%-100%+ of the reasoning gap for GPT-5.4-mini on held-out tasks -- exceeding the reasoning mode outright on two of four -- while emitting 2.7-6x fewer output tokens and zero reasoning tokens. Notably, reasoning traces are not a prerequisite: skills distilled from non-reasoning trajectories alone remain competitive with skills distilled from paired reasoning/non-reasoning corpora, with domain-dependent differences between the two sources. We interpret these results through a search lens: test-time reasoning is deep search inside a single episode, re-paid at every deployment, while corpus distillation is wide search across episodes, paid once. The two recover overlapping procedural knowledge, and width over cheap trajectories is often the better buy -- with the residual gap on some domains (telecom, SpreadsheetBench) delineating where genuinely per-instance deep search remains necessary.

Agamdeep Singh, Srishti Gautam, Priyanshu Gupta et al. · 0 citations
Open access 2026

Asking language models how to represent data for fine-tuning

It is shown that format choice remains important even after fine-tuning; models learn more efficiently with specific formats rather than adapting to any format; this finding allows format selection to be done via inference alone, avoiding costly trial-and-error fine-tuning runs.

Usneek Singh, Ananya Singha, Abhijeet Awasthi et al. · 0 citations
Preprint Jul 2026

Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents

A prototype of a Plan Mode for spreadsheet programming is built and evaluated against a non-planning baseline and it is found that using Plan Mode led to a reduction in refinement and a better perception of the tool across dimensions of creativity support and human-machine collaboration.

Aayush Kumar, Avik Dutta, Sumit Gulwani et al. · 0 citations

Towards Autonomous Software Development

A three-level taxonomy inspired by autonomous driving that distinguishes degrees of autonomy along a roadmap from today’s AI-assisted development workflows to fully autonomous software development in which AI systems autonomously identify demands and design, implement, verify, and maintain software without human oversight is introduced.

Hao Wang, Ruijie Meng, Zhe Ye et al. · 0 citations