Jun 2026
MultModLM: A multi-modal benchmark for Large-Language Model based hardware schematic generation
This work introduces MultModLM, a benchmark for evaluating LLMs on the task of generating hardware schematics from RTL (Register Transfer Level) descriptions, and finds that LLM-based evaluators exhibit near-zero agreement with human raters, revealing that LLM-as-a-judge paradigms are unreliable in structurally precise domains.
Dhruva Kulkarni, Sai Manoj Pudukotai Dinkarrao
· arXiv.org · 0 citations