Using Vision Language Foundation Models to Generate Plant Simulation Configurations via In-Context Learning
A benchmark for evaluating whether vision-language models (VLMs) can generate plant simulation configurations from imagery using in-context learning and establishing a benchmark for studying how multimodal reasoning, prompt design, and the sim-to-real domain gap affect plant phenotyping tasks is established.