Implied Expressive Motions in Paintings: An Early Study on the Limits of Text-to-Motion Generation
Abstract
Paintings encode what cognitive science calls implied motion: the movement that the artist freezes at a chosen instant and that the viewer’s perceptual system completes. This motion also carries expressive and affective qualities, a property we term Implied Expressive Motion (IEM). Generating this motion computationally could open new pathways for art understanding and cultural welfare, but it is not yet clear whether current generative models can handle the expressive demands of artistic content. In this paper, we present an envisioned pipeline for extracting IEM from paintings and conduct a feasibility study of its text-to-motion stage in isolation. Four recent motion in-betweening models are compared using target poses and hand-authored prompts derived from three figures in two paintings by Luca Cambiaso, later evaluated by a panel of four experts. The results show a 67% hallucination rate, a near-total failure on the most culturally loaded gesture, and, even in the best-performing model, smooth but expressively neutral executions. These findings characterize where current models break down: expressive control, artistic vocabulary, and cultural grounding; and define open challenges for culturally-aware multimodal interaction.