The results distinguish cross-sensor consistency, multimodal fusion, and generalization to new action orders as separate questions as separate questions in evaluating multimodal physical representations.
Kai-Zhen Tan, Xin Xu, Si-Ru Tao et al.· 0 citations
Metric questions about video require vision-language models to use supplied real-world references to convert visual measurements into physical units. Yet we find that current models use this scale information only partially. When every world-space quantity in a prompt is rescaled by a common factor, the video remains e...
Kai-Zhen Tan, Yang Feng, Heqing Du et al.· 0 citations
Extending compression-based memorization analysis to the frozen-base setting, this work measures directly, in bits, how much a low-rank adapter writes into a model it never changes, finding that the answer is both smaller than full fine-tuning and less lawful than parameter counting would predict.
Kaizhen Tan, Heqing Du, Yang Feng· arXiv.org· 0 citations
A central premise of latent world models is that predicting the future forces a representation to internalize the physics of its environment. Which physical quantities does a trained latent actually contain, and what decides this? We answer with controlled interventions in POKEWORLD, an interactive environment whose vi...
Kai-Zhen Tan, Xin Xu, Siru Tao et al.· arXiv.org· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.