The results distinguish cross-sensor consistency, multimodal fusion, and generalization to new action orders as separate questions as separate questions in evaluating multimodal physical representations.
Kai-Zhen Tan, Xin Xu, Si-Ru Tao et al.· 0 citations
Pretrained vision embeddings are increasingly used as general-purpose representations for modelling how people appraise urban scenes, and are validated almost entirely by how well they predict human ratings. High predictive accuracy does not establish that these embeddings organise scenes as human perception does. We t...
A disjunctive policy allows a value to depend on at most one of two secrets and never on both: an analyst may consult one client's file or the other's, a share of a split secret may be released but not its sibling. Such policies are not lattice-shaped, and Hunt and Sands introduced the quantale of information to give t...
Metric questions about video require vision-language models to use supplied real-world references to convert visual measurements into physical units. Yet we find that current models use this scale information only partially. When every world-space quantity in a prompt is rescaled by a common factor, the video remains e...
Kai-Zhen Tan, Yang Feng, Heqing Du et al.· 0 citations
Vision-language models are increasingly used to measure urban change from repeated street-level imagery, but their longitudinal reliability is not well understood. We test how much a perception score can change when the street itself does not undergo substantial redevelopment. Using 4,648 consecutive-epoch image pairs...
Extending compression-based memorization analysis to the frozen-base setting, this work measures directly, in bits, how much a low-rank adapter writes into a model it never changes, finding that the answer is both smaller than full fine-tuning and less lawful than parameter counting would predict.
Kaizhen Tan, Heqing Du, Yang Feng· arXiv.org· 0 citations
A central premise of latent world models is that predicting the future forces a representation to internalize the physics of its environment. Which physical quantities does a trained latent actually contain, and what decides this? We answer with controlled interventions in POKEWORLD, an interactive environment whose vi...
Kai-Zhen Tan, Xin Xu, Siru Tao et al.· arXiv.org· 2 citations
Empirical work on algorithmic collusion asks one question of the data: are prices supracompetitive? We show this can be answered"no"by a conspiracy that is nonetheless profitable. Consider bidding agents that couple only through the joint distribution of their unexplained bid components, leaving every agent's own bid l...
Xin Xu, Cheng-Rui Wu, Jiayu Lu et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.