GeoCo-SAVi: Geometry-Consistent Slot Attention for Explicitly Editable Object Representations
Abstract
Object-centric video models represent scenes with slots, yet exposed geometry can vary in meaning with appearance. In Invariant Slot Attention (ISA), explicit position and scale can disagree with the decoded center and extent; edits can yield unexpected motion or resizing, and replacing appearance can shift geometry. GeoCo-SAVi promotes geometric authority and semantic alignment. Its spatially equivariant, object-wise decoder makes position and scale effective commands: changing them moves or resizes the rendered support. Factual position alignment ties position to the decoded center, and normalized attention overlap discourages duplicate allocation. Appearance transplantation aligns geometry semantics across objects, so recipient geometry governs layout while donor appearance supplies shape. On Obj3D, GeoCo-SAVi matches ISA reconstruction, reduces latent-position-to-decoded-centroid error by nearly 90%, and reduces appearance-induced size variation while producing the expected translation and scale responses. On 250 MOVi-C videos, it improves reconstruction and instance grouping over both same-protocol references, and video editing over STAITUS. GeoCo-SAVi transforms explicit geometry into compositional control, making both position and scale more readable and editable. Project page: https://GeoCo-SAVi.github.io/.