FAVE: Foveated Adaptive Visual Encoding for Efficient Fine-Grained Visual Understanding
FAVE (Foveated Adaptive Visual Encoding), a lightweight variable-resolution ViT that encodes externally selected regions at high acuity while preserving native geometry, is introduced and integrated as a complementary local branch in FastVLM.
Amitangshu Mukherjee, Kaushik Roy
· 0 citations