Not Another Text Benchmark: Putting the"Visual"Back in Visual Question Answering for Large Video Models
Large video models have exhibited impressive performance on a wide range of visual question answering tasks, owing to the rise of powerful, pretrained text and vision encoders. The usefulness of such models have also been demonstrated on a wide range of benchmarks, with an important caveat - the dominant approach in th...