RAG-enhanced MLLMs are a viable, low-barrier solution for inspection tasks in remanufacturing scenarios with limited expert supervision and diverse object classes and demonstrate their viability in low-data, multi-class remanufacturing scenarios.
A novel paradigm called Human vs. LLM Identification (HLI) is proposed which introduces a Retrieval-Augmented Generation (RAG)-inspired evidence-based detection strategy alongside a fine-tuned transformer classifier.
Ibtasam Ur Rehman, Muhammad Islam, Muhammad Yousaf Rehman et al.· Knowledge· 0 citations
This research provides a highly accurate, scalable, and reliable framework for automated bridge defect analysis, offering a practical methodology to enhance data utilization in bridge management.
Lu-yang Zhang, Xuzhao Lu, Fengzong Gong et al.· Advances in Structural Engin...· 0 citations
This paper presents a multimodal pipeline that combines vision-language captioning models and openvocabulary object detectors to investigate the impact of automatically generated textual prompts on semantic image understanding and reveals complementary behaviors between Grounding DINO and OWLv2.
Xin Gao, Madjid Maidi, B. Daachi· NLP & Big Data· 0 citations
AutoDavis is introduced, a first-of-its-kind automatic and dynamic evaluation protocol that enables on-demand benchmarking of LVLMs across specific capability dimensions and shows effectiveness and reliability, offering a new paradigm for dynamic benchmarking of multimodal intelligence.
Han Bao, Yue Huang, Yan-Bo Wang et al.· Proceedings of the 32nd ACM...· 0 citations
Industrial anomaly detection is a critical component of modern manufacturing. Most traditional unsupervised methods rely on modelling normal feature distributions, inherently limiting generalization to unknown categories. To improve generalizability, some recent methods incorporate vision-language models (VLMs) for zer...
Weifeng Chen, Hong-Hao Zhang, Zhiyuan You et al.· 1 citation
It is observed that attention scores from both vision and text tokens peak at modality separator tokens, suggesting that these separators bridge the two modalities and proposes SepPrune, an efficient, training-free, plug-and-play pruning method that uses the separator token as a unified query to rank and select informa...
Yucheng Wang, Qihui Zhu, Yang Liu et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.