Industrial visual sensing systems play a critical role in automated quality inspection. However, deploying data-driven perception models on industrial visual sensors faces significant bottlenecks due to the extreme scarcity of annotated anomaly samples and the pronounced long-tailed distribution of defect types. These factors render traditional supervised models prone to overfitting and catastrophic failure when sensing rare anomalies. To address these limitations in the sensor data processing pipeline, this article proposes a novel few-shot perception framework leveraging multimodal large language models (MLLMs) as a cognitive backend, without task-specific fine-tuning of the sensing model. Our approach introduces two key innovations: 1) a vision-text paired in-context learning (VTP-ICL) mechanism that constructs structured multimodal prompts to activate pretrained knowledge for precise defect localization and 2) a two-stage reflection pipeline that decouples initial prediction from verification, enabling the model to self-correct outputs based on explicit criteria. Extensive experiments on a real-world industrial fabric dataset demonstrate that our method achieves state-of-the-art performance under strict few-shot conditions (two shots per class). Notably, our approach maintains an exceptional balance between precision and recall, significantly outperforming traditional detectors, particularly on hard-to-sense, low-contrast defects. These results validate the potential of MLLMs as robust and generalizable engines for next-generation industrial sensing applications.
Scientific reasoning requires language models to retrieve specialized knowledge and incorporate it reliably into multi-step computation. Conditional memory provides an explicit lookup pathway that complements dense neural representations, but its usefulness is inherently input- and computation-dependent: retrieved information may repair missing scientific associations, yet it may also introduce distracting shortcuts or interfere with reasoning that the base model can already perform correctly. In this work, we systematically investigate when, where, and to what extent conditional memory should participate in scientific reasoning. We characterize the scientific knowledge boundary and controlled interventions on memory-enabled knowledge-circuit nodes. Based on these analyses, we propose a Knowledge Boundary-Aware Router that uses task-specific input proxies available before generation to determine whether memory is activated, which layer-stage nodes receive memory signals, and how strongly these signals contribute. Experiments on biological and chemical reasoning benchmarks, covering two backbone families and six task types, show that memory effects vary substantially across inputs, tasks, and injection locations. Compared with static and activation-rate-matched random routing, our approach more consistently preserves beneficial memory contributions while suppressing memory-induced regressions, establishing selective memory allocation as an important principle for reliable scientific reasoning.
Zhen Bi, Xueshu Chen, Yan Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.