Jul 2026· International Conference on Control, Decision and Information Technologies· pp. 1299-1304· 0 citations· 16 references
Abstract
Person search in real-world surveillance and robotic perception often fails when facial cues are unreliable due to low resolution, motion blur, or occlusion. We propose a clothing-centric person search framework that represents people using appearance attributes (e.g., garment type, dominant colors, accessories) instead of biometrics. Given person tracklets, the system samples frames, extracts clothing cues, and generates schema-constrained structured descriptions using a large language model, enabling consistent semantic indexing. Retrieval is performed with attribute-aware semantic search over these descriptions and compared against appearance-based baselines in controlled multi-camera experiments. A full-dataset analysis of the generated descriptions highlights common failure modes under real-world conditions, including color ambiguity under lighting changes (e.g., black vs dark), posture hallucinations under occlusion (e.g., standing behind a table described as sitting), and occasional prompt deviations that introduce spurious attributes. Runtime results show that detection is lightweight (about 30 ms per frame), while semantic extraction dominates; with parallel workers, the full pipeline processes roughly 0.25-0.34 s of computation per second of video, supporting practical deployment.
A mask-aware tri-modal framework that improves the quality of superpoint representations by retrieving a scene-level structural context from a pretrained PointSAM encoder to enhance object-centric evidence and predicting a soft mask weight to suppress unreliable superpoints.
Feng Zhou, Hui Wang, Kaida Ning et al.· The Visual Computer· 0 citations
A Multi-level Semantic-Guided (MSG) framework that integrates contextual and fine-grained visual information to eliminate clothing variance across both conceptual and pixel dimensions is proposed.
Shijuan Huang, Hefei Ling, Zongyi Li et al.· ACM Transactions on Multimed...· 0 citations
Person search is challenging due to limitations in identity representation. Existing methods rely on one-hot encoding, ignoring semantic relationships among pedestrians. This leads to a fragmented feature space and reduces generalization ability, especially in large-scale scenarios with a significant proportion of unla...
Xi Yang, He-Xun Zhou, Hai-Yang Zhu et al.· Neural Networks· 0 citations
In real-world surveillance, situations arise in which two people wear nearly identical clothing or in which identifying features are obscured by heavy occlusions and shifting poses. In these environments, traditional uni-modal systems that rely on static appearance do not perform well and often produce false matche...