Visually impaired people (VIP) need assistive technology for smooth and confident navigation on pavements along the roads. VIPs find it hard to follow the speech generated through a text-to-audio converter due to the presence of advertisement text. Automating the process of audio conversion from the signboards by integrating the capabilities of emerging AI models in a fraction of a second will be useful for their confident navigation. In this context, SignSpeak, an AI-based assistive utility, is designed to convert signboard text to speech using visual inputs, which may be captured through smart sensors or canes. The proposed system follows a three-step (detection-extraction-conversion) pipeline. Initially, a fine-tuned lightweight YOLOv11s model is used to generate a bounded box for the relevant text region, which is input to Gemini model, a multimodal large language model for text extraction. Lastly, the extracted text is converted into audio using the Google Text-to-Speech tool. We fine-tuned the YOLOv11s model on 200 signboard images and tested its performance on an additional curated dataset of 42 real-world signboard images. Experimental findings indicate that the proposed integrated pipeline achieved superior performance compared to the solitary usage of the Gemini model for text extraction followed by its audio conversion.
Sharanjit Kaur, Manju Bhardwaj, Kriti Misra et al.· ITEGAM- Journal of Engineeri...· 0 citations
Steganography attempts to conceal messages in plain sight while steganalysis seeks to identify them or, more importantly, to extract the embedded data. Low-payload and spatially localized steganographic embedding is increasingly used to evade detection by classical steganalysis methods. While such strategies preserve global image statistics and remain visually imperceptible, they can disrupt natural pixel-level behavior. This work proposes a behavioral steganalysis framework inspired by process mining that detects image steganography by analyzing localized behavioral deviation using regional behavioral contrast and behavioral amplification. Experiments on lossless grayscale PNG images from the USC SIPI database and 10,000 images from the BOWS2 dataset using 1-bit LSB embedding show that the proposed framework reliably identifies steganographic embedding. On the USC SIPI dataset, conventional statistical detectors, including chi-square analysis and the StegExpose tool, showed limited detection capability under the evaluated localized embedding settings. Despite high perceptual quality of stego images (PSNR > 55 dB), significant behavioral deviation is consistently observed within embedded regions. These results demonstrate that the proposed process mining-inspired framework provides an interpretable and complementary direction for image steganalysis, particularly under low-payload and localized embedding scenarios.
Shikha Badhani, Vinita Verma, Manju Bhardwaj et al.· International Journal of Mat...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.