Jul 2026· International Conference on Future Internet of Things and Cloud· pp. 441-448· 0 citations· 11 references
Abstract
This paper proposes a well-rounded media consumption system that utilizes speech recognition and generative Artificial Intelligence to create a user-influenced slideshow. The system runs on a Raspberry Pi, using a microphone for input and a touchscreen for output. Spoken user requests are transcribed and categorized into one of five content classes: news, daily activities, social media trends, interest-based topics, and storytelling. These categories are then converted into prompts suitable for image generation. The system uses AI to condense unstructured speech into descriptive, image-ready content, reducing cognitive overhead and minimizing screen time. Unlike conventional browsing, this approach enables passive, voice-controlled consumption of highly relevant media. Qualitative and quantitative evaluation shows that the system reliably transcribes varied speech inputs, classifies user intent with high interpretability, and generates coherent, category-aligned visuals. Specifically, the LLM-based intent classifier achieved 90% accuracy and a Macro-F1 of 0.931 on the tested prompts, while on-device transcription operated at an average CPU utilization of 7.3% with a peak SoC temperature of 48.7°C, confirming the feasibility of the pipeline on a Raspberry Pi. The use of AI in this context enhances personalization, reduces interaction friction, and supports timeefficient engagement with digital media.
An automated pipeline for transcription and summarization of video conferences held on the Jitsi Meet platform that requires no proprietary cloud API keys and is deployable on-premise via Docker Compose, making it suitable for organizations with strict data-privacy requirements.
G. Amirkhanova, L. Bektemir, Shyrailym Adilkyzy et al.· International Conference on...· 0 citations
Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by multimodal assistants,avatars, and embodied agents. While recent audio-video generative models can synthesizehigh-fidelity synchronized content, existing supervision is...
Chi Zhang, Hao-Yan Shi, Yueyi Liu et al.· 2 citations
People with physical and motor disabilities face significant barriers when interacting with image creation applications that rely on keyboard input, touchscreens, or complex graphical interfaces. Although modern AI image generation models produce high-fidelity images from textual descriptions, they remain dependent on...
Dasari Sri Krishna, Vaddi Radhesyam, Lakshmi Aiswarya Narikimilli et al.· 2026 4th International Confe...· 0 citations
Taking notes during meetings sounds easy, but in reality, people often miss key points—especially when multiple people are talking. Managing follow-ups in separate apps only adds to the hassle. MinuteMaster was built to simplify this. It's a web app that brings transcription, speaker identification, summarization, and...
B. Banik, Pampari Yukitha, Charan Sai Venna et al.· Next-Generation Computing Sy...· 0 citations
A user study comparing two modalities for writing prompts for generative AI tasks reveals that input modality significantly influenced prompting behaviour but did not lead to measurable differences in subjective evaluations.
Nishant Rathore, Tushar Billakanti, J. Ceha et al.· International Conference on...· 0 citations
The generative artificial intelligence has made tremendous advancement in visual content generation; but the issue of intuitive and user-friendly adjustment of the current video scenes is a difficult question to answer. The paper describes a personalized movie reimagining system based on AI that reimagines input video...
Thupakula Venkata Sumanth, Kriti Gupta, Charanjit Singh et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.