Skip to content
Open access

Vision-Enabled Virtual Robotic Head with Multi-Modal Communication Capability.

2026 · International Journal of Mathematical and Computer Sciences · 0 citations · 10 references

Abstract

AI-enabled intelligent virtual humanoids are critical to ensure natural and socially aware human-machine interactions in areas such as education, customer services, and companionship. In this paper, we develop a vision-enabled virtual robotic head which is entirely running inside a web browser and includes modules like Speech-to-Text, Large Language Models, Text-to-Speech, and vision processing using FastAPI and WebSocket architecture. The proposed approach generates a dynamic robot-like face capable of realistic talking and facial expressions along with a vision module for face detection and visitor recognition and personalized interaction based on their memory. The proposed architecture ensures low-latency (less than one second) despite the active use of vision processing capabilities. User evaluation on 50 people resulted in 92% face recognition accuracy and high satisfaction scores.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.