Skip to content
Open access

LIP Reading Using Neural Network and Deep Learning

Jul 2026 · International Journal for Research in Applied Science and Engineering Technology · Vol 14, pp. 1873-1876 · 0 citations

TL;DR

An automated lip reading system built using Convolutional Neural Networks to classify spoken words from video sequences of a speaker's mouth region and is able to reliably identify isolated words such as “Hello”, “Start”, “Stop”, and “Previous” in real time.

Abstract

Lip reading, the task of inferring spoken words purely from visual observation of lip movements, remains a challenging problem even for trained human lip readers, who are typically able to correctly identify only about every second word. This paper presents an automated lip reading system built using Convolutional Neural Networks (CNNs) to classify spoken words from video sequences of a speaker's mouth region. The system uses a Haar Cascade classifier to localize the face and mouth region in each video frame, followed by a dlib-based facial landmark detector to extract a precise lip Region of Interest (ROI). The cropped and normalized lip images are then passed to a trained CNN, which was benchmarked against a Long Short-Term Memory (LSTM) network and a Temporal Convolutional Network (TCN) to model the sequential nature of lip movement, with the CNN architecture selected for deployment based on superior classification accuracy on the Lip Reading in the Wild (LRW) dataset. The trained model was integrated into a real-time application capable of capturing live webcam video, isolating the speaker's lip region frame-by-frame, and predicting the spoken word with an associated confidence score. Experimental evaluation on a ten-class subset of the LRW dataset shows that the proposed CNN model is able to reliably identify isolated words such as “Hello”, “Start”, “Stop”, and “Previous” in real time. The system demonstrates the practical feasibility of vision-only speech recognition and its potential applications in assistive hearing technology, security and surveillance, and human–computer interaction in audio-degraded environments

Read PDF

Similar papers

Open access 2019

Speech Emotion Recognition using Convolutional Neural Networks and Recurrent Neural Networks with Attention Model

Speech emotion recognition is an upcoming subfield of automatic speech recognition that shares multiple similarities with mood recognition in music signals. Audio signals containing human speech are used as input to classification algorithms trained to recognize emotions in the form of audio features. This thesis outline...

G. Tomas, S. Weinzierl, Athanasios Lykartsis · 2 citations
Open access 2026

MultiModal Deep Learning Framework for Missing Children Detection using Vision-Language Mode

As report of missing children continue to rise around the world there is an urgent requirement of some smart and scalable systems which can help in timely and accurate identification. In this paper, we propose a multimodal deep learning system that combines Vision-Language Transformers like CLIP, Face Re-identification...

K. U. Maheswari, P. Dhanalakshmi · 0 citations
Open access Sep 2026

A Hybrid Deep Learning Architecture for Image Classification Across Diverse Visual Recognition Applications

The core computer vision problem of picture categorization has several domain-specific applications. This paper's goal is to talk about the significance of picture categorization in modern technology and society, as well as its ideas, techniques, and applications. Many computer vision systems rely on image classificati...

C. Patel · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.