MUMBAI, India, July 30 -- Intellectual Property India has published a patent application (202641088171 A) filed by Karpaga Vinayaga College Of Engineering And Technology; Dr. A. B. Hajira Be; and Ms. Mahalakshmi S on July 20, 2026, for A Multimodal Ai Music Platform Using Face, Gesture, Text, And Voice Analytics Thereof (emoges).
Inventors include Dr. A. B. Hajira Be; and Ms. Mahalakshmi S.
The application for the patent was published on July 24, 2026, under issue no. 30/2026.
Abstract: Music plays a significant role in influencing human emotions and mental well-being. Traditional music recommendation systems primarily rely on user listening history, ratings, and preferences, which may not accurately reflect a user's current emotional state. To address this limitation, the proposed system, EmoGes: Emotion and Gesture-Based Music Recommendation System, introduces an intelligent and interactive approach to music recommendation by analyzing real-time human emotions and gestures. The system utilizes multiple input modalities, including facial expression recognition, hand gesture detection, text sentiment analysis, and voice- based emotion identification, to determine the user's emotional condition. Facial emotions are detected using a Convolutional Neural Network (CNN) trained on facial expression datasets, while hand gestures are recognized using the MediaPipe Hand Landmarker framework. Textual inputs and voice recordings are processed through keyword-based sentiment analysis and speech-to-text conversion techniques. The identified emotions are classified into three primary emotional states: Happy, Neutral, and Sad. Based on the detected emotional state, the recommendation engine filters and retrieves suitable songs from predefined music datasets using valence and tempo attributes. A unique Mood Booster feature is incorporated to assist users experiencing negative emotions by suggesting uplifting or relaxing music according to their preferences. The system supports both English and Tamil music recommendations, thereby enhancing accessibility and user engagement. The application is developed using Streamlit as the frontend framework, TensorFlow and Keras for deep learning-based emotion recognition, OpenCV for image processing, MediaPipe for gesture tracking, and SpeechRecognition for voice analysis. The modular architecture ensures scalability, maintainability, and efficient real-time performance. The proposed system offers a personalized music listening experience by understanding human emotions through multiple channels rather than relying solely on historical data. This approach improves recommendation accuracy, enhances user satisfaction, and demonstrates the potential of artificial intelligence in creating emotionally aware multimedia applications for entertainment and mental wellness.
Disclaimer: Curated by HT Syndication.