MUMBAI, India, Aug. 12 -- Intellectual Property India has published a patent application (202611078898 A) filed by Dr. Bajarang Prasad Mishra; Dr. Aakunuri Manjula; Barnali Barman; Smt. Pemma Radhika; Dr. Murthy Ravaleedhar Reddy; Srisailanath; Sailen Dutta Kalita; Dr. Saurabh Kumar; Dr. Arun Kumar Sao; Chitri Rami Naidu; Dr. S. Muthuselvan; and Lakshmi S on June 26, 2026, for A Multimodal Generative Artificial Intelligence System For Unified Text, Image, Audio, And Video Content Synthesis.

Inventors include Dr. Bajarang Prasad Mishra; Dr. Aakunuri Manjula; Barnali Barman; Smt. Pemma Radhika; Dr. Murthy Ravaleedhar Reddy; Srisailanath; Sailen Dutta Kalita; Dr. Saurabh Kumar; Dr. Arun Kumar Sao; Chitri Rami Naidu; Dr. S. Muthuselvan; and Lakshmi S.

The application for the patent was published on August 07, 2026, under issue no. 32/2026.

Abstract: The present invention relates to a multimodal generative artificial intelligence system for unified text, image, audio, and video content synthesis. The system is configured to receive one or more user inputs, including text prompts, documents, reference images, voice samples, audio files, video clips, style preferences, and output requirements, and generate coordinated multimedia content through an integrated artificial intelligence framework. The invention includes a multimodal input processing module, contextual understanding module, text generation module, image generation module, audio generation module, video generation module, synchronization module, personalization module, validation module, and output rendering module. The multimodal input processing module extracts semantic, visual, acoustic, and temporal features from the input data. The contextual understanding module interprets user intention, content purpose, theme, tone, target audience, and output format, and prepares a unified content generation plan. Based on the plan, the generative modules create written content, visual scenes, speech narration, sound elements, animations, subtitles, and video sequences. The synchronization module aligns text, visuals, audio, and video frames to ensure smooth and contextually consistent multimedia output. The personalization module adapts the output according to language, style, voice, duration, resolution, platform, and user preference. The validation module checks quality, relevance, safety, clarity, and timing alignment before final rendering. The invention reduces dependency on multiple independent content creation tools, minimizes manual effort, improves cross-modal consistency, and enables scalable digital content generation. Therefore, the proposed system provides an intelligent, efficient, adaptive, and context-aware platform for generating high-quality multimodal content for education, advertising, entertainment, healthcare communication, business promotion, social media, and other digital applications.

Disclaimer: Curated by HT Syndication.