MUMBAI, India, July 13 -- Intellectual Property India has published a patent application (202641082124 A) filed by Cvr College Of Engineering on July 03, 2026, for Hybrid Visual-Text Transformer For Explainable Multimodal Decision Making.

Inventors include Cindhe Ramesh; and More Swami Das.

The application for the patent was published on July 10, 2026, under issue no. 28/2026.

Abstract: The present invention relates to a multimodal artificial intelligence framework that integrates Vision Transformers (ViTs) and Large Language Models (LLMs) for enhanced retrieval and reasoning across heterogeneous data sources. The framework receives visual inputs, textual queries, or combinations thereof and converts them into unified semantic representations using a dedicated vision and language encoding modules. A cross-modal retrieval engine identifies relevant information from multimodal knowledge repositories, while an adaptive reasoning module employs attention mechanisms to correlate visual entities, textual concepts and contextual knowledge. The framework dynamically prioritizes relevant visual patches and language tokens to reduce computational overhead while improving inference accuracy. The invention further supports explainable decision generation by producing reasoning traces that associate retrieved evidence with generated outputs. The proposed architecture improves multimodal understanding, question answering, document analysis, image interpretation, and decision-support applications, offering enhanced retrieval precision, contextual awareness, scalability, and reasoning capability compared with conventional vision-language systems.

Disclaimer: Curated by HT Syndication.