MUMBAI, India, Aug. 10 -- Intellectual Property India has published a patent application (202631050171 A) filed by Indian Institute Of Technology Patna on April 20, 2026, for A Transformer-Based Sparse Range-Aware Audio-Visual Question Answering System And Method Thereof.

Inventors include Suman Kumar Maji; and Debashis Das.

The application for the patent was published on July 31, 2026, under issue no. 31/2026.

Abstract: The invention relates to the field of artificial intelligence-based multi-modal processing systems, and more particularly to a computer-implemented Audio-Visual Question Answering (AVQA) system and method for performing structured spatio-temporal reasoning over synchronized visual, audio, and linguistic inputs. The invention discloses a transformer-based architecture comprising a semantic-preserving temporal pre-processing module (106) for aligning and selectively retaining salient temporal information, a Sparse Range-Aware Multi-Head Attention mechanism employing dilation-adaptive attention heads for efficient multi-range contextual modeling, a transposed tensor-space attention mapping strategy for reduced computational overhead, a Multi-Scale Feedforward Network for multi-resolution feature refinement, and a question-guided hierarchical cross-modality fusion pipeline for progressive semantic alignment across modalities. An answering module (112) generates a probability distribution over a predefined answer space to output an interpretable natural language response. Compared to conventional uni-modal or dense- attention audio-visual frameworks, the present invention provides enhanced cross-modal alignment, reduced quadratic complexity, improved computational efficiency, and robust contextual reasoning in real-time multimedia environments, thereby achieving significant technical advancement in multi-modal machine intelligence systems. Refer Fig. 1

Disclaimer: Curated by HT Syndication.