MUMBAI, India, July 30 -- Intellectual Property India has published a patent application (202641087912 A) filed by Vellore Institute Of Technology on July 17, 2026, for Method And System For Processing Devotional Audio Files.

Inventors include Valarmathi B; N. Srinivasa; A. R Ganesh; and Anbu Selvan V.

The application for the patent was published on July 24, 2026, under issue no. 30/2026.

Abstract: WE CLAIM: 1. A method for processing a devotional audio file, comprising: receiving, by a processor (100), an input audio file and a selection of a source language and a target language from a plurality of supported languages; preprocessing, by the processor (100), the input audio file to generate a normalized audio file (102) in a standard format; separating, by a vocal separation module (104), a vocal track (106) from background music (108) in the normalized audio file (102); dividing, by the processor (100), the separated vocal track (106) into a plurality of audio chunks (110); transcribing, by a speech recognition module (112), each audio chunk (110) of the plurality of audio chunks (110) to generate extracted lyrics (114) in the target language using acoustic-to-semantic word mapping; generating, by a language model (116), a devotional meaning output (118) based on the extracted lyrics (114), wherein the devotional meaning output (118) comprises a spiritual interpretation in the target language; training, by a voice conversion module (120), a voice model (122) using the separated vocal track (106); performing, by the voice conversion module (120), voice conversion on the separated vocal track (106) using the trained voice model (122) to generate converted vocals (124); and combining, by an audio mixing module (126), the converted vocals (124) with the background music (108) to generate a final output audio file (128). 2. The method of claim 1, wherein separating the vocal track (106) from the background music (108) comprises: applying a hybrid transformer model to isolate a human voice component from instrumental accompaniment in the normalized audio file (102). 3. The method of claim 1, wherein dividing the separated vocal track (106) into the plurality of audio chunks (110) comprises: segmenting the separated vocal track (106) into chunks of a predetermined duration with an overlap between consecutive chunks. 4. The method of claim 1, wherein transcribing each audio chunk (110) further comprises: detecting a transcription error in a transcribed output; applying a fallback strategy to re-transcribe the audio chunk (110) when the transcription error is detected; and injecting verified text into a context memory for processing a subsequent audio chunk (110). 5. The method of claim 1, wherein transcribing each audio chunk (110) further comprises: applying a phonetic distance optimization algorithm to correct spelling variations in the extracted lyrics (114); and performing a Unicode verification to validate that the extracted lyrics (114) belong to a script associated with the target language. 6. The method of claim 1, wherein generating the devotional meaning output (118) comprises: identifying a deity associated with the extracted lyrics (114) based on a deity identification guide; and translating an explanation of the devotional meaning from an intermediate language to the target language. 7. The method of claim 1, wherein training the voice model (122) comprises: dividing the separated vocal track (106) into a plurality of training clips (130); extracting a fundamental frequency contour from the plurality of training clips (130) for melody preservation; extracting speaker embeddings from the plurality of training clips (130) for capturing vocal characteristics; and training a synthesis model using the extracted fundamental frequency contour and the speaker embeddings. 8. The method of claim 7, wherein performing voice conversion comprises: applying the fundamental frequency contour to the converted vocals (124) to preserve an original melody of the input audio file; and applying the speaker embeddings to transfer an emotional expression from the input audio file to the converted vocals (124). 9. The method of claim 1, wherein combining the converted vocals (124) with the background music (108) comprises: adjusting a volume level of the converted vocals (124) relative to the background music (108); and mixing the converted vocals (124) and the background music (108) using an audio filter. 10. The method of claim 1, wherein the plurality of supported languages comprises Tamil, Hindi, Telugu, Malayalam, Marathi, and English.

Disclaimer: Curated by HT Syndication.