MUMBAI, India, Aug. 17 -- Intellectual Property India has published a patent application (202511008867 A) filed by Gdm Holding LLC on February 03, 2025, for Training Generative Neural Network Systems Using Multiple Reward Models.
Inventors include Srivastava, Pragya; Madhavan, Rahul; Shanmugam, Karthikeyan; and Raghuveer, Aravindan.
The application for the patent was published on August 14, 2026, under issue no. 33/2026.
Abstract: Methods, systems, and computer storage media are provided for training generative neural network systems, such as Large Language Models (LLMs) or Vision- Language Models (VLMs), using multiple reward models. The process involves generating one or more output sequences from a training input sequence and processing these examples using a plurality of different reward models to generate respective reward values. These reward values are combined to obtain an aggregated reward, which may be calculated as a product of the values, a weighted geometric mean, or a Nash score. The generative neural network is trained using an objective function determined using the aggregated reward, such as a cross-entropy loss or a contrastive objective function, often utilizing log score differences between the system and a reference model. These techniques allow for multi-objective alignment and are adapted for implementation on parallel processing computer systems. Figure 1 is the representative figure.
Disclaimer: Curated by HT Syndication.