MUMBAI, India, July 24 -- Intellectual Property India has published a patent application (202647084542 A) filed by Qualcomm Incorporated on July 09, 2026, for Reduced Latency For Mixed-Precision Quantized Machine Learning Models.
Inventors include Bartan, Burak; Zeng, Weiliang; Beletchi, Andrian; Patel, Chirag Sureshbhai; Kota, Yathindra; Hsieh, Kevin Lishing Number; and Khobare, Abhijit.
The application for the patent was published on July 17, 2026, under issue no. 29/2026.
Abstract: Certain aspects of the present disclosure provide techniques and apparatus for improved machine learning. In an example method, a quantization profile for a machine learning model is accessed, the quantization profile indicating, for each respective operation of a plurality of operations of the machine learning model, a respective quantization precision of a plurality of quantization precisions. A set of modifications for the quantization profile is generated based on conversion latency for converting tensors among the plurality of quantization precisions, where each respective modification of the set of modifications indicates to increase a respective quantization precision of a respective operation of the plurality of operations. A modified quantization profile is generated based on modifying the quantization profile using the set of modifications, and the machine learning model is quantized in accordance with the modified quantization profile.
Disclaimer: Curated by HT Syndication.