Publication: The Context-Aware Quantization Design Space: Unlocking Scalable Training and Inference for Large AI Models
Open/View Files
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
The rapid development of large AI models capable of remarkable performance in a diverse array of complex tasks has inspired its widespread application and deployment, making the demand for scalable AI models that support fast training and inference increasingly intense. Model quantization has emerged as a critical and widely applied technique for reducing the computational and memory demands of training and inference of large deep learning models. Recently, advances in GPU tensor cores for acceleration of low-precision floating point computation have driven adoption of quantized training, in addition to quantization at inference time. In tandem, the emergence of new architectures like state space models as alternatives to transformers and quadratic attention demand a diversification of our understanding of quantization dynamics beyond ad-hoc, model-specific solutions. For a given model, dataset, and downstream task--the context of quantization--we are faced with a combinatorially large and complex design space, within which any choice can have drastic implications on the stability of quantized training and accuracy degradation at inference time. This thesis proposes advances towards a unified framework for structuring the problem space of quantization, evaluating and realizing the computational gains of quantization, and for exhaustively and comprehensively characterizing quantization in a context-aware manner. We define two design principles--systems performance and model performance--as foundational objectives for quantized optimization and inference, and propose evaluative metrics and diagnostic experiments for understanding quantization dynamics from every dimension for a wide distribution of data regimes. As a proof of concept of our framework, we conduct an empirical study of FP8 quantized training of the Mamba-2 state space model, revealing new insights into the impact of quantization on its numerical stability, gradient norm dynamics, temporal and layer-wise quantization error propagation, and catastrophic degradation phenomena.