Publication:

The Context-Aware Quantization Design Space: Unlocking Scalable Training and Inference for Large AI Models

Loading...
Thumbnail Image

Date

2025-05-16

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Yang, Emma. 2025. The Context-Aware Quantization Design Space: Unlocking Scalable Training and Inference for Large AI Models. Bachelors Thesis, Harvard University Engineering and Applied Sciences.

Abstract

The rapid development of large AI models capable of remarkable performance in a diverse array of complex tasks has inspired its widespread application and deployment, making the demand for scalable AI models that support fast training and inference increasingly intense. Model quantization has emerged as a critical and widely applied technique for reducing the computational and memory demands of training and inference of large deep learning models. Recently, advances in GPU tensor cores for acceleration of low-precision floating point computation have driven adoption of quantized training, in addition to quantization at inference time. In tandem, the emergence of new architectures like state space models as alternatives to transformers and quadratic attention demand a diversification of our understanding of quantization dynamics beyond ad-hoc, model-specific solutions. For a given model, dataset, and downstream task--the context of quantization--we are faced with a combinatorially large and complex design space, within which any choice can have drastic implications on the stability of quantized training and accuracy degradation at inference time. This thesis proposes advances towards a unified framework for structuring the problem space of quantization, evaluating and realizing the computational gains of quantization, and for exhaustively and comprehensively characterizing quantization in a context-aware manner. We define two design principles--systems performance and model performance--as foundational objectives for quantized optimization and inference, and propose evaluative metrics and diagnostic experiments for understanding quantization dynamics from every dimension for a wide distribution of data regimes. As a proof of concept of our framework, we conduct an empirical study of FP8 quantized training of the Mamba-2 state space model, revealing new insights into the impact of quantization on its numerical stability, gradient norm dynamics, temporal and layer-wise quantization error propagation, and catastrophic degradation phenomena.

Description

Other Available Sources

Research Data

Keywords

artificial intelligence, FP8 training & inference, GPUs, machine learning, quantization, state space models, Artificial intelligence, Computer science

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories