Publication:

Quantization, Sparsity, Reliability, and Their Interactions

Loading...
Thumbnail Image

Date

2025-05-22

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Park, Joshua Huanjia. 2025. Quantization, Sparsity, Reliability, and Their Interactions. Bachelors Thesis, Harvard University Engineering and Applied Sciences.

Abstract

In this thesis, we will explore and contribute to three seemingly disjoint areas of machine learning: quantization, sparsity, and reliability. While each of these areas are well-understood individually, there is very little existing literature on the interaction of these research areas. They are typically viewed as orthogonal to each other and solved independently. In this thesis, we would like to shed light on the fact that many practical systems employ elements from all three fields. Therefore, it is important to understand how they can influence and interact with each other.

This thesis makes the following three key contributions to the existing literature. First, it presents GoldenEye, a functional simulator for modeling fault injections into machine learning models. In particular, it targets models that use novel number formats to quantize or emulate to different data types. The flexible API design makes it easy to add new number formats as the area evolves. Furthermore, we propose a mathematical framework for using reliability analysis to inform quantization. Second, we present our work on EdgeBERT, an accelerator for accelerating transformer inference. EdgeBERT features both quantization and sparsity optimizations that together achieve strong speedups when profiled on true silicon. We also write compiler code that allows other transformer workloads to be compiled to EdgeBERT. Finally, we attempt to understand the problem of sequential applications of quantization and sparsity. We prove a theoretical result that shows that at the tensor-level sparsity before quantization is preferred over quantization before sparsity. However, we show that at the model level, the order is not too important, so long as the algorithms are "synergetic". Then, we propose a novel quantization-aware sparsity algorithm that consider the problem of sparsity when the weights are already quantized. Together, we believe that these contributions provide valuable insights to understanding the interactions between these distinct problems.

Description

Other Available Sources

Research Data

Keywords

Computer science, Mathematics

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories