LogFlex: Flexible-Bit Log Arithmetic Accelerator for Language Models on Edge

Abstract

Deploying language models on resource-constrained mobile/wearable devices while maintaining output quality is challenging. To address such challenges, many floating-point (FP) and integer (INT) quantization methods have been explored. FP arithmetic provides high quality at the cost of heavy area and energy costs, while INT-quantized models deliver superior efficiency at the cost of accuracy/perplexity loss. In addition to the design choices between FP and INT, we explore an alternative option based on a logarithmic number system (LNS), which delivers FP-like accuracy/perplexity at an efficiency close to INT. In addition to applying low-precision (8-bit) LNS, we adaptively assign bits for the INT and the fraction depending on data distribution, which enables near-FP16 accuracy/perplexity. We also co-design the LNS arithmetic and accelerator architecture, which leads to 33% less energy than the FP8 (E4M3) accelerator with similar area as an INT8 accelerator, while delivering 30% lower perplexity compared to FP8 (E4M3).

Publication
IEEE Micro
Yujin Kim
Yujin Kim
Master Student
Gunjae Koo
Gunjae Koo
Associate Professor