← All posts

SVDQuant and Nunchaku

Introduction

This is a super brief summary of SVDQUant. Please refer to the original paper if interested.

QX=round⁡(X/sX) Q_X = \operatorname{round}(X / s_X)

sX=max⁡(∣X∣)/qmaxs_X = \max(\vert X \vert ) / q_{max} and qmax=possible max value in reprq_{max} = \text{possible max value in repr}

Q(X)=dequantization of X=sX QX Q(X) = \text{dequantization of }X = s_X\,Q_X

XWXW can be approximated by

XW=Q(X)Q(W)=SX SW QX QW XW = Q(X)Q(W) = S_X\,S_W\,Q_X\,Q_W

SVDQuant introduces two-path quantization.

Let's introduce a smoothing factor λ{\lambda}: X^=X diag⁡(λ)−1 \hat{X} = X\,\operatorname{diag}({\lambda})^{-1}

Then,

XW=X^ W diag⁡(λ)=X^ W^ XW = \hat{X}\,W\,\operatorname{diag}({\lambda}) = \hat{X}\,\hat{W}

Use SVD to decompose W^\hat{W} as W^=L1 L2+R,where L1=SΣ and L2=V \hat{W} = L_1\,L_2 + R,\quad \text{where } L_1 = S\Sigma \text{ and } L_2 = V

Thus, XW=X^ W^=X^ L1 L2+X^ R XW = \hat{X}\,\hat{W} = \hat{X}\,L_1\,L_2 + \hat{X}\,R

  • L1L_1 and L2L_2 are low-rank (32 in actual implementations), preserved in 16 bits.
  • Quantize X^\hat{X} and RR using W4A4 quantization.

This is open-sourced in Nunchaku. I'm also a maintainer of the project, responsible for Python engine–related tasks such as caching and adding new modules. In my last post, I mentioned my interest in contributing—and the chance has arrived.

Comments