ICLR Poster LevAttention: Time, Space and Streaming Efficient Algorithm for Heavy Attentions

Poster

LevAttention: Time, Space and Streaming Efficient Algorithm for Heavy Attentions

Ravindran Kannan · Chiranjib Bhattacharyya · Praneeth Kacham · David Woodruff

Hall 3 + Hall 2B #462

[ Abstract ]

Fri 25 Apr midnight PDT — 2:30 a.m. PDT

Abstract: A central problem related to transformers can be stated as follows: given two

n \times d

$n \times d$ matrices

Q

$Q$ and

K

$K$ , and a non-negative function

f

$f$ , define the matrix

A

$A$ as follows: (1) apply the function

f

$f$ to each entry of the

n \times n

$n \times n$ matrix

Q K^{T}

$Q K^T$ , and then (2) normalize each of the row sums of

A

$A$ to be equal to

1

$1$ . The matrix

A

$A$ can be computed in

O (n^{2} d)

$O(n^2 d)$ time assuming

f

$f$ can be applied to a number in constant time, but the quadratic dependence on

n

$n$ is prohibitive in applications where it corresponds to long context lengths. For a large class of functions

f

$f$ , we show how to find all the "large attention scores", i.e., entries of

A

$A$ which are at least a positive value

ε

$\varepsilon$ , in time with linear dependence on

n

$n$ (i.e.,

n \cdot poly (d / ε)

$n \cdot \textrm{poly}(d/\varepsilon)$ ) for a positive parameter

ε > 0

$\varepsilon > 0$ . Our class of functions include all functions

f

$f$ of the form

f (x) = | x |^{p}

$f(x) = |x|^p$ , as explored recently in transformer models. Using recently developed tools from randomized numerical linear algebra, we prove that for any

K

$K$ , there is a "universal set"

U \subset [n]

$U \subset [n]$ of size independent of

n

$n$ , such that for any

Q

$Q$ and any row

i

$i$ , the large attention scores

A_{i, j}

$A_{i,j}$ in row

i

$i$ of

A

$A$ all have

j \in U

$j \in U$ . We also find

U

$U$ in

n \cdot poly (d / ε)

$n \cdot \textrm{poly}(d/\varepsilon)$ time. Notably, we (1) make no assumptions on the data, (2) our workspace does not grow with

n

$n$ , and (3) our algorithms can be computed in streaming and parallel settings. We empirically show the benefits of our scheme for vision transformers, showing how to train new models that use our universal set while training as well, showing that our model is able to consistently select "important keys'" during training. We also provide theoretical motivation by formulating a planted model in which our efficient algorithms provably identify relevant keys for each query.

Live content is unavailable. Log in and register to view live content