Positional Embeddings
This paper posits there are three traits of positional embeddings which should be studied (not necessarily all useful) 1) Monotonicity: The proximity of embeddings decreases as positions grow further apart Formally, for some distance function \(\phi(x,y)\), \(|a - b | < | a - c | \implies \phi(\vec{a} - \vec{b}) < \phi(\vec{a} - \vec{c})\) where \(\vec{a}, \vec{b}\) are the positional embeddings. 2) Translation invariance: The proximity of embeddings are translation-invariant The distance between PE(1) -- PE(5) = PE(7) -- PE(13) Positional embeddings which consider relative position rather than absolute position satisfy the property of translation invariance 3) symmetry: distance between a, b = distance between b,a
Monotonicity¶
Absolute vs Relative PE¶
$$APE = [Q_x, K_x, V_x] = (WE_x + P_x) \odot [W^Q, W^K, W^V] $$ $$ RPE = [Q_x, K_x, V_x] = (WE_x) \odot [W^Q, W^K, W^V] + [0, P_{x-y}, P{x-y}] $$ for my task of constructing a time-dependent PE, i should think more about whether i want translation invariance. i.e should i treat a transplant from year 1-2 equivalent to one from year 3-4? Methinks not, since rejection risk decreases with absolute time
Questions¶
Why is #positional_embedding additive? For my purpose maybe i would want something that multiplicatively scales the input embedding, similar to a gating mechanism.
Or i could just adapt the attention scores according to some time-discounted-gated factor.