๊ฐ์ ๊ณผ ๋ ์ฐฝ์ฑ, ํต์ฌ ์๊ณ ๋ฆฌ์ฆ, ๊ทธ๋ฆฌ๊ณ ํ๊ณ
๋ฌผ์ด๋ณธ ๊ฒ: ์ด ๋ ผ๋ฌธ์ ์์ฃผ ์์ธํ๊ฒ ์ฝ๊ณ ๋ ผ๋ฌธ์ ๊ฐ์ ๊ณผ ๋ ์ฐฝ์ ์ธ ์ง์ ์ ์ค๋ช ํด์ฃผ๊ณ ํต์ฌ ์๊ณ ๋ฆฌ์ฆ์ ์์ ์ ๋ ฅ์ ๋ค์ด์ ์ ์ฒด์ ์ธ ๊ณผ์ ์ ์ค๋ช ํด์ค ์ถ๊ฐ์ ์ผ๋ก ๋ ผ๋ฌธ์ ํ๊ณ์ ์ ๋ํด์๋ ์๋ ค์ค
๐ ๊ฒฐ๋ก ์์ฝ
๋ ผ๋ฌธ “MELODI: Exploring Memory Compression for Long Contexts"๋ Transformer์ ๊ธด ๋ฌธ๋งฅ ์ฒ๋ฆฌ ๋ฌธ์ ๋ฅผ ํด๊ฒฐํ๊ธฐ ์ํด ๊ณ์ธต์ ๋ฉ๋ชจ๋ฆฌ ์์ถ(hierarchical memory compression) ๊ตฌ์กฐ๋ฅผ ์ ์ํฉ๋๋ค. ํต์ฌ์ ๋ค์ธต ๋ฐ๋ณต ์์ถ ๊ธฐ๋ฐ์ ๋จ๊ธฐ ๋ฉ๋ชจ๋ฆฌ(SM)์ ๋จ์ผ์ธต ์ถ๊ฐ ์์ถ ๊ธฐ๋ฐ์ ์ฅ๊ธฐ ๋ฉ๋ชจ๋ฆฌ(LM)๋ฅผ ์กฐํฉํ โ์๋์์น ๊ตฌ์กฐโ๋ฅผ ์ฌ์ฉํ์ฌ ๊ธด ๋ฌธ์๋ฅผ ์งง์ ์๋์ฐ(์: 512 tokens)๋ก ํจ์จ์ ์ผ๋ก ์ฒ๋ฆฌํ๋ ๊ฒ์ ๋๋ค. ๋ฉ๋ชจ๋ฆฌ ์ฌ์ฉ๋์ ๊ธฐ์กด Dense Memory ๋ฐฉ์์ธ Memorizing Transformer๋ณด๋ค ์ต๋ 8๋ฐฐ ์ ๊ฐํ๋ฉด์๋ ์ฑ๋ฅ(PPL ๊ธฐ์ค)์ ์คํ๋ ค ํฅ์๋ฉ๋๋ค.
๐ MELODI์ ๊ตฌ์กฐ ์์ฝ
| ๊ตฌ์ฑ ์์ | ํน์ง |
|---|---|
| Short-Term Memory | - ์๋์ฐ ๋จ์๋ก Recurrentํ๊ฒ ์ ๋ณด๋ฅผ ์์ถ - Layer๋ฅผ ๋ฐ๋ผ ์ ๋ณด๋ฅผ ์ ๋ฌ (vertical) |
| Long-Term Memory | - Layer ์ค๊ฐ ์ง์ ์์ ์ถ๊ฐ ์์ถํ์ฌ ์ ์ฅ - ์๊ฐ ์์ ๋ฐ๋ผ ์ ๋ณด๋ฅผ ์ถ์ (horizontal) |
| ๊ตฌ์กฐ | - [SM ร M์ธต] + [LM ร 1์ธต] + [SM ร (NโMโ1)์ธต] ์ ์๋์์น ๊ตฌ์กฐ ์ฌ์ฉ |
์์ ์ ๋ ฅ ํ๋ฆ (๋จ์ผ ์๋์ฐ ๊ธฐ์ค)
์
๋ ฅ: xโ (k๋ฒ์งธ context window, ์: 512 tokens)
SM Layer 1
zโโโ(์ด์ SM token)๊ณผxโ์ causal attention- Transformer block โ
xโ โ xโ', summary tokenuโ'๊ณ์ฐ - Linear Mixer:
uโ = Mโ(xโ', uโ'),zโ = Mโ(xโ', uโ')
… SM Layer M ๋ฐ๋ณต โ recurrent ์์ถ ์งํ
LM Layer (์ค๊ฐ layer)
- ์ง๊ธ๊น์ง ์ ์ฅ๋ LM (
mโ:โโโ)์ cross-attention ์ํ - self-attention๊ณผ cross-attention์ gating ฮฑ๋ก ํฉ์ฑ
- LM token
mโ์์ฑ ํ KV pair ํํ๋ก ๋ฉ๋ชจ๋ฆฌ์ append
- ์ง๊ธ๊น์ง ์ ์ฅ๋ LM (
์ดํ Layer N๊น์ง ๋ค์ SM ๋ฐ๋ณต ์ํํ์ฌ ์ต์ข ์ถ๋ ฅ
๐ง ํต์ฌ ์๊ณ ๋ฆฌ์ฆ: ์์ ๊ธฐ๋ฐ ์ค๋ช
๊ฐ์ :
- context window ๊ธธ์ด: 512 tokens
- short-term memory: 128 tokens
- long-term memory: 64 tokens, ์ต๋ 128 window ์ ์ฅ
์
๋ ฅ ๋ฌธ์ฅ (k=3๋ฒ์งธ ์๋์ฐ): "In the middle of the night, he found a strange box."
==> Layer 1์์:
- ์ด์ zโ (128 token)์ ํ์ฌ context xโ์ attention
- context xโ โ xโโ ๊ณ์ฐ
- summary token uโโ ์์ฑ
- Linear Mixer Mโ, Mโ๋ฅผ ๊ฑฐ์ณ:
uโ = summary for ๋ค์ layer
zโ = short-term memory token (โ window 4 ์
๋ ฅ ์ ์ฌ์ฉ)
==> ์ค๊ฐ LM Layer:
- ํ์ฌ xโโ, uโ์ mโ:โ (์ด์ ๊น์ง์ LM KV pair)์ cross-attention
- ์๋ก์ด long-term token mโ ์์ฑ ํ long-term memory์ append
==> ์ดํ layer์์๋ zโ, uโ ์ ๋ณด ์ด์ฉํ์ฌ inference ์ด์ด๊ฐ๐ ์ฑ๋ฅ ๋น๊ต (Perplexity ๊ธฐ์ค)
| Model | PG19 (T5 vocab) | arXiv (Meena) | Memory Usage |
|---|---|---|---|
| Transformer-XL | 11.41 | 2.60 | 13.6M |
| Block Recurrent Transf. | 10.98 | 2.26 | 13.1M |
| Memorizing Transf. | 10.62 | 2.14 | 147.8M |
| MELODI S128+L64 | 10.44 | 2.11 | 18.5M |
| MELODI S192+L96 | 10.29 | 2.09 | 27.8M |
๐ Dense attention ์์ด๋ ์ฑ๋ฅ์ ๋ ๋๊ณ , ๋ฉ๋ชจ๋ฆฌ ์ฌ์ฉ๋์ ํ๊ธฐ์ ์ผ๋ก ์ค์
๐งช Ablation์ผ๋ก ๋ฐํ์ง ์ค๊ณ์ ํจ๊ณผ
- SM๊ณผ LM์ ์ํธ๋ณด์์ : ๋ ๋ค ํค์ฐ๋ฉด ์ฑ๋ฅ ์์น (Fig. 4)
- Context window ์๊ฒ ์ค์ฌ๋ ์ฑ๋ฅ ์ ์ง: LM์ด ์ฅ๊ธฐ ๋ฌธ๋งฅ์ ๋ณด์กดํจ (Fig. 6)
- Summary branching: cross-window + cross-layer ํ๋ฆ ๋์ โ ์ฑ๋ฅ ํฅ์ (+0.3 PPL)
- LM ์์น๋ 5~11์ธต ์ด๋๋ ๋น์ทํ๊ฒ ๋์ โ ์ ์ฐํ ๊ตฌ์กฐ ์ค๊ณ ๊ฐ๋ฅ
๐ ๋ ผ๋ฌธ์ ๊ฐ์
| ๊ฐ์ | ์ค๋ช |
|---|---|
| ๋ฉ๋ชจ๋ฆฌ ํจ์จ | Dense KV ์ ์ฅ ๋์ ์์ถ๋ token๋ง ์ ์ฅํด ์ต๋ 8x ์ ๊ฐ |
| ๊ตฌ์กฐ ์ผ๋ฐ์ฑ | ๊ธฐ์กด Transformer์ ๊ฑฐ์ ์๋์ง ์๊ณ ํ์ฅ ๊ฐ๋ฅ |
| ์ฑ๋ฅ ์ ์ง | ๊ธฐ์กด state-of-the-art๋ณด๋ค ์ข์ perplexity |
| ๋จ๊ธฐ-์ฅ๊ธฐ ๋ฉ๋ชจ๋ฆฌ ๊ฒฐํฉ | LSTM ์คํ์ผ recurrent compression + ๋จ์ผ์ธต summary ์ ์ฅ |
| ๋ชจ๋ํ | Token mixer, gating, memory branching ๋ฑ์ ๊ตฌ์กฐ๊ฐ ๊น๋ํ ๋ถ๋ฆฌ๋์ด ์์ |
โ ํ๊ณ์
| ํ๊ณ | ์ค๋ช |
|---|---|
| ์ฌ์ ํ์ต ๋ชจ๋ธ ์ ์ฉ X | MELODI๋ ์ฒ์๋ถํฐ ํ์ตํจ. ๊ธฐ์กด ์ฌ์ ํ์ต ๋ชจ๋ธ์ plug-in ํ๋ ๋ฐฉ์์ ์์ง ์์ |
| ์ ์ฉ ๋ณต์ก๋ | Short-term summary token flow, token mixer ๋ฑ ๊ตฌํ ๋ณต์ก์ฑ์ด ๋์ |
| Memory queue ๊ณ ์ | LM์ FIFO ํ์ KV pair ์ ์ฅ โ ํ์ต ์ธ ๊ธฐ์ต ์ ์ง ํ์ ์ ํ๊ณ ์์ |
| ์์ถ ์ ๋ณด ์์ค ๊ฐ๋ฅ์ฑ | ์ง๋์น๊ฒ aggressiveํ ์์ถ์ ์ ๋ณด ์ ์ค ์ ๋ฐ ๊ฐ๋ฅ (ablation์์ ํ์ธ๋จ) |
๐ ๊ฒฐ๋ก ๋ฐ ํ์ฉ ์ ์
MELODI๋ long-context language modeling์์ ํจ์จ์ฑ๊ณผ ์ฑ๋ฅ์ ๊ท ํ์ ๋งค์ฐ ์ ๋ง์ถ ๊ตฌ์กฐ์ ๋๋ค. ํนํ:
- ๋ฉ๋ชจ๋ฆฌ ์ ํ์ด ์๋ ๋๋ฐ์ด์ค (e.g., ๋ชจ๋ฐ์ผ inference, edge GPU)
- ๊ธด ๋ฌธ์๋ฅผ chunk ๋จ์๋ก ์ฒ๋ฆฌํด์ผ ํ๋ task (e.g., RAG, summarization)
- ์ปจํ ์คํธ window ํ์ฅ์ด ์ด๋ ค์ด ๋ชจ๋ธ ํ์ธํ๋ ์, hierarchical memory ์ถ๊ฐ ๊ตฌ์กฐ
๋ฑ์ ๋งค์ฐ ์ ํฉํฉ๋๋ค.
Q : gemma์ ๊ฐ์ ๋ชจ๋ธ์์ ์ฌ์ฉํ๋ slice window attention๊ณผ ๋น์ทํ ๊ฐ๋ ์ธ๊ฑฐ์ผ? ์ด๋ค ์ฐจ์ด์ ์ด ์๋์ง ์๋ ค์ค ์ ์๋? ๊ทธ๋ฆฌ๊ณ ์๊ฐ์ ์ธ ๋ค์ด์ด๊ทธ๋จ๋ ๋ถํํด
MELODI๋ Gemma ๋ฑ์์ ์ฌ์ฉํ๋ Slice Window Attention (SWA)๊ณผ ๋น์ทํ๊ฒ ์งง์ context window๋ง์ ์ฌ์ฉํ์ง๋ง, ํต์ฌ ์ฐจ์ด์ ์ “์ด๋ป๊ฒ ๊ณผ๊ฑฐ ์ ๋ณด๋ฅผ ์์ถยท์ถ์ ํ๋๋”์ ์์ต๋๋ค.
๐ MELODI vs Slice Window Attention: ๋น๊ต
| ํญ๋ชฉ | MELODI | Slice Window Attention (์: Gemma) |
|---|---|---|
| ํต์ฌ ๊ฐ๋ | ๊ณ์ธต์ ๋ฉ๋ชจ๋ฆฌ ์์ถ (SM+LM) | sliding ๋๋ sliced attention window |
| ๋ฌธ๋งฅ ์ฐ๊ฒฐ ๋ฐฉ์ | - ๋จ๊ธฐ: recurrent compression - ์ฅ๊ธฐ: middle-layer์์ ์์ถ๋ KV ์ ์ฅ | - ์ต๊ทผ n๊ฐ์ window๋ง sliding ์ฐธ์กฐ |
| ์ฅ๊ธฐ ์ ๋ณด ์ ์ง | โ long-term memory (128 window๊น์ง) | โ ๊ธด ๊ฑฐ๋ฆฌ ์ ๋ณด๋ ์ํ์ง |
| Attention ๋์ | SM: ์ด์ memory (z) LM: ์ ์ฅ๋ KV pool (mโ:โโโ) | ํ์ฌ ๋ฐ ์ผ๋ถ ๊ณผ๊ฑฐ context tokens |
| ๊ตฌ์กฐ ๋ณ๊ฒฝ ์ ๋ | transformer layer์ memory path ์ฝ์ (๋ฏธ์ธ ๋ณ๊ฒฝ) | attention mask๋ง ์กฐ์ ํ์ฌ ๊ฐ๋จํ ๊ตฌํ ๊ฐ๋ฅ |
| ์์ถ ์ฌ๋ถ | โ context โ summary token โ KV pair๋ก ์์ถ | โ ์์ถ ์์, token ๊ทธ๋๋ก ์ฌ์ฉ |
| ๋ฉ๋ชจ๋ฆฌ ์ฌ์ฉ๋ | ๋ฎ์ (์์ถ๋ token๋ง ์ ์ฅ) | ์ค๊ฐ (token ์ง์ ์ ์ฅ) |
ํต์ฌ ์ฐจ์ด
- SWA๋ “์ต๊ทผ ํ ํฐ๋ง์ ๊ธฐ์ต"ํ๋ ๋ก์ปฌ context ์ฐ๊ฒฐ ๋ฐฉ์
- MELODI๋ “์ ๋ณด ์์ฝ ๋ฐ ์ถ์ "์ ํตํด ์ ์ญ context ์ฐ๊ฒฐ ์ ์ง
์ฆ, SWA๋ locality ๊ธฐ๋ฐ, MELODI๋ hierarchical memory abstraction ๊ธฐ๋ฐ์ ๋๋ค.
์๋๋ ์์ฒญํ์ MELODI ๊ตฌ์กฐ์ ์๊ฐํ์
๋๋ค (์๋์์น ๊ตฌ์กฐ, SM๊ณผ LM์ ํ๋ฆ ํฌํจ):
๊ฑฐ์ ๋ง์ต๋๋ค. ๊ฐ๋ ์ ์ ๋ฆฌํ์๋ฉด ์๋์ ๊ฐ์ด ์ ๋ฆฌํ ์ ์์ต๋๋ค.
โ ํต์ฌ ์ฐจ์ด ์์ฝ: SWA vs MELODI
| ํญ๋ชฉ | Slice Window Attention (SWA) | MELODI |
|---|---|---|
| ์ ์ฅ ๋จ์ | ๐น token-level (์๋์ token sequence ์ผ๋ถ) | ๐ธ compressed embedding-level (context window ์์ฝ๋ณธ) |
| ์ ์ฅ ๋์ | - ์ต๊ทผ window์ raw token์ ์ง์ attention ๋์์ ํฌํจ (์: ์ต๊ทผ 2~4๊ฐ window) | - ๊ฐ window์์ ์์ถ๋ summary token ๋๋ KV embedding์ ์ ์ฅ |
| ์ ๋ณด ํํ | - ์ ๋ณด ์์ค ์์ด ๊ทธ๋๋ก ์ ์ฅ - ํ์ง๋ง memory size โโ | - ์ ๋ณด ์์ฝ/์์ถ ํ ์ ์ฅ - memory size โโ |
| Memory ๊ตฌ์กฐ | - Sliding ๋ฐฉ์์ cache | - FIFO queue ํํ์ long-term memory |
| Attention ๋ฐฉ์ | - ํ์ฌ token์ด ์ง์ ์ด์ token๋ค๊ณผ attention | - ํ์ฌ token์ด ์์ฝ๋ memory representation์ attention |
์์๋ก ์ดํดํ๊ธฐ
SWA:
PLAINTEXTWindowโ = [tโโโ, tโโโ, tโ] ์ค tโโโ, tโโโ๋ ๊ทธ๋๋ก ์ ์ฅ โ attention ๋์Windowโ = [tโโโ, tโโโ, tโ] ์ค tโโโ, tโโโ๋ ๊ทธ๋๋ก ์ ์ฅ โ attention ๋์MELODI:
PLAINTEXTWindowโโโ โ summary token โ linear mixer โ compressed KV โ ์ ์ฅ๋จ ํ์ฌ window๋ ์ด summary KV์๋ง cross-attentionWindowโโโ โ summary token โ linear mixer โ compressed KV โ ์ ์ฅ๋จ ํ์ฌ window๋ ์ด summary KV์๋ง cross-attention
๊ฒฐ๋ก
“SWA๋ ๊ณผ๊ฑฐ raw token์ ์ง์ ์ ์ฅํด์ ๋ณด๋ ๋ฐฉ์์ด๊ณ , MELODI๋ ๊ณผ๊ฑฐ context๋ฅผ ์์ฝํ ์์ถ embedding์ ์ ์ฅํด์ ๋ณด๋ ๋ฐฉ์์ด๋ค.”
๋ฐ๋ผ์ MELODI๋ memory usage ๋ฉด์์ ํจ์ฌ ํจ์จ์ ์ด๋ฉฐ, ์ฅ๊ธฐ dependency ์ ์ง์ ๋ ์ ๋ฆฌํฉ๋๋ค.
ํต์ฌ ์๊ณ ๋ฆฌ์ฆ์ ์์๋ก ๋ฐ๋ผ๊ฐ๊ธฐ
๋ฌผ์ด๋ณธ ๊ฒ: ๋ ผ๋ฌธ์ ํต์ฌ์๊ณ ๋ฆฌ์ฆ์ ์์๋ฅผ ๋ค์ด ์์ฃผ ์์ธํ๊ฒ ์ค๋ช ํด์ฃผ๊ณ ์ ๋ฆฌํด ์ค
๋ ผ๋ฌธ MELODI์ ํต์ฌ ์๊ณ ๋ฆฌ์ฆ์ ๊ธด ๋ฌธ์๋ฅผ ์งง์ context window๋ก ๋๋ ์ ์ฒ๋ฆฌํ๋ฉด์๋ ์ด์ ๋ฌธ๋งฅ์ ์ ๋ณด๋ฅผ ํจ๊ณผ์ ์ผ๋ก ์์ฝยท๋ณด์กดํ๊ธฐ ์ํ ๊ณ์ธต์ ๋ฉ๋ชจ๋ฆฌ ์์ถ(Hierarchical Memory Compression) ๊ธฐ๋ฒ์ ๋๋ค. ์๋์ ์๊ณ ๋ฆฌ์ฆ์ ์์ ์ค์ฌ์ผ๋ก ๋จ๊ณ๋ณ๋ก ์์ฃผ ์์ธํ ์ค๋ช ๋๋ฆฌ๊ฒ ์ต๋๋ค.
โ ์ ์ฒด ๊ตฌ์กฐ ์์ฝ
์
๋ ฅ ์ํ์ค X = [xโ, ..., x_T]๋ 512 tokens ๋จ์์ context window xโ๋ก ๋๋ฉ๋๋ค. ๊ฐ window๋ ๋ค์ ๋ ๊ฐ์ง memory ๊ตฌ์กฐ๋ฅผ ์ฌ์ฉํฉ๋๋ค:
Short-Term Memory (STM):
- ๊ฐ window ๋ด์์ layer๋ณ๋ก recurrent compression
- ์: 512 tokens โ 128 tokens
- window ๊ฐ zโ (์์ถ๋ ๋ฉ๋ชจ๋ฆฌ) ์ ๋ฌ
Long-Term Memory (LTM):
- ํน์ ์ค๊ฐ layer์์ 64๊ฐ ์๋ฒ ๋ฉ์ผ๋ก window ์ ์ฒด๋ฅผ ์์ฝ
- ์ด์ window์ ์์ฝ๊ฐ๋ค์ FIFO queue๋ก ์ ์ฅ
- ์ด ๋ฉ๋ชจ๋ฆฌ์ ๋ํด cross-attention ์ํ
๐งช ์์: ์ ์ฒด ์๊ณ ๋ฆฌ์ฆ ๋์ ํ๋ฆ
๊ฐ์
- context window ๊ธธ์ด: 512 tokens
- short-term memory token ์ S = 128
- long-term memory token ์ L = 64
- window index:
k = 3(์ธ ๋ฒ์งธ context window ์ฒ๋ฆฌ ์ค)
โถ๏ธ Step 1: Input ์ค๋น
์ ๋ ฅ
xโ = [tokenโ, tokenโ, ..., tokenโ
โโ]์ด์ memory
zโ: ์ด์ window์ short-term memory (128 vectors)mโ:โ: ์ด์ ๋ window์ long-term memory (64ร2 = 128 KV pair)
โถ๏ธ Step 2: Short-Term Memory ์ฒ๋ฆฌ (M๊ฐ์ layer ๋ฐ๋ณต)
๊ฐ short-term layer์์:
zโ์ ํ์ฌ tokenxโ์ ๋ํด causal attention- context token โ transformer โ
xโโฒ๋ก ์ ๋ฐ์ดํธ - summary token
uโ์์ฑ (128 tokens) - summary token๊ณผ context token์ ํตํด ๋ค์ window์ฉ memory token
zโ์์ฑ
๐ก ์์ ์ ๋ฆฌ:
xโโฒ = T(xโ | zโ)
uฬโ = T(uโ | xโ, zโ)
zโ = Mโ(xโโฒ, uฬโ)
uโ = Mโ(xโโฒ, uฬโ)โถ๏ธ Step 3: Long-Term Memory ์ฒ๋ฆฌ (๋จ์ผ middle layer์์๋ง)
- xโโฒ, uโ๊ฐ ์ด์ LTM (mโ:โ)์ ๋ํด cross-attention ์ํ
- ๊ฒฐ๊ณผ๋ฅผ self-attention ๊ฒฐ๊ณผ์ gatingํ์ฌ ์กฐํฉ:
A = ฮฑ โ
A_cross + (1 โ ฮฑ) โ
A_self- ์ด ๊ฒฐ๊ณผ๋ฅผ
L = 64๊ฐ์ vector๋ก ์ถ๊ฐ ์์ถ - ์์ถ๋ vector๋ฅผ KV๋ก ๋ณํํ์ฌ
mโ๋ก long-term memory์ ์ถ๊ฐ
โถ๏ธ Step 4: Output ์ ๋ฌ
- ์ ๊ณผ์ ์ ํตํด ๋ค์ window์์๋
zโ,mโ:โ๋ฅผ ์ฌ์ฉ - ํด๋น window ๋ด layer๋
uโ์xโโฒ๋ฅผ ๋ค์ layer๋ก ๋๊น
๐ ์ ์ฒด ๋ฐ๋ณต
์ด ๊ณผ์ ์ ๊ฐ window๋ง๋ค ๋ฐ๋ณตํ๋ฉฐ ๊ธด ๋ฌธ๋งฅ์ ์ด์ด๊ฐ๊ณ , ํ์ํ ๋ฌธ๋งฅ์:
- ์ง์ short memory
zโโโ - ์ฅ๊ธฐ ์์ฝ ๋ฉ๋ชจ๋ฆฌ
mโ:โโโ
๋ฅผ ํตํด ์ ๋ฐฉํฅ์ผ๋ก context๋ฅผ ์ฐ๊ฒฐํฉ๋๋ค.
๐ ์ ๋ฆฌ: ์๊ณ ๋ฆฌ์ฆ ํ๋ฆ๋
| ๋จ๊ณ | ๋์ | ์ ๋ ฅ | ์ถ๋ ฅ |
|---|---|---|---|
| โ | Short-term attention | xโ, zโโโ | xโโฒ, uโ |
| โก | Summary compression | xโโฒ, uโ | zโ (๋ค์ window์ฉ) |
| โข | Long-term memory attention | xโโฒ, uโ, mโ:โโโ | Aโ (gated attention result) |
| โฃ | Long-term token ์์ฑ | Aโ โ linear mixer | mโ (KV ํํ ์ ์ฅ) |
| โค | Forwarding | xโโฒ, uโ, zโ, mโ:โ | ๋ค์ window ์ฒ๋ฆฌ์ ์ฌ์ฉ |
๐ง ์ฃผ์ ํน์ฑ
- ๐พ Memory-efficient: long memory๋ KV๋ง ์ ์ฅ
- ๐ Recurrent compression: LSTM ์คํ์ผ ์ ๋ณด๋ฅผ ์์ฝ
- โ๏ธ Global context retention: ์ ์ฒด window history ๋ณด์กด ๊ฐ๋ฅ
์คํ ๊ฒฐ๊ณผ์ ๋ค๋ฅธ ๋ฐฉ๋ฒ๋ก ๊ณผ์ ์ฐจ์ด
๋ฌผ์ด๋ณธ ๊ฒ: ์ด ๋ ผ๋ฌธ์์ ์ ์ํ ๊ฒฐ๊ณผ๋ฅผ ์์ธํ๊ฒ ๋ณด๊ณ ๋ค๋ฅธ ๋ฐฉ๋ฒ๋ก ์ ๋นํ์ฌ ํน์ถ๋ ์ ๊ณผ ๋ ผ๋ฌธ์์ ์ ๊ธฐํ๋ ์ด๋ ํ ๋ฐฉ๋ฒ์ด ์ด๋ฌํ ๊ฒฐ๊ณผ๋ฅผ ๋์ถํ๊ฒ ๋์๋์ง ๋ ผ๋ฌธ์์ ์ ์ํ๋ ์ด์ ์ ๋์ ์๊ฐ์ ์๋ ค์ค
์ข์ต๋๋ค. ์๋์ ๋ ผ๋ฌธ MELODI๊ฐ ์ ์ํ ์คํ ๊ฒฐ๊ณผ ๋ถ์, ๊ธฐ์กด ๋ฐฉ๋ฒ๋ก ๊ณผ์ ๋น๊ต, ๊ทธ๋ฆฌ๊ณ ๊ทธ ๊ฒฐ๊ณผ๋ฅผ ๋์ถํ ์ค๊ณ ์์ธ๊ณผ ํ์์ ๊ฒฌํด๋ฅผ ์์น ๊ธฐ๋ฐ์ผ๋ก ์ ๋ฆฌํด ๋๋ฆฌ๊ฒ ์ต๋๋ค.
โ 1. ์คํ ๊ฒฐ๊ณผ ์ ๋ฆฌ (Perplexity ๊ธฐ์ค ์ฑ๋ฅ ๋น๊ต)
| ๋ชจ๋ธ | PG19 (T5 vocab) | arXiv (Meena) | C4(4K+) | ์ด ๋ฉ๋ชจ๋ฆฌ ์ฌ์ฉ๋ | ์ฅ๊ธฐ ๋ฉ๋ชจ๋ฆฌ | ๋จ๊ธฐ ๋ฉ๋ชจ๋ฆฌ |
|---|---|---|---|---|---|---|
| Transformer XL | 11.41 | 2.60 | 18.22 | 13.6M | โ ์์ | 13.6M |
| Block Recurrent Transf. | 10.98 | 2.26 | 17.82 | 13.1M | โ ์์ | 13.1M |
| Memorizing Transformer | 10.62 | 2.14 | 17.37 | 147.8M | 134.2M | 13.6M |
| MELODI (S192+L96) | 10.29 | 2.09 | 17.25 | 27.8M | 25.2M | 2.6M |
๐ ํต์ฌ ์ฑ๊ณผ ์์ฝ:
- Memorizing Transformer๋ณด๋ค ์ฑ๋ฅ ํฅ์ (PG19: โ0.33 PPL)
- Memory ์ฌ์ฉ๋ 5.3๋ฐฐ ๊ฐ์
- Transformer-XL ๋๋น ๋ฉ๋ชจ๋ฆฌ ์ฌ์ฉ๋ โ 19%, ์ฑ๋ฅ์ โ
โ 2. ๊ธฐ์กด ๋ฐฉ๋ฒ๋ก ๋๋น MELODI์ ํน์ถ๋ ์
| ํญ๋ชฉ | Memorizing Transformer (MT) | MELODI |
|---|---|---|
| KV ์ ์ฅ ๋ฐฉ์ | context token ์ ์ฒด KV ์ ์ฅ (dense) | context window๋ฅผ ์์ถํ low-dim KV ์ ์ฅ |
| ๋จ๊ธฐ ๋ฉ๋ชจ๋ฆฌ | ์์ or top-layer LSTM์ ์ฉ | ์ ์ธต์ ๊ฑธ์น multi-layer recurrent compression |
| ๋ฉ๋ชจ๋ฆฌ ์ฉ๋ | 64K token KV ์ ์ฅ | 64 compressed KV / window ร 128 windows |
| ์ฑ๋ฅ ํจ์จ tradeoff | long-term memory ํฌ๋ฉด ์ฑ๋ฅ ์ฆ๊ฐ โ but ๋ฉ๋ชจ๋ฆฌ ๊ธ์ฆ | ์์ memory footprint๋ก๋ ์ฑ๋ฅ ์ ์ง |
| ๋ฉ๋ชจ๋ฆฌ ๊ตฌ์กฐ ํตํฉ | ๋จ์ผ layer๋ง ์ฌ์ฉ | SM + LM์ ๊ณ์ธต ๊ตฌ์กฐ ์ค๊ณ |
| ์์ฝ ํ ํฐ ์ฌ์ฉ | X | summary token์ผ๋ก SM-๊ฐ ์ฐ๊ฒฐ ๊ฐํ |
โ 3. ๋ ผ๋ฌธ์ด ์ค๋ช ํ๋ ์ฑ๋ฅ ํฅ์์ ์ด์
๋ ผ๋ฌธ์ ์๋ ์ธ ๊ฐ์ง๋ฅผ ์ฑ๋ฅ ํฅ์์ ํต์ฌ ์ด์ ๋ก ์ ์ํฉ๋๋ค.
โ ๊ณ์ธต์ ๋ฉ๋ชจ๋ฆฌ ๊ตฌ์กฐ (hierarchical compression)
- SM: window ๋ด๋ถ ์ ๋ณด๋ฅผ ์ฌ๋ฌ ์ธต์ ํตํด ์์ถ โ
512 โ 128 tokens - LM: window ์ ์ฒด๋ฅผ ์์ฝํ์ฌ ๋จ์ผ layer์์
โ 64 tokens - โ ์์ฝ ์์ค์ ๊ณ์ธต์ ์ผ๋ก ๋ณด์, ๋จ์ผ ๋ฐฉ์๋ณด๋ค ์ ๋ณด ๋ณด์กด ์ฐ์
๐ Ablation์์ ํ์ธ๋จ: SM ๋๋ LM๋ง ์ฌ์ฉ ์ PPL โ, ๋ ํจ๊ป ์ธ ๋ ๊ฐ์ฅ ๋ฎ์
โก Summary branching (๋จ๊ธฐ ๊ธฐ์ต์ layer ๊ฐ ์ ํ ๊ฐํ)
- summary token์ ๋ค์ layer๋ฟ ์๋๋ผ ๋ค์ window์๋ ์ ๋ฌ
- โ memory ๊ฐ cross-layer / cross-window ์ ๋ณด ํ๋ฆ ํ์ฑ
- โ PPL ์ฝ 0.3 ๊ฐ์ ํจ๊ณผ
๋ ผ๋ฌธ Table 4:
PLAINTEXTw/o branching: 11.24 w/ branching: 10.95w/o branching: 11.24 w/ branching: 10.95
โข Gated cross-attention in LM
- LM layer์์ self-attn๊ณผ cross-attn (long-term memory)์ ๊ฐ์ค ์กฐํฉ
A = ฮฑโ A_cross + (1โฮฑ)โ A_self- โ ์ฅ๊ธฐ ๊ธฐ์ต์ ๊ณผ๋ํ๊ฒ ์์กดํ์ง ์๋๋ก ์กฐ์ ๊ฐ๋ฅ
- ํ์ต ์ค ๊ฐ head๋ง๋ค ฮฑ ํ์ต ๊ฐ๋ฅ
๐ค ๋ด ๊ฒฌํด ๋ฐ ํ๊ฐ
์ด๋ก ์ ์์ฑ๋ MELODI๋ ๊ธฐ์กด ๋ฐฉ์๋ค๋ณด๋ค Transformer ์ํคํ ์ฒ์ ์์ฐ์ค๋ฝ๊ฒ ํตํฉ๋๋ฉฐ, ๊ตฌ์กฐ์ ๋ณ๊ฒฝ์ด ์ต์ํ๋๋ฉด์๋ ์ฅ๋จ๊ธฐ ์ ๋ณด๋ฅผ ๋ชจ๋ ํฌ๊ดํ๋ค๋ ์ ์์ ์ด๋ก ์ ์์ฑ๋๊ฐ ๋๋ค๊ณ ํ๊ฐํ ์ ์์ต๋๋ค.
์์ถ ๋ฐฉ์์ ์คํจ์ฑ Summary token ๊ธฐ๋ฐ์ ์์ถ ๋ฐฉ์์ ์ ๋ณด๋ฅผ ์์ ํ ์ ์งํ์ง ๋ชปํ ์ ์์์๋, ๊ณ์ธต์ ๋ฐ๋ณต ๊ตฌ์กฐ์ summary branching์ ํตํด ์ด๋ฅผ ๊ทน๋ณตํ ๊ฒ์ด ํจ์จ์ฑ๊ณผ ์ฑ๋ฅ ๊ฐ ๊ท ํ์ ์ ๋ง์ถ ์ค๊ณ๋ผ๊ณ ๋ด ๋๋ค.
ํ์ค์ ์ ์ฉ์ฑ Memory size๋ฅผ 8๋ฐฐ ์ค์ด๋ฉด์๋ MT๋ณด๋ค ๋ ์ข์ ์ฑ๋ฅ์ ๋ณด์ธ๋ค๋ ์ ์ ๋ชจ๋ฐ์ผยท์์ง ๋๋ฐ์ด์ค๋ ์ ์์ ํ๊ฒฝ์์ ๋งค์ฐ ํฐ ์ฅ์ ์ ๋๋ค. ํนํ context window๋ฅผ ์ค์ด๋ฉด์๋ ์ฑ๋ฅ ์ ์ง ๊ฐ๋ฅ์ฑ๋ ๋์ต๋๋ค.
๐ ๊ฒฐ๋ก
MELODI์ ๊ฐ์ ์ ๋จ์ํ ๋ฉ๋ชจ๋ฆฌ๋ฅผ ์ค์ด๋ ๊ฒ์ด ์๋๋ผ, ์ฅ๋จ๊ธฐ ๋ฌธ๋งฅ ์ ๋ณด๋ฅผ ์์ถํ๊ณ ์ ๋ฌํ๋ ๋ฐฉ์ ์์ฒด๋ฅผ ์ต์ ํํ๋ค๋ ์ ์ ๋๋ค. ์ด๋ ์ด์ ๋ฐฉ๋ฒ๋ค์ด ๋ฉ๋ชจ๋ฆฌ๋ฅผ ๋๋ฆฌ๋ฉด ์ฑ๋ฅ์ด ์ข์์ง๋ค๋ ๋ฐฉํฅ์ฑ์ ๊ฐ์ก๋ ๊ฒ๊ณผ๋ ๋ค๋ฅธ, ์์ถยท์์ฝ ๊ธฐ๋ฐ์ ์๋ก์ด ๋ฉ๋ชจ๋ฆฌ ์ค๊ณ ํจ๋ฌ๋ค์์ ์ ์ํ ๊ฒ์ด๋ผ ๋ณผ ์ ์์ต๋๋ค.
ํ์ํ์๋ค๋ฉด ์ด ์คํ ๊ฒฐ๊ณผ๋ฅผ ์๊ฐํํ ํ๋ ๊ทธ๋ํ๋ ์ ์ํด ๋๋ฆด ์ ์์ด์. ์ํ์๋์?
์ค์ ๋ก ์ ์ฉํ๋ ค๋ฉด: ์ถ๊ฐ ๊ตฌํ๊ณผ ์์ ์ฐ์
๋ฌผ์ด๋ณธ ๊ฒ: ์ด ๋ ผ๋ฌธ์์ ์ ์๋ ๋ฐฉ๋ฒ์ ์ค์ ๋ก ์ฌ์ฉํ๋ ค๋ฉด ๊ธฐ์กด ๋ฐฉ๋ฒ์์ ์ด๋ ํ ๊ตฌํ์ด ์ถ๊ฐ์ ์ผ๋ก ํ์ํ๊ณ ์ด๋ฌํ ๊ตฌํ์ ์๋ชจ๋๋ ๊ณต์ ๋ฐ ์ฐ์ฐ์ ํ์ํ ์ปดํจํ ์์์ ํฌ๊ธฐ์ ๋ํด์ ๊ณ์ฐํด์ค ์ ์๊ฒ ๋?
์ข์ต๋๋ค. ์๋์ ๋ ผ๋ฌธ MELODI์ ๊ตฌ์กฐ๋ฅผ ๊ธฐ์กด Transformer (์: GPT, LLAMA ๋ฑ) ์์ ์ฌ๋ฆฐ๋ค๊ณ ๊ฐ์ ํ ๋, ์ค์ ๊ตฌํ์ ํ์ํ ๊ตฌ์กฐ์ ๋ณ๊ฒฝ, ์ฐ์ฐ๋ ์ฆ๊ฐ, ํ๋ผ๋ฏธํฐ ์ ์ฆ๊ฐ, ๊ทธ๋ฆฌ๊ณ ๋ฉ๋ชจ๋ฆฌ/์ฐ์ฐ ๊ณต์๋ฅผ ์ ๋์ ์ผ๋ก ๊ณ์ฐํ์ฌ ๋ถ์ํ๊ฒ ์ต๋๋ค.
โ 1. ๊ธฐ์กด Transformer ๋๋น MELODI์ ๊ตฌํ ๋ณ๊ฒฝ์ฌํญ
| ๋ณ๊ฒฝ ํญ๋ชฉ | ์ค๋ช | ๊ตฌํ ๋์ด๋ |
|---|---|---|
| Short-Term Memory (STM) | - ๊ฐ layer๋ง๋ค summary token, token mixer ์ถ๊ฐ- zโโโ๋ฅผ ๋ฐ์์ attention์ ํฌํจ | ์ค |
| Summary Branching | - ๊ฐ layer์์ summary โ ๋ค์ window๋ก ์ ๋ฌ ๊ฒฝ๋ก ์ถ๊ฐ | ์ค |
| Long-Term Memory (LTM) | - ํน์ ์ค๊ฐ layer์์ context window ์์ถ ํ KV ์ ์ฅ - ๋ค์ window์์ cross-attention ์ํ | ์ค~์ |
| Cross-Attn Gating | - self-attn, cross-attn ๊ฒฐ๊ณผ๋ฅผ ฮฑ๋ก ์กฐํฉ (ฮฑ: ํ์ต ๊ฐ๋ฅ scalar per head) | ๋ฎ์ |
| KV ๋ฉ๋ชจ๋ฆฌ ์ ์ฅ ๊ตฌ์กฐ | - FIFO queue ํํ๋ก ์์ถ๋ KV ์ ์ฅ ๋ฐ ์ฐธ์กฐ | ์ค |
| Position Embedding ๋ณ๊ฒฝ | - zโ, uโ์๋ ์์น ์๋ฒ ๋ฉ ์ ์ฉ ํ์ | ๋ฎ์ |
๐ ์์ฝ: ๊ธฐ์กด Transformer์ ๊ตฌ์กฐ๋ฅผ ์ ์งํ๋ฉด์ ์ฝ๊ฐ์ ๋ชจ๋ ์ฝ์ ๋ฐ routing ๊ตฌํ์ด ํ์ํ ์์ค. GPT๋ LLAMA ๊ณ์ด์์๋ ์ถฉ๋ถํ ํ์ฅ ๊ฐ๋ฅํจ.
โ 2. ์ฐ์ฐ๋ ๋ฐ ํ๋ผ๋ฏธํฐ ์ฆ๊ฐ๋ (์์น ๊ธฐ๋ฐ)
๊ธฐ์ค:
- Transformer depth = 13
- Embedding dim = 1024
- Context window = 512 tokens
- Short-term token ์
S = 128 - Summary token ์
U = 128 - Long-term token ์
L = 64 - Long-term memory depth
Q = 128windows
๐ง [A] ์ถ๊ฐ๋๋ ํ๋ผ๋ฏธํฐ ์ (์ด๋์ transformer ํ๋๋น โ 100M ์์ค)
1. Linear token mixer (2๊ฐ per STM layer)
๊ฐ layer๋ง๋ค:
Input: (W + U) x d = (512 + 128) x 1024
Output: S = 128 tokens โ Weight: (640 ร 128) x 2 mixers
์ด = 164,864 params/layerโ 6 STM layer ์๋ค๊ณ ํ๋ฉด:
์ด = 165K ร 6 = **~1M params**2. Gating scalar (ฮฑ per head)
์: 8 heads โ ฮฑ 8๊ฐ (ํ์ต ๊ฐ๋ฅ scalar) โ ๋ฌด์ ๊ฐ๋ฅํ ์์ค (8 ร N layer โ ์๋ฐฑ)
๐ ๊ฒฐ๋ก : ์ ์ฒด์ ์ผ๋ก ์ฝ 1% ๋ฏธ๋ง์ ํ๋ผ๋ฏธํฐ ์ฆ๊ฐ๋ก ์ ํ๋จ
โ๏ธ [B] ์ฐ์ฐ๋ (FLOPs) ์ฆ๊ฐ
1. ์ถ๊ฐ attention ์ ๋ ฅ ์ ์ฆ๊ฐ (STM)
๊ธฐ์กด:
Attention over 512 tokens (self-attn)
โ QK: (512ร1024) ร (512ร1024) = O(512ยฒรd)MELODI (STM):
Attention over [512 + 128] tokens = 640
โ O(640ยฒ ร d) = ์ฝ 56% ์ฆ๊ฐ2. LTM cross-attention
ํ layer์์ 512 tokens๊ฐ 128ร64๊ฐ์ memory KV์ cross-attn ์ํ
long memory ์ด ํฌ๊ธฐ:
PLAINTEXT64 tokens ร 128 windows = 8192 tokens โ attention: 512 ร 8192 ร d = O(4M ร d)64 tokens ร 128 windows = 8192 tokens โ attention: 512 ร 8192 ร d = O(4M ร d)
์ด๋ self-attn์ 512ยฒ = 0.25M ๋ณด๋ค ~16๋ฐฐ ํฌ์ง๋ง, ๋จ 1๊ฐ layer์์๋ง ์ํ โ ์ ์ฒด ์ฐ์ฐ์์ ๋ณด๋ฉด ์ฝ 10~15% ์ฆ๊ฐ
๐ ์ด FLOPs ์ฆ๊ฐ๋ ์ถ์ :
- ์ ์ฒด ๋ชจ๋ธ ๊ธฐ์ค ์ฝ 30% ๋ด์ธ ์ฐ์ฐ๋ ์ฆ๊ฐ
๐พ 3. ๋ฉ๋ชจ๋ฆฌ ์ฌ์ฉ๋
| ํญ๋ชฉ | ๊ธฐ์กด Transformer | MELODI |
|---|---|---|
| ๋จ๊ธฐ ๋ฉ๋ชจ๋ฆฌ (zโ) | ์์ | 128 tokens/layer ร 13 = 1.6M float |
| ์ฅ๊ธฐ ๋ฉ๋ชจ๋ฆฌ (mโ:โ) | ์์ | 64 tokens ร 2 (KV) ร 128 windows = 16.8M float |
| ํฉ๊ณ | - | ์ฝ 18.4M float = ~73.6MB (FP32) |
FP16 ์ฌ์ฉ ์ ~36.8MB ์์ค
๐ทโโ๏ธ 4. ๊ตฌํ ๊ณต์ (์์ง๋์ด๋ง ๊ด์ )
| ๊ตฌ์ฑ | ์์ ์์ | ๋์ด๋ |
|---|---|---|
| STM token mixer | Linear layer ์ถ๊ฐ ๋ฐ branching logic ๊ตฌํ | ์ค |
| Summary token flow | layer ๊ฐ routing + next window๋ก ๋ณต์ฌ | ์ค |
| LM queue | FIFO ๋ฉ๋ชจ๋ฆฌ ๊ด๋ฆฌ ๊ตฌ์กฐ, KV indexing ๊ด๋ฆฌ | ์ค~์ |
| Cross-attn with gating | multihead attention ๋ชจ๋ ์์ | ๋ฎ์ |
| ํฌ์ง์ ์๋ฒ ๋ฉ ํ์ฅ | relative pos embed ์ถ๊ฐ | ๋ฎ์ |
์ด ๊ณต์:
- ๊ธฐ์กด GPT-style ์ฝ๋ ๊ธฐ๋ฐ์ด๋ผ๋ฉด ์ฝ 1~2์ฃผ ๊ฐ๋ฐ ์ธ๋ ฅ 1~2๋ช ์ ๋
- Flax/PyTorch ๋ชจ๋ ๊ตฌํ ๊ฐ๋ฅ, JAX ์ฝ๋๋ก๋ ์ด๋ฏธ ์กด์ฌํจ (๋ ผ๋ฌธ์ Flax ์ฌ์ฉ)
โ ๊ฒฐ๋ก ์์ฝ
| ํญ๋ชฉ | ์์น ๋๋ ํ๊ฐ |
|---|---|
| ํ๋ผ๋ฏธํฐ ์ฆ๊ฐ | ์ฝ 1% (token mixer ๋ฑ) |
| FLOPs ์ฆ๊ฐ | ์ฝ 30% ์ด๋ด |
| ๋ฉ๋ชจ๋ฆฌ ์ฌ์ฉ๋ | ์ฝ 70MB (FP32 ๊ธฐ์ค) |
| ๊ตฌํ ๊ณต์ | ์์ง๋์ด 1~2์ธ ๊ธฐ์ค 1~2์ฃผ ์์ค |
| ์ ์ฉ ๊ฐ๋ฅ์ฑ | ๊ธฐ์กด GPT, LLAMA ๋ฑ transformer ๊ณ์ด์ ์ฝ๊ฒ ํ์ฅ ๊ฐ๋ฅ |
MELODI๋ ํ๋ผ๋ฏธํฐ ์ ์ฆ๊ฐ๋ฅผ ๊ฑฐ์ ์ผ์ผํค์ง ์์ผ๋ฉด์, ๋ฉ๋ชจ๋ฆฌ์ ์ฐ์ฐ์ ์ฌ์ฉ ํจ์จ์ ๊ฐ์ ํ๊ณ , long-context capability๋ฅผ ๋งค์ฐ ํจ๊ณผ์ ์ผ๋ก ๊ฐํํ ์ ์๋ ์ค์ฉ์ ์ด๊ณ ํ์ฅ ๊ฐ๋ฅํ ๊ตฌ์กฐ์ ๋๋ค.
Q : ๊ฐ ์ปดํผ๋ํธ๋ณ๋ก PyTorch-style ๋ชจ๋ ๊ตฌํ์ ๋ง๋ค์ด์ค ์ ์๊ฒ ๋?
# MELODI ํต์ฌ ์ปดํฌ๋ํธ PyTorch-style ๋ชจ๋ ๊ตฌํ
import torch
import torch.nn as nn
import torch.nn.functional as F
from typing import Optional
# -----------------------------
# 1. Summary Token Mixer (Mโ and Mโ)
# -----------------------------
class LinearTokenMixer(nn.Module):
def __init__(self, input_tokens: int, output_tokens: int, dim: int):
super().__init__()
self.mixer = nn.Linear(input_tokens, output_tokens, bias=False)
self.proj = nn.Linear(dim, dim) # to map channel-wise
def forward(self, x): # x: [B, T, D]
x = x.transpose(1, 2) # [B, D, T]
x = self.mixer(x) # [B, D, T']
x = x.transpose(1, 2) # [B, T', D]
return self.proj(x)
# -----------------------------
# 2. Gated Cross-Attention Layer
# -----------------------------
class GatedCrossAttention(nn.Module):
def __init__(self, dim, num_heads):
super().__init__()
self.self_attn = nn.MultiheadAttention(dim, num_heads, batch_first=True)
self.cross_attn = nn.MultiheadAttention(dim, num_heads, batch_first=True)
self.alpha = nn.Parameter(torch.ones(num_heads))
def forward(self, x, long_mem_kv):
self_out, _ = self.self_attn(x, x, x)
cross_out, _ = self.cross_attn(x, long_mem_kv, long_mem_kv)
alpha = self.alpha.view(1, 1, -1)
out = alpha * cross_out + (1 - alpha) * self_out
return out
# -----------------------------
# 3. Short-Term Memory Layer
# -----------------------------
class ShortTermLayer(nn.Module):
def __init__(self, dim, num_heads, token_count=512, summary_count=128):
super().__init__()
self.attn = nn.MultiheadAttention(dim, num_heads, batch_first=True)
self.ffn = nn.Sequential(
nn.Linear(dim, dim * 4),
nn.GELU(),
nn.Linear(dim * 4, dim)
)
self.token_mixer_up = LinearTokenMixer(token_count + summary_count, summary_count, dim)
self.token_mixer_right = LinearTokenMixer(token_count + summary_count, summary_count, dim)
def forward(self, x, summary, prev_memory):
input_seq = torch.cat([prev_memory, x], dim=1)
x_out, _ = self.attn(x, input_seq, input_seq)
x_out = x_out + self.ffn(x_out)
combined = torch.cat([x_out, summary], dim=1)
new_summary = self.token_mixer_up(combined)
new_memory = self.token_mixer_right(combined)
return x_out, new_summary, new_memory
# -----------------------------
# 4. Long-Term Memory Update Layer
# -----------------------------
class LongTermMemoryLayer(nn.Module):
def __init__(self, dim, num_heads, token_count=512, summary_count=128, long_tokens=64):
super().__init__()
self.cross_attn = GatedCrossAttention(dim, num_heads)
self.token_mixer_long = LinearTokenMixer(token_count + summary_count, long_tokens, dim)
def forward(self, x, summary, long_mem):
x = self.cross_attn(x, long_mem)
combined = torch.cat([x, summary], dim=1)
long_tokens = self.token_mixer_long(combined)
return x, long_tokens
# -----------------------------
# 5. FIFO-based Long-Term Memory Bank
# -----------------------------
class LongTermMemoryBank:
def __init__(self, max_windows: int):
self.max_windows = max_windows
self.bank = []
def append(self, kv):
self.bank.append(kv)
if len(self.bank) > self.max_windows:
self.bank.pop(0)
def get_memory(self):
if len(self.bank) == 0:
return None
return torch.cat(self.bank, dim=1)
# -----------------------------
# 6. ์ ์ฒด MELODI ๋ธ๋ก ๋ฐ Inference ๋ฃจํ
# -----------------------------
class MELODIBlock(nn.Module):
def __init__(self, dim=1024, num_heads=8, num_layers=13, sm_layers=[0,1,2,3,5,6,8,10], lm_layer=7):
super().__init__()
self.dim = dim
self.num_layers = num_layers
self.sm_layers = sm_layers
self.lm_layer = lm_layer
self.layers = nn.ModuleList()
for l in range(num_layers):
if l == lm_layer:
self.layers.append(LongTermMemoryLayer(dim, num_heads))
elif l in sm_layers:
self.layers.append(ShortTermLayer(dim, num_heads))
else:
self.layers.append(nn.TransformerEncoderLayer(d_model=dim, nhead=num_heads, batch_first=True))
self.memory_bank = LongTermMemoryBank(max_windows=128)
def forward(self, x_seq_list):
"""
x_seq_list: List of [B, T, D] segments, one per context window
"""
z_prev = torch.zeros(x_seq_list[0].shape[0], 128, self.dim, device=x_seq_list[0].device) # STM ์ด๊ธฐ๊ฐ
u_prev = torch.zeros_like(z_prev) # Summary token ์ด๊ธฐ๊ฐ
out_all = []
for x in x_seq_list:
for l, layer in enumerate(self.layers):
if isinstance(layer, ShortTermLayer):
x, u_prev, z_prev = layer(x, u_prev, z_prev)
elif isinstance(layer, LongTermMemoryLayer):
long_mem = self.memory_bank.get_memory()
x, long_token = layer(x, u_prev, long_mem if long_mem is not None else x)
self.memory_bank.append(long_token)
else:
x = layer(x)
out_all.append(x)
return torch.cat(out_all, dim=1) # ์ ์ฒด output concat์ ๋ ฅ ๋ฐ์ดํฐ์ ์ถ๋ก ๊ณผ์ , ๋ชจ๋ธ ๊ตฌ์กฐ์ ์์ ์๊ตฌ๋
๋ฌผ์ด๋ณธ ๊ฒ: ์ด ๋ ผ๋ฌธ์ ์ ๋ ฅ๋ฐ์ดํฐ์ ์ถ๋ก ๊ณผ์ ์ ๋ํด์ ์์๋ฅผ ๋ค์ด ์์ฃผ ์์ธํ๊ฒ ์ค๋ช ํด ์ฃผ๊ฒ ๋? ์ถ๊ฐ์ ์ผ๋ก ๋ชจ๋ธ์ํคํ ์ฒ์ ๊ตฌ์ฑ ๋ฐ ๋ชจ๋ธ์ ๊ตฌ์ฑํ๋ ์ฐ์ฐ๊ณผ ๋ฉ๋ชจ๋ฆฌ ์๊ตฌ๋ ์ปดํจํ ์๊ตฌ๋ ๋ํด์๋ ์ค๋ช ํด์ค
์์ฃผ ์ข์ต๋๋ค. ์ง๊ธ๊น์ง ์ ์๋ MELODI ์ํคํ ์ฒ๋ฅผ ๋ฐํ์ผ๋ก, ์ด ๊ตฌ์กฐ๊ฐ ์ ๋ ฅ ๋ฐ์ดํฐ๋ฅผ ์ด๋ป๊ฒ ์ฒ๋ฆฌํ๊ณ , ์ถ๋ก ์ด ์ด๋ป๊ฒ ๋์ํ๋ฉฐ, ์ฐ์ฐ๋/๋ฉ๋ชจ๋ฆฌ ์๊ตฌ๊ฐ ์ด๋ค์ง๋ฅผ ๊ตฌ์ฒด์ ์ธ ์์๋ฅผ ๋ค์ด ์ค๋ช ํ๊ฒ ์ต๋๋ค.
โ ์ ๋ ฅ ๋ฐ์ดํฐ ์์ ๋ฐ ์ ์ฒ๋ฆฌ
์ ๋ ฅ ์์ (๊ธด ๋ฌธ์)
"In the middle of the night, he found a strange box hidden beneath the floorboards. ..."
โ ๊ธธ์ด: 8,192 tokens (์: ๊ธด ์์ค)์ ์ฒ๋ฆฌ
# Assume tokenized to shape [B, 8192]
# Split into 512-token windows โ 16๊ฐ window
x_windows = torch.split(input_tensor, 512, dim=1) # ๊ฐ window: [B, 512, D]์ด x_windows๋ MELODIBlock์ ์ ๋ฌ๋ฉ๋๋ค.
๐ง ์ถ๋ก ํ๋ฆ (forward logic ์์)
melodi = MELODIBlock(dim=1024, num_heads=8)
output = melodi(x_seq_list=x_windows)๋ด๋ถ ์ฒ๋ฆฌ ์์
ShortTermLayer๋ ๊ฐ window๋ง๋ค 512-token์ ์ฒ๋ฆฌํ๋ฉด์- ์ด์ window์์ ์ ๋ฌ๋ฐ์
z_{k-1}memory ์ฌ์ฉ - summary token์ ๋ง๋ค์ด ๋ค์ layer/๋ค์ window๋ก ์ ๋ฌ
- ์ด์ window์์ ์ ๋ฌ๋ฐ์
LongTermMemoryLayer(์: 7๋ฒ์งธ layer)์์๋- ์ง๊ธ๊น์ง ์ ์ฅ๋ long-term memory (
mโ:โโโ)์ ๋ํด cross-attention ์ํ - window ์ ์ฒด๋ฅผ ์์ถํด์ 64๊ฐ long-token ์์ฑ โ ๋ฉ๋ชจ๋ฆฌ bank์ ์ถ๊ฐ
- ์ง๊ธ๊น์ง ์ ์ฅ๋ long-term memory (
๋ง์ง๋ง๊น์ง ์ฒ๋ฆฌ๋
x_k๋ ์ถ๋ ฅ์ผ๋ก ์ฌ์ฉ๋จ
๐๏ธ ๋ชจ๋ธ ์ํคํ ์ฒ ๊ตฌ์ฑ
| ๊ตฌ์ฑ ์์ | ์์น (๊ธฐ์ค config) |
|---|---|
| ์ด layers | 13 |
| ShortTermLayer | 8 (e.g., layer 0~3, 5~6, 8, 10) |
| LongTermLayer | 1 (์: layer 7) |
| dim (hidden size) | 1024 |
| attention heads | 8 |
| context window size | 512 tokens |
| summary token ์ | 128 |
| long token ์ | 64 |
| long mem depth | 128 windows |
๐พ ๋ฉ๋ชจ๋ฆฌ ์๊ตฌ๋ ๊ณ์ฐ
1. Short-Term Memory
- zโ: 128 tokens/layer ร 8 layers ร 1024 dim = 1,048,576 float
- = 4.0MB (FP32) or 2.0MB (FP16)
2. Long-Term Memory
- 64 tokens ร 2 (KV) ร 1024 dim ร 128 windows = 16,777,216 float
- = 64.0MB (FP32) or 32.0MB (FP16)
3. Input/Activation buffer (xโ: 16 windows)
- 512 ร 16 ร 1024 = 8,388,608 float = 32MB (FP32)
๐ ์ด ๋ฉ๋ชจ๋ฆฌ (FP16 ๊ธฐ์ค, ์ถ๋ก ์):
โ 2MB (STM) + 32MB (LTM) + 16MB (x buffer) = **50MB ์์ค**โ๏ธ ์ฐ์ฐ๋ (FLOPs ๊ธฐ์ค, ๋จ์ผ window ์ฒ๋ฆฌ ๊ธฐ์ค)
1. ShortTermLayer ร 8 layers
- Self-attn: O((512+128)ยฒ ร 1024) โ 0.4G FLOPs/layer
- FFN: 512 ร 1024 ร 4 ร 1024 ร 2 โ 4.2G FLOPs/layer
- TokenMixer: 2 ร 640 ร 128 ร 1024 โ 0.17G FLOPs/layer
โ ์ด โ ~4.7G ร 8 = 37.6 GFLOPs
2. LongTermLayer (1 layer)
- cross-attn: 512 ร 8192 ร 1024 โ 4.3G FLOPs
- token mixer (LTM): โ 0.08G FLOPs
โ ์ด โ ~4.4 GFLOPs
๐ ์ด ์ถ๋ก ์ฐ์ฐ๋ โ 42 GFLOPs / window โ 16๊ฐ window โ โ 670 GFLOPs (long-context ์ถ๋ก ๊ธฐ์ค)
โ ์์ฝ ์ ๋ฆฌ
| ํญ๋ชฉ | ์์น / ์ค๋ช |
|---|---|
| ์ ๋ ฅ ๊ตฌ์กฐ | [B, 8192, 1024] โ 16 window ๋ถํ |
| ์ฐ์ฐ ๊ตฌ์ฑ | 13-layer Transformer (SM + LM ํผํฉ) |
| ๋ฉ๋ชจ๋ฆฌ ์๊ตฌ๋ (FP16) | ์ฝ 50MB |
| ์ฐ์ฐ๋ (์ถ๋ก ์ ์ฒด) | ์ฝ 670 GFLOPs |
| ํน์ง | - SM์ผ๋ก layer ๊ฐ ์ ๋ณด ํ๋ฆ - LM์ผ๋ก window ๊ฐ ์ ๋ณด ์์ถ ์ ๋ฌ - ๋์ ์ฑ๋ฅ ๋๋น ์ ์ ๋ฉ๋ชจ๋ฆฌ |
ํ๊ณ๋ฅผ ๋์ผ๋ ค๋ ์ฐ๊ตฌ ํ๋ฆ
๋ฌผ์ด๋ณธ ๊ฒ: ์ด ๋ ผ๋ฌธ์ ํ๊ณ๋ฅผ ๊ทน๋ณตํ๊ธฐ ์ํ ๋ฐฉ๋ฒ์ผ๋ก ์ด๋ค ์ฐ๊ตฌํ๋ฆ์ด ์๋์ง ์ ๋ฆฌํด์ ์์ธํ๊ฒ ์ค๋ช ํด ์ค
MELODI๋ ๋ฐ์ด๋ ํจ์จ๊ณผ ์ฑ๋ฅ์ ๋ณด์ฌ์ฃผ๋ ๊ตฌ์กฐ์ด์ง๋ง, ๋ ผ๋ฌธ์์๋ ์ผ๋ถ ํ๊ณ์ ์ ๋ช ์ํ๊ณ ์์ต๋๋ค. ์ฌ๊ธฐ์ ๊ธฐ๋ฐํ์ฌ, ํด๋น ํ๊ณ๋ฅผ ๊ทน๋ณตํ๊ธฐ ์ํ ์ฐ๊ตฌ ํ๋ฆ ๋ฐ ์์ด๋์ด๋ฅผ ๋ค์๊ณผ ๊ฐ์ด ์ ๋ฆฌํด ๋๋ฆฌ๊ฒ ์ต๋๋ค:
โ MELODI์ ํต์ฌ ํ๊ณ
| ๊ตฌ๋ถ | ์ค๋ช |
|---|---|
| 1. Pretrained ๋ชจ๋ธ ์ ์ฉ ๋ถ๊ฐ | MELODI๋ ํ์ต์ ์ฒ์๋ถํฐ ์งํํด์ผ ํ๋ฉฐ, ๊ธฐ์กด ์ฌ์ ํ์ต ๋ชจ๋ธ(GPT ๋ฑ)์ ์ง์ ์ ์ฉ์ด ์ด๋ ต๋ค. |
| 2. Fixed compression ratio | SM/LM์์ ์ฌ์ฉํ๋ token ์๊ฐ ๊ณ ์ ๋์ด ์์ด, ๋ค์ํ ๋ฌธ์ ๊ธธ์ด๋ ๋๋ฉ์ธ์ ์ ์ฐํ๊ฒ ๋์ํ์ง ๋ชปํจ. |
| 3. ์ ๋ณด ์์ค ๊ฐ๋ฅ์ฑ | ์์ฝ ๊ธฐ๋ฐ ๋ฉ๋ชจ๋ฆฌ๋ ์์ถ ๊ณผ์ ์์ ์ค์ ์ ๋ณด๋ฅผ ๋๋ฝํ ์ํ์ด ์๋ค. |
| 4. ์์ฐจ ์ถ๋ก ๊ตฌ์กฐ | window ๊ฐ ์์ฐจ์ ์ถ๋ก ์ด ํ์ํด, parallelism์ด ์ ํ๋๋ค. |
๐ ํ๊ณ ๊ทน๋ณต์ ์ํ ์ฐ๊ตฌ ํ๋ฆ
1. ๐ Pretrained ๋ชจ๋ธ์ MELODI memory ์ฝ์
์ฐ๊ตฌ ํ๋ฆ: ๊ธฐ์กด ์ฌ์ ํ์ต๋ ๋ชจ๋ธ(GPT, LLaMA)์ MELODI-style memory module์ ์ฝ์ ํ๋ ๋ฐฉ์ (plug-and-play)
๋ฐฉ๋ฒ: LoRA, Adapter, QLoRA ๋ฑ๊ณผ ๊ฒฐํฉํด fine-tuning
์์:
- ๐ธMemory-Augmented Fine-tuning
- ๐ธAdapter with Memory Bank
- ๐ธFlash-Memory Injection (streamable memory)
โก๏ธ ์ ์ฉ์ฑ ๊ฐํ + ํ๋ผ๋ฏธํฐ ํจ์จ์ฑ ํ๋ณด
2. ๐ง Adaptive Compression Memory (์์ถ ์ ์ํ)
์ฐ๊ตฌ ํ๋ฆ: token importance score๋ฅผ ๊ธฐ๋ฐ์ผ๋ก ์์ฝ ๋น์จ์ ๋์ ์ผ๋ก ์กฐ์
๊ด๋ จ ์ฐ๊ตฌ:
- ๐ธAutoCompressor (Chevalier et al. 2023): attention-based token compression
- ๐ธICAE (2024): adaptive context encoder with LoRA
๊ฐ๋ฅํ ํ์ฅ:
- ์ ๋ณด๋ ๋ง์ window๋ ๋ ๋ง์ long-token ํ ๋น
- summary token ์๋ฅผ layer depth๋ attention entropy์ ๋ฐ๋ผ ๋ณ๊ฒฝ
โก๏ธ ์ ๋ณด ์์ค ์ต์ํ + ๊ณ ํจ์จ ์์ฝ
3. ๐ Recursive Memory Compression
๊ฐ๋ : LM์ ๋จ์ผ์ธต์ด ์๋ ๋ค์ธต์ผ๋ก ๊ตฌ์ฑํ๊ณ , ์ค๋๋ memory๋ ์ฌ์์ถ (recompress) โ compact KV ์ ์ง
์๊ฐ: Compressive Transformer (Rae et al. 2019)
์ ์ฉ ๋ฐฉ๋ฒ:
- ์ผ์ window ์ด์ ์ค๋๋ memory โ ๋ค์ ํ ๋ฒ token mixer์ ์ ๋ ฅํ์ฌ ์ฌ์์ฝ
- ๊ธฐ์กด FIFO โ multi-tier cache ๊ตฌ์กฐ
โก๏ธ long-context ์ ์ง ๋ฅ๋ ฅ ํฅ์ + memory ํญ์ฆ ์ต์
4. โก ๋น์์ฐจ์ memory ์ ๊ทผ ์ง์
๋ฌธ์ : MELODI๋ window ์์ฐจ ์ฒ๋ฆฌ โ parallelization ์ด๋ ค์
ํด๊ฒฐ ๋ฐฉํฅ:
- long-term memory๋ read-only + cross-attn๋ง ์ํํ๋๋ก ๋ง๋ค๋ฉด ์ฌ๋ฌ window์์ ๋ณ๋ ฌ ์ ๊ทผ ๊ฐ๋ฅ
- memory indexing์ positional RoPE ๊ธฐ๋ฐ์ผ๋ก ์ ๊ทํํ์ฌ ๋น์์ฐจ window์๋ ๋์ ๊ฐ๋ฅ
โก๏ธ ๋ฉํฐ-GPU ๋๋ ๋ถ์ฐ ์ถ๋ก ํจ์จ ๊ทน๋ํ
5. ๐ Multi-modal ํ์ฅ
๋๊ธฐ: MELODI ๊ตฌ์กฐ๋ modality-independent (์์ถ/๋ฉ๋ชจ๋ฆฌ ๋ฐฉ์์ text ์ธ์๋ ์ ์ฉ ๊ฐ๋ฅ)
์ ์ฉ ๊ฐ๋ฅ ์์ญ:
- ์์/์์ฑ์ temporal compression
- multi-modal summary token (์: vision-text)
โก๏ธ GEMINI-style ๋ชจ๋ธ์์ ํจ์จ์ long-memory ๊ตฌ์ฑ ๊ฐ๋ฅ
๐งฉ ์ข ํฉ ์ ๋ฆฌ: ์ฐ๊ตฌ ๋ฐฉํฅ ๋งคํ
| ํ๊ณ | ๊ทน๋ณต ์ฐ๊ตฌ ํ๋ฆ | ๋ํ ์์ด๋์ด |
|---|---|---|
| Pretrained ๋ถ๊ฐ | Memory-finetune, LoRA | memory-injection |
| ๊ณ ์ ์์ถ๋น | Adaptive compression | token importance โ compression ratio |
| ์ ๋ณด ์์ค | Multi-pass summarization | ์ฌ์์ฝ / ์ ๋ณด ์ ํ |
| ์์ฐจ์ ๊ตฌ์กฐ | Memory-as-Cache ๋ฐฉ์ | long-memory ๋ณ๋ ฌ ์ ๊ทผ |
| ๋๋ฉ์ธ ๊ณ ์ | Multi-modal ํ์ฅ | ์์ฝ ๋ฐฉ์ ์ผ๋ฐํ |
Q : MELODI์ short-term memory์ long-term memory๋ ๊ฐ๊ฐ ์ด๋ค ๋ฐฉ์์ผ๋ก ์ ๋ณด๋ฅผ ์์ถํ๊ณ ์ ๋ฌํ๋ฉฐ, ์ด ๊ตฌ์กฐ๊ฐ ๊ธฐ์กด memory-augmented transformer๋ค๊ณผ ๋น๊ตํด ์ด๋ค ์ด์ ์ ์ ๊ณตํ๋๊ฐ?
์์ฃผ ์ข์ ์ง๋ฌธ์ ๋๋ค. ์๋์ MELODI์ short-term memory (STM), long-term memory (LTM)๊ฐ ์ด๋ป๊ฒ ์ ๋ณด๋ฅผ ์์ถํ๊ณ ์ ๋ฌํ๋์ง ๊ทธ๋ฆฌ๊ณ ๊ธฐ์กด memory-augmented transformer์์ ๋น๊ต๋ฅผ ์ ๋ฆฌํด๋๋ฆฝ๋๋ค.
โ MELODI์ Memory ๊ตฌ์กฐ์ ๋์ ๋ฐฉ์
1. Short-Term Memory (STM): Layer-wise recurrent compression
๋ชฉ์ : ํ์ฌ window ๋ด ์ ๋ณด + ์ด์ window ์์ฝ ์ ๋ณด๋ฅผ ์ฒ๋ฆฌ
๊ตฌํ ๋ฐฉ์:
๊ฐ context window
xโ๋ฅผ ์ฌ๋ฌ ShortTermLayer์ ํต๊ณผ์ํค๋ฉฐ ๋ฐ๋ณต ์์ถ๊ฐ layer์์๋:
- context token
xโ, summary tokenuโ, ์ด์ memory tokenzโโโ์ ์ ๋ ฅ์ผ๋ก ์ฌ์ฉ - attention + FFN ํ,
xโโฒ์์ฑ - summary token๊ณผ ํจ๊ป linear token mixer๋ฅผ ํตํด ๋ค์ layer์ฉ
uโ, ๋ค์ window์ฉzโ์์ฑ
- context token
์ ๋ณด ํ๋ฆ:
- layer ๊ฐ:
uโ(summary token) - window ๊ฐ:
zโ(compressed STM token)
- layer ๊ฐ:
๐ ์ด ๊ตฌ์กฐ๋ Transformer ๋ด๋ถ์ LSTM์ฒ๋ผ layer-recurrent ํ๋ฆ์ ๋ง๋ ๋ค๊ณ ๋ณผ ์ ์์
2. Long-Term Memory (LTM): Mid-layer compression + FIFO stacking
๋ชฉ์ : ๊ณผ๊ฑฐ ์ฌ๋ฌ window์ ์ ์ฒด ์์ฝ ์ ๋ณด๋ฅผ ์ฅ๊ธฐ์ ์ผ๋ก ์ ์ง
๊ตฌํ ๋ฐฉ์:
- ์ค๊ฐ layer (์: 7์ธต)์์, ํ์ฌ
xโ,uโ๋ฅผ long-term memory์ cross-attention - self-attn๊ณผ cross-attn์
ฮฑgating์ผ๋ก ํฉ์ฑ - ์ด์ด์ token mixer๋ฅผ ํตํด 512-token window๋ฅผ 64๊ฐ long-token์ผ๋ก ์์ถ
- ์ด KV์์
mโ๋ก ์ ์ฅ, memory queue (mโ:โ)์ append
- ์ค๊ฐ layer (์: 7์ธต)์์, ํ์ฌ
๐ ์ ๋ณด๋ KV ํํ๋ก ์ ์ฅ๋๋ฉฐ, ๋ค์ window ์ฒ๋ฆฌ ์ cross-attention ๋์์ด ๋จ
๐ ๊ธฐ์กด memory-augmented ๋ชจ๋ธ๊ณผ์ ๋น๊ต
| ํญ๋ชฉ | Memorizing Transformer | MELODI |
|---|---|---|
| memory ์ ์ฅ ๋ฐฉ์ | ๋จ์ผ layer์์ ๋ชจ๋ token์ KV ์ ์ฅ | ์ค๊ฐ layer์์ ์์ถ๋ KV ์ ์ฅ |
| ๋จ๊ธฐ ๋ฌธ๋งฅ ์ ์ง | ์์ (์ง์ token attention) | STM ์ฌ์ฉ์ผ๋ก ์ฐ์์ฑ ๋ณด์กด |
| memory ์ฉ๋ | ๋งค์ฐ ํผ (dense KV ์ ์ฅ) | 8~10๋ฐฐ ์ ์ (64-token ์์ค) |
| ์ ๋ณด ํ๋ฆ | ๋จ์ผ ๋ฐฉํฅ (context โ memory) | ๊ณ์ธต์ ํ๋ฆ (SM + LM) |
| attention ๋ฐฉ์ | cross-attn (top-k ๋๋ dense) | gated self + cross-attn |
| ์์ฝ ๋ฐฉ์ | ์์ (raw token ์ ์ฅ) | token mixer ๊ธฐ๋ฐ ์์ฝ/์์ถ |
๐ฏ MELODI์ ๊ตฌ์กฐ์ ์ด์ ์์ฝ
๋ฉ๋ชจ๋ฆฌ ํจ์จ์ฑ ์์ถ๋ memory token (e.g., 64๊ฐ)๋ง ์ ์ฅ โ MT ๋๋น ์ต๋ 8๋ฐฐ ๊ฐ์
์ฅ๋จ๊ธฐ ๋ฌธ๋งฅ ๊ฒฐํฉ STM: ์ง์ window์ smoothํ ์ฐ๊ฒฐ LTM: ์ ์ฒด history ์์ฝ ์ ์ฅ โ local + global ๋ฌธ๋งฅ ๋ชจ๋ ๋ณด์กด
์ฐ์ฐ ์ต์ ํ cross-attn์ ๋จ 1 layer์์๋ง ์ํ, top-k pruning ์์ด๋ ๋ฎ์ cost
ํ์ต๋ memory ์ ์ด (gating) self vs. cross attention์ ฮฑ๋ก ์กฐ์ โ ๋ฉ๋ชจ๋ฆฌ ์์กด๋ ํ์ต ๊ฐ๋ฅ
๊ฒฐ๋ก ์ ์ผ๋ก, MELODI๋ ๊ธฐ์กด memory transformer๋ค์ memory ํญ์ฆ ๋ฌธ์ ๋ฅผ ํด๊ฒฐํ๋ฉด์๋, ์ ๋ณด ๋ณด์กด๊ณผ ์ถ๋ก ์ฑ๋ฅ์ ์๋ฆฝํ ๊ณ์ธต์ ๋ฉ๋ชจ๋ฆฌ ๊ตฌ์กฐ๋ฅผ ์ค๊ณํ ๊ฒ์ด ๊ฐ์ฅ ํฐ ๊ฐ์ ์ ๋๋ค.
๋ ๊ถ๊ธํ์ ์ ์์ผ์ค๊น์? ๋๋ ์ด ๊ตฌ์กฐ๋ฅผ ํ์ฉํ downstream task ์ ์ฉ ๋ฐฉ์์ด ๊ถ๊ธํ์ ๊ฐ์?
Q : ๋ ผ๋ฌธ์์ ์ ์ํ๋ hierarchical memory compression ๋ฐฉ์์ด long-context language modeling์ ์ฑ๋ฅ ํฅ์์ ์ด๋ค ๊ธฐ์ฌ๋ฅผ ํ๋์ง, ablation ๊ฒฐ๊ณผ๋ฅผ ํตํด ์ด๋ป๊ฒ ๊ฒ์ฆ๋์๋๊ฐ?
๋ ผ๋ฌธ์์ ์ ์ํ๋ Hierarchical Memory Compression์ MELODI์ ํต์ฌ ๊ธฐ์ฌ๋ก, short-term memory (STM)๊ณผ long-term memory (LTM)๋ฅผ ๊ณ์ธต์ ์ผ๋ก ๊ฒฐํฉํจ์ผ๋ก์จ long-context language modeling์ ์ฑ๋ฅ์ ๋์์ต๋๋ค. ์ด ๊ตฌ์กฐ๊ฐ ์ค์ ๋ก ์ด๋ป๊ฒ ์ฑ๋ฅ ํฅ์์ ๊ธฐ์ฌํ๋์ง๋ Ablation Study๋ฅผ ํตํด ๋ช ํํ ๊ฒ์ฆ๋์์ต๋๋ค.
์๋์ ๊ตฌ์กฐ์ ์ดํด์ ablation ๊ฒฐ๊ณผ ๊ธฐ๋ฐ์ ๋ถ์์ ์ ๋ฆฌํด๋๋ฆฝ๋๋ค.
โ Hierarchical Memory Compression์ด๋?
Short-Term Memory (STM):
- context window ๋ด ์ ๋ณด๋ฅผ ์ฌ๋ฌ layer๋ฅผ ํตํด ๋ฐ๋ณต์ ์ผ๋ก ์์ถ
- ์ด์ window์ memory token
z_{k-1}์ summary tokenu_{k-1}ํ์ฉ - ๋ก์ปฌ ๋ฌธ๋งฅ ์ ์ง์ ํจ๊ณผ์
Long-Term Memory (LTM):
- ํ ์ค๊ฐ layer์์ context window ์ ์ฒด๋ฅผ ์์ฝ โ 64๊ฐ token์ผ๋ก ์์ถ
- ๊ณผ๊ฑฐ window๋ค์ ์์ถ๋ KV๋ฅผ FIFO queue์ ์ ์ฅ
- ์ ์ญ ๋ฌธ๋งฅ ์ ์ง์ ํจ๊ณผ์
์์ฝ: โ STM์ ์ต๊ทผ ๋ฌธ๋งฅ์ ์ธ๋ฐํ๊ฒ ๋ณด์กด, LTM์ ๋จผ ๊ณผ๊ฑฐ๋ฅผ ์์ฝํด ๊ธฐ์ต โ ์ด ๋์ ๊ณ์ธต์ ์ผ๋ก ๊ฒฐํฉํ์ฌ short+long dependency ๋์ ์ฒ๋ฆฌ
๐งช Ablation ์คํ์ผ๋ก ํ์ธ๋ ๊ธฐ์ฌ
1. STM + LTM ์กฐํฉ์ ์ฑ๋ฅ ํฅ์ (Fig. 4)
์คํ ์ค์ : PG-19 (T5 vocab) ๊ธฐ์ค
์กฐ๊ฑด:
- ๋ค์ํ short memory (
S) - ๋ค์ํ long memory (
L) ํฌ๊ธฐ - ์ด perplexity ๋น๊ต
- ๋ค์ํ short memory (
๊ฒฐ๊ณผ ์์ฝ:
| ๊ตฌ์กฐ | Perplexity (PG-19) |
|---|---|
| STM only (S192+L0) | 11.0+ (๋์) |
| LTM only (S0+L64) | 11.2+ (๋์) |
| STM + LTM (S128+L64) | 10.44 (์ต์ ) |
โก๏ธ STM๊ณผ LTM์ ์ํธ๋ณด์์ ์ด๋ฉฐ, ๋์ ํจ๊ป ์จ์ผ ์ฑ๋ฅ ์ต์ ํ
2. LTM coverage๊ฐ ์ฑ๋ฅ์ ๋ฏธ์น๋ ์ํฅ (Fig. 5)
- ๊ณ ์ ๋ L = 64 long-token, S = 128 short-token
- LTM์ด ์ปค๋ฒํ๋ window ์: 2 โ 128๊น์ง ์คํ
๊ฒฐ๊ณผ ์์ฝ:
- 2~4 window๋ง ํฌํจํ LTM์ ์ฑ๋ฅ ๊ฑฐ์ ๋ณํ ์์
- 32 window ์ด์ ํฌํจํ ๊ฒฝ์ฐ๋ถํฐ PPL ๊ธ๊ฒฉํ ๊ฐ์
- 128 window ์ด์์์๋ ์ฑ๋ฅ ๊ฐ์ ์ ์ฒด
โก๏ธ STM์ ์ต๊ทผ ๋ช window๊น์ง๋ง ํจ๊ณผ์ , ๋ฉ์ด์ง ๋ฌธ๋งฅ์ LTM์ด ํ์
3. Summary Branching ๊ธฐ๋ฒ์ ์ํฅ (Table 4)
| ๊ตฌ์กฐ | PPL (ST only) | PPL (ST+LT) |
|---|---|---|
| without branching | 11.68 | 11.24 |
| with branching | 11.39 | 10.95 |
โก๏ธ layer ๊ฐ summary token ํ๋ฆ์ ๊ฐํํ ๊ฒ์ด ์ฑ๋ฅ ํฅ์์ ์ง์ ๊ธฐ์ฌํจ์ ๋ณด์ฌ์ค
๐ ์ ๋ฆฌ: ์ฑ๋ฅ ํฅ์์ ๊ธฐ์ฌํ ํต์ฌ ํฌ์ธํธ
| ๊ธฐ๋ฒ | ๊ธฐ์ฌ |
|---|---|
| STM (multi-layer compression) | ์ต๊ทผ ๋ฌธ๋งฅ ์ ๋ณด ์ธ๋ฐํ๊ฒ ์ ์ง โ short-range dependency ๊ฐํ |
| LTM (window-level compression) | ์ค๋๋ ๋ฌธ๋งฅ์ ๊ฐ๊ฒฐํ๊ฒ ๋ณด์กด โ long-range dependency ์ ์ง |
| Summary token + branching | ์ ๋ณด ํ๋ฆ ๊ฐ์ โ ์์ถ ํจ์จ ๋ฐ ๊ณ์ธต๊ฐ ์ ๋ณด ์ ๋ฌ๋ ฅ ์ฆ๊ฐ |
| Gated cross-attn | LTM ์์กด๋ ์กฐ์ ๊ฐ๋ฅ โ ์ ์ ํ memory usage ํ์ต ๊ฐ๋ฅ |
๐ฏ ๊ฒฐ๋ก
Ablation ๊ฒฐ๊ณผ๋ MELODI์ hierarchical memory ๊ตฌ์กฐ๊ฐ ๋จ์ memory stacking ๋ฐฉ์๋ณด๋ค ํจ์ฌ ํจ์จ์ ์ด๋ฉฐ, STM๊ณผ LTM์ ๊ณ์ธต์ ๊ฒฐํฉ์ด long-context ์ดํด ๋ฅ๋ ฅ ํฅ์์ ํต์ฌ์ ์ธ ๊ธฐ์ฌ๋ฅผ ํ๋ค๋ ๊ฒ์ ์ค์ฆ์ ์ผ๋ก ๋ณด์ฌ์ค๋๋ค.
Q : MELODI ๊ตฌ์กฐ๊ฐ ๊ฐ๋ ํ๊ณ(์: pretrained ๋ชจ๋ธ ์ ์ฉ ์ด๋ ค์, ์์ฐจ์ ์ฒ๋ฆฌ ๋ฑ)๋ ์ค์ ์์ฉ์์ ์ด๋ค ์ ์ฝ์ ๋ง๋ค๋ฉฐ, ์ด๋ฅผ ๋ณด์ํ๊ธฐ ์ํ ๊ธฐ์ ์ ํ์ฅ ๊ฐ๋ฅ์ฑ์ ์ด๋ค ๊ฒ์ด ์์๊น?
MELODI๋ ํจ์จ์ ์ด๊ณ ๊ฐ๋ ฅํ long-context ์ฒ๋ฆฌ ๋ฅ๋ ฅ์ ๋ณด์ฌ์ฃผ์ง๋ง, ๋ช ๊ฐ์ง ๊ตฌ์กฐ์ ํ๊ณ๊ฐ ์ค์ ์์ฉ์ ์ ์ฝ์ ์ค ์ ์์ต๋๋ค. ์๋์ MELODI์ ๊ตฌ์กฐ์ ํ๊ณ, ์์ฉ์์ ์ ์ฝ, ๊ทธ๋ฆฌ๊ณ ์ด๋ฅผ ๋ณด์ํ ์ ์๋ ๊ธฐ์ ์ ํ์ฅ ๊ฐ๋ฅ์ฑ์ ์ ๋ฆฌํด๋๋ฆฝ๋๋ค.
โ MELODI ๊ตฌ์กฐ์ ์ฃผ์ ํ๊ณ์ ์ค์ ์ ์ฝ
1. ์ฌ์ ํ์ต(pretrained) ๋ชจ๋ธ ์ ์ฉ ์ด๋ ค์
๋ฌธ์ : MELODI๋ Transformer ๊ตฌ์กฐ๋ฅผ ๋ฐ๊พผ ๊ตฌ์กฐ์ด๊ธฐ ๋๋ฌธ์ ๊ธฐ์กด GPT, LLaMA ๋ฑ์ pretrained weight๋ฅผ ์ง์ ์ฌ์ฉํ ์ ์์
์ค์ ์ ์ฝ:
- ๊ธฐ์กด ๋๊ท๋ชจ ์ฌ์ ํ์ต ์์์ ํ์ฉํ ์ ์์ด from scratch training ํ์
- ๋น์ฉ, ๋ฐ์ดํฐ ํ๋ณด, ์ฑ๋ฅ ์ฌํ ์ธก๋ฉด์์ ํ์ค์ ์ธ ์ฅ๋ฒฝ ์กด์ฌ
2. ์์ฐจ์ ์ฒ๋ฆฌ (window-by-window)
๋ฌธ์ : context window๋ฅผ ์์๋๋ก ์ฒ๋ฆฌํ๋ฉด์ memory๋ฅผ ๊ฐฑ์ ํ๋ ๊ตฌ์กฐ โ ๋ณ๋ ฌ์ฑ ์ ํ
์ค์ ์ ์ฝ:
- batch-level parallelism ๋ถ๊ฐ โ ์ถ๋ก latency ์ฆ๊ฐ
- GPU ๋ค์ค์ฒ๋ฆฌ๋ ๋ถ์ฐ์ถ๋ก ์ ๋ถ๋ฆฌ โ inference throughput ๋ฎ์
3. ๋ฉ๋ชจ๋ฆฌ ์์ถ์ ์ ๋ณด ์์ค ๊ฐ๋ฅ์ฑ
๋ฌธ์ : long-term memory๋ summary token์ ํตํด ์์ถ ์ ์ฅ
- โ ๋ชจ๋ ์ค์ํ ์ ๋ณด๊ฐ ๋ณด์กด๋๋ค๋ ๋ณด์ฅ์ ์์
์ค์ ์ ์ฝ:
- ์ผ๋ถ downstream task (์: QA, reasoning)์์๋ ์น๋ช ์ ์ ๋ณด ์ ์ค ๊ฐ๋ฅ
- ํน์ window์ ํต์ฌ ๋ด์ฉ์ด ์ถ๋ก ์ ๋๋ฝ๋ ์ํ
๐ง ๊ธฐ์ ์ ํ์ฅ ๊ฐ๋ฅ์ฑ ๋ฐ ๋ณด์ ๋ฐฉ์
1. Pretrained ๋ชจ๋ธ๊ณผ์ ํธํ์ ์ํ Adapter-based ์ฝ์
์ ๊ทผ๋ฒ: ๊ธฐ์กด GPT ๋ฑ์ ์ฌ์ ํ์ต ๋ชจ๋ธ์ MELODI memory block์ LoRA, Adapter ํํ๋ก ์ฝ์
์์ ์์ด๋์ด:
MELODI-Attentionโ ๊ธฐ์กด self-attn ํ cross-attn to memory ์ถ๊ฐ- ๊ธฐ์กด weight freezing + memory block๋ง ํ์ต
์ฅ์ :
- ์ฌ์ ํ์ต weight ํ์ฉ ๊ฐ๋ฅ
- few-shot tuning ๊ฐ๋ฅ
2. ๋น์์ฐจ์ memory ์ ๊ทผ ๊ตฌ์กฐ๋ก parallelism ํ๋ณด
์ ๊ทผ๋ฒ:
- Long-term memory๋ฅผ KV cache์ฒ๋ผ read-only ๊ตฌ์กฐ๋ก ๊ตฌ์ฑ
- ์ฌ๋ฌ window๋ฅผ ๋ณ๋ ฌ์ ์ผ๋ก long memory์ cross-attn ์ํ ๊ฐ๋ฅํ๊ฒ ๋ณ๊ฒฝ
์์:
- โFlashMemory-styleโ ์ ๊ทผ
- Routing index์ positional RoPE ๋ณด์ ํ์ฉ
์ฅ์ :
- window ๊ฐ ๋ณ๋ ฌ ์ฒ๋ฆฌ ๊ฐ๋ฅ โ GPU ํจ์จ ๊ฐ์
3. ์ ๋ณด ์์ค ์ต์ํ๋ฅผ ์ํ Adaptive Memory Compression
์ ๊ทผ๋ฒ:
- token importance ๊ธฐ๋ฐ์ผ๋ก summary token ๋น์จ์ ๋์ ์ผ๋ก ์กฐ์
- ๋๋ retrieval ๊ธฐ๋ฐ memory reinforcement (RAG hybrid)
๊ธฐ์ ํ๋ฆ ์์:
- AutoCompressor (Chevalier et al. 2023)
- Gisting (Mu et al. 2024)
์ฅ์ :
- ์์ถ ์์ค์ ์ต์ํ โ ๊ณ ์ summary token ๊ตฌ์กฐ๋ณด๋ค ์ ๋ฐ
4. Memory ์ฌํ์ฉ ๊ตฌ์กฐ ์ถ๊ฐ (Compressive Memory)
์ ๊ทผ๋ฒ:
- ์ค๋๋ long-term memory๋ฅผ ์ฌ์์ฝํ์ฌ compactํ๊ฒ ๋ณด๊ด
- Compressive Transformer์ ์ ์ฌํ multi-tier memory stack ๊ตฌ์ฑ
์ฅ์ :
- memory ์ฉ๋ ์ ํ ๊ทน๋ณต + ์ ๋ณด ์ ์ง
5. Multi-modal ํ์ฅ์ฑ ํ๋ณด
์ ๊ทผ๋ฒ:
- MELODI์ memory block์ vision/audio encoder ๋ค์ ์ฐ๊ฒฐ
- multi-modal token๋ memory token์ฒ๋ผ ์ฒ๋ฆฌ
์ฅ์ :
- Gemini, Flamingo ๋ฑ multi-modal long context ๋ชจ๋ธ๋ก ํ์ฅ ๊ฐ๋ฅ
๐ ๊ฒฐ๋ก ์์ฝ
| ํ๊ณ | ์์ฉ ์ ์ฝ | ๋ณด์ ๋ฐฉ์ |
|---|---|---|
| pretrained ์ ์ฉ ์ด๋ ค์ | ๊ธฐ์กด weight ์ฌ์ฌ์ฉ ๋ถ๊ฐ | memory block โ adapter/LoRAํ |
| ์์ฐจ ์ฒ๋ฆฌ ๊ตฌ์กฐ | GPU ๋ณ๋ ฌ์ฑ ์ ํ | LTM์ read-only cross-attn์ผ๋ก ๊ตฌ์กฐํ |
| memory ์ ๋ณด ์์ค | ์ถ๋ก ์ ๋ณด ์ ์ค ๊ฐ๋ฅ์ฑ | token ์ค์๋ ๊ธฐ๋ฐ adaptive compression |
| memory ํญ์ฆ | ๊ธด ๋ฌธ์์์ ๋ฉ๋ชจ๋ฆฌ ํ๊ณ | compressive memory or hybrid routing |
| modality ํ์ | text ์ ์ฉ ๊ตฌ์กฐ | multi-modal memory ํ์ฅ ๊ฐ๋ฅ |
๋๊ธ