๊ฐ์ ๊ณผ ๋ ์ฐฝ์ฑ, ํต์ฌ ์๊ณ ๋ฆฌ์ฆ, ๊ทธ๋ฆฌ๊ณ ํ๊ณ
๋ฌผ์ด๋ณธ ๊ฒ: ์ด ๋ ผ๋ฌธ์ ์์ฃผ ์์ธํ๊ฒ ์ฝ๊ณ ๋ ผ๋ฌธ์ ๊ฐ์ ๊ณผ ๋ ์ฐฝ์ ์ธ ์ง์ ์ ์ค๋ช ํด์ฃผ๊ณ ํต์ฌ ์๊ณ ๋ฆฌ์ฆ์ ์์ ์ ๋ ฅ์ ๋ค์ด์ ์ ์ฒด์ ์ธ ๊ณผ์ ์ ์ค๋ช ํด์ค ์ถ๊ฐ์ ์ผ๋ก ๋ ผ๋ฌธ์ ํ๊ณ์ ์ ๋ํด์๋ ์๋ ค์ค
๐ ๊ฒฐ๋ก ์์ฝ (ํต์ฌ ๊ธฐ์ฌ ๋ฐ ์ฑ๋ฅ)
MegaScale-Infer๋ ๋๊ท๋ชจ Mixture-of-Experts (MoE) ๋ชจ๋ธ ์๋น์ ์ํ ํจ์จ์ ์์คํ ์ผ๋ก, Attention๊ณผ FFN ๋ชจ๋์ ๋ถ๋ฆฌ(disaggregate) ํ์ฌ GPU ํ์ฉ๋ฅ ์ ๊ทน๋ํํ๊ณ ์ต๋ 1.9ร์ GPU throughput ๊ฐ์ ๋ฐ 1.86ร ๋น์ฉ ๋๋น ์ฑ๋ฅ ํฅ์์ ๋ฌ์ฑํฉ๋๋ค.
โ ๋ ผ๋ฌธ์ ๊ฐ์ ๊ณผ ๋ ์ฐฝ์ ์ธ ๊ธฐ์ฌ
| ๊ตฌ๋ถ | ๋ด์ฉ |
|---|---|
| ํต์ฌ ๊ธฐ์ฌ | Attention๊ณผ FFN์ ๋ถ๋ฆฌํ์ฌ ๋ ๋ฆฝ์ ์ธ ๋ณ๋ ฌ ์ ๋ต ์ ์ฉ |
| ์ฑ๋ฅ ์ต์ ํ | Ping-Pong Pipeline + M2N ํต์ ๊ตฌ์กฐ๋ก ๊ณ์ฐ/ํต์ ์ค๋ฒ๋ฉ |
| ํ๋์จ์ด ์ ์์ฑ | ์ด๊ธฐ์ข (Heterogeneous) GPU ํ๊ฒฝ์ ์ต์ ํ๋ ๋ฐฐ์น ์ ๋ต ์ง์ |
| ํต์ ์ต์ ํ | NCCL ๋๋น ์ต๋ 96.2% latency ๊ฐ์, 4.2ร throughput ์ฆ๊ฐ |
| ์ด์ ํจ์จ์ฑ | ์์คํ ์์ค ๋ฐฐ์น ๊ณํ ์ต์ ํ (GPU ์, ๋ณ๋ ฌ๋, micro-batch ์ ๋ฑ ํฌํจ) |
โ๏ธ ํต์ฌ ์๊ณ ๋ฆฌ์ฆ ๋ฐ ์์ ์ ๋ ฅ ๊ธฐ๋ฐ ๋์ ๊ณผ์
์์ ์ค์
- ๋ชจ๋ธ: Mixtral 8ร22B
- GPU: A100 80GB
- Batch size: 156
- top-k experts: 2, ์ด expert ์: 8 โ ๊ฐ expert๋น ํ๊ท 39๊ฐ์ ํ ํฐ๋ง ์ฒ๋ฆฌ
๋ฌธ์ ์ : FFN์ compute-intensive์ธ๋ฐ, MoE sparsity๋ก ์ธํด ๋ฐฐ์น ํฌ๊ธฐ๊ฐ ์์์ ธ GPU ํ์ฉ๋ฅ โ
MegaScale-Infer ๋์ ํ๋ฆ
1. Attention/FFN ๋ถ๋ฆฌ (Disaggregated Expert Parallelism)
- ๊ฐ layer์์ Attention์ A GPU ๊ทธ๋ฃน, FFN์ E GPU ๊ทธ๋ฃน์ ๋ฐฐ์น
- Attention โ FFN โ Attention ๊ฐ ํต์ ํ์
2. Ping-Pong Pipeline Parallelism
- ์ ์ฒด ๋ฐฐ์น๋ฅผ micro-batch๋ก ๋๋ (์: m=4)
- ๊ฐ micro-batch๋ Attention โ FFN ์์ผ๋ก ์ฒ๋ฆฌ๋๋ฉฐ, ๊ฐ ๋จ๊ณ์์ pipeline์ด ์ค๋ฒ๋ฉ๋จ
$\text{Condition 1:}\ T_a \approx T_e \quad \text{(๊ณ์ฐ์๊ฐ ์ ์ฌ)}$
$\text{Condition 2:}\ T_c < T_f \quad \text{(ํต์ ์๊ฐ < ๊ณ์ฐ์๊ฐ)}$
$m \geq 2\left(1+\frac{T_c}{T_f}\right) \quad \text{(micro-batch์ ๊ฐ์ ์กฐ๊ฑด)}$
3. M2N ํต์ ์ต์ ํ (Attention M๊ฐ โ Expert N๊ฐ)
- ๊ธฐ์กด NCCL์ All2All์ ์ต์ ํ๋์ด MoE token routing์๋ ๋ถ์ ํฉ
- ์๋ก ๊ตฌํํ M2N ๋ผ์ด๋ธ๋ฌ๋ฆฌ๋:
- GPU-to-CPU copy ์ ๊ฑฐ
- GPU sync ์ ๊ฑฐ
- RDMA + GPUDirect ํ์ฉ
- ACK ์ฐ์ ์ ์ก + ํผ์ก์ ์ด ์ต์ ํ
๐ ์ฑ๋ฅ ๋น๊ต (vLLM, TensorRT-LLM ๋๋น)
| ๋ชจ๋ธ | MegaScale-Infer vs. vLLM | MegaScale-Infer vs. TensorRT-LLM |
|---|---|---|
| Mixtral 8x22B | 2.56ร โ | 1.28ร โ |
| DBRX | 1.70ร โ | 1.30ร โ |
| Scaled-MoE (317B) | 7.11ร โ | 1.90ร โ |
Heterogeneous Deployment (H20: Attention / L40S: Expert)์์๋ ์ต๋ 3.24ร throughput/cost ํฅ์ ๊ด์ธก
๐งฉ ํ๊ณ์ ๋ฐ ๊ฐ์ ๊ฐ๋ฅ์ฑ
| ํ๊ณ | ์ค๋ช |
|---|---|
| Pipeline ์ ์ฝ | Ta โ Te ์กฐ๊ฑด์ด ๋ง์ง ์์ผ๋ฉด ping-pong ํ์ดํ๋ผ์ธ์ ํจ์จ์ฑ ์ ํ |
| ํต์ ๋ณ๋ชฉ | M2N ํต์ ์ต์ ํ์๋ ๋ถ๊ตฌํ๊ณ ๋๊ท๋ชจ ๋ ธ๋ ์์์๋ ์ฌ์ ํ ๋ณ๋ชฉ ๊ฐ๋ฅ |
| Load imbalance | Expert ๊ฐ ์ ๋ฆผ ๋ฌธ์ ํด๊ฒฐ์ ์ํด redundant expert ๋ฐฐ์น ํ์ (๋ณต์ก์ฑโ) |
| ๋ชจ๋ธ ์ ์ฐ์ฑ | attention๊ณผ FFN ๊ตฌ์กฐ๊ฐ ํฌ๊ฒ ๋ณํ ๋ชจ๋ธ(ex. Perceiver ๊ณ์ด)์๋ ์ ์ฉ ์ด๋ ค์ |
โณ๏ธ ์์ฝ ์ ๋ฆฌ
| ํญ๋ชฉ | ๋ด์ฉ |
|---|---|
| ์ฃผ์ | ๋๊ท๋ชจ MoE ๋ชจ๋ธ ์๋น ์ต์ ํ |
| ํต์ฌ ์์ด๋์ด | Attention๊ณผ FFN ๋ถ๋ฆฌ + ํ์ดํ๋ผ์ธ + M2N ํต์ ์ต์ ํ |
| ์ฑ๋ฅ | ๊ธฐ์กด ๋๋น ์ต๋ 1.9ร throughput, 1.86ร cost-efficiency ํฅ์ |
| ์๊ณ ๋ฆฌ์ฆ | micro-batching ๊ธฐ๋ฐ ping-pong pipeline + ์ต์ ๋ฐฐ์น ๊ณํ ํ์ ์๊ณ ๋ฆฌ์ฆ |
| ํ๊ณ | load imbalance, ํต์ ๋ณ๋ชฉ, pipeline ๊ท ํ ์๊ตฌ |
Q : ์๊ณ ๋ฆฌ์ฆ 1 (๋ฐฐ์น ์ต์ ํ ํ์)์ด๋ M2N ํต์ ๊ตฌ์กฐ์ ๋ํ ๋ ๊ตฌ์ฒด์ ์ธ ๋ถ์์ ํด์ฃผ๊ฒ ๋?
์ข์ต๋๋ค. ์์ฒญํ์ ๋ ๊ฐ์ง ์ค ๋จผ์ Algorithm 1 (๋ฐฐ์น ์ต์ ํ ํ์ ์๊ณ ๋ฆฌ์ฆ)๊ณผ ๊ทธ ๊ตฌ์กฐ๋ฅผ ๋จผ์ ์ฌ๋ ์๊ฒ ๋ถ์ํ๊ณ , ์ด์ด์ M2N ํต์ ๊ตฌ์กฐ์ ์ค๊ณ์ ๋ณ๋ชฉ ์ ๊ฑฐ ๋ฐฉ์์ ๋ค๋ฃจ๊ฒ ์ต๋๋ค.
โ Algorithm 1: MegaScale-Infer์ ๋ฐฐ์น ์ต์ ํ ํ์ (Deployment Plan Search)
๐ฏ ๋ชฉํ
์ฃผ์ด์ง ํ๋์จ์ด ํ๊ฒฝ(CPU, GPU, ๋ฉ๋ชจ๋ฆฌ)๊ณผ ๋ชจ๋ธ ์ค์ (MoE ๊ตฌ์กฐ, latency ์๊ตฌ ์กฐ๊ฑด) ํ์์,
- ping-pong pipeline ๋ณ๋ ฌ์ฑ
- tensor parallelism ์์ค
- attention/FFN ๊ฐ ๋ฐธ๋ฐ์ค
- GPU ๋ฉ๋ชจ๋ฆฌ ์ ์ฝ ๋ฑ์ ๋ง์กฑํ๋ฉด์ cost-efficiency (Throughput per Dollar)๊ฐ ์ต๋ํ๋๋ ๋ฐฐ์น ๊ณํ(plan)์ ํ์.
๐ฃ ์ฃผ์ ํ๋ผ๋ฏธํฐ (from Table 1)
| ๊ธฐํธ | ์๋ฏธ |
|---|---|
| \( tpa, tpe \) | attention, expert์ ํ ๋นํ tensor parallelism ์์ค |
| \( Ca, Ce \) | attention, expert ๋ ธ๋์ GPU ๋ฉ๋ชจ๋ฆฌ ์ฉ๋ |
| \( Pa, Pe \) | attention, expert ํ๋์ weight ํ๋ผ๋ฏธํฐ ํฌ๊ธฐ |
| \( m \) | micro-batch ๊ฐ์ |
| \( B \) | global batch size |
| \( tpd \) | throughput per dollar (์ต์ ํ ๋์) |
๐ ์๊ณ ๋ฆฌ์ฆ ๊ตฌ์กฐ ์์ฝ
for tpe in [1, 2, ..., Me]: # expert์ TP ์์ค ๋ฐ๋ณต
for tpa in [1, 2, ..., Ma]: # attention์ TP ์์ค ๋ฐ๋ณต
if GPU memory ์ ์ฝ ๋ง์กฑ:
na = balance(G, tpa, tpe) # attention node ๊ฐ์ ๊ณ์ฐ (Ta โ Te ๋ง์กฑ)
for m in [3, 4, ..., Nm]: # micro-batch ์
plan = (tpe, E), (tpa, na), m
B, tpd = simulate(plan, SLO)
if plan.tpd > current_best:
plan* = plan๐ ํต์ฌ ๋
ผ๋ฆฌ: balance(G, tpa, tpe)
- ๋ชฉํ: attention๊ณผ expert์ forward ์ฐ์ฐ ์๊ฐ ์ผ์น (Ta โ Te)
- ์์ ๊ธฐ๋ฐ์ผ๋ก attention node ์ \(n_a\) ๊ณ์ฐ:
์ฌ๊ธฐ์
- \(k_1\): attention micro-batch ์ฒ๋ฆฌ ์๊ฐ์ ๊ณ์ (ํ๋กํ์ผ๋ง ๊ธฐ๋ฐ)
- \(k_3\): expert micro-batch ์ฒ๋ฆฌ ์๊ฐ์ ๊ณ์
- \(K\): top-k expert ์ (๋ณดํต 2 ๋๋ 4)
- \(E\): ์ ์ฒด expert ์
์ด ์์์ ์ค์ ์คํ ๊ธฐ๋ฐ \(k_i\) ๊ณ์๋ฅผ ์ ๋ ฅํ์ฌ attention๊ณผ expert ๊ฐ์ pipeline ๊ท ํ์ ๋ง์ถ๊ธฐ ์ํ ๊ฒ.
๐ ์ ์ฝ ์กฐ๊ฑด ์ ๋ฆฌ
| ์กฐ๊ฑด | ์ค๋ช |
|---|---|
| \( T_a \approx T_e \) | attention, expert compute ์๊ฐ ์ ์ฌํด์ผ pipeline์ idle ์๊ฒ |
| \( T_c < T_f \) | ํต์ ์๊ฐ๋ณด๋ค ๊ณ์ฐ ์๊ฐ์ด ๋ ์ปค์ผ communication hiding ๊ฐ๋ฅ |
| \( m \ge 2(1 + \frac{T_c}{T_f}) \) | pipeline ์์ ํ์ฑํ๋ฅผ ์ํ ์ต์ micro-batch ์ |
| \( T_{iter} \le SLO \) | ์ ์ฒด iteration latency๊ฐ latency SLO ๋ง์กฑํด์ผ ํจ |
| \( \text{Memory usage} < \text{GPU capacity} \) | KV cache ๋ฐ parameter size ๊ณ ๋ คํ ๋ฉ๋ชจ๋ฆฌ ์ ์ฝ |
๐ SIMULATE ํจ์: Throughput per Dollar ํ๊ฐ
- ๊ฐ plan์ ๋ํด latency ์ธก์ (์ 5: \(T_{total}\)), throughput ๊ณ์ฐ (\(B / T_{total}\))
- ๋น์ฉ์ GPU ์ * ๋จ๊ฐ๋ก ๊ณ์ฐ
- ์ต์ข objective:
์ค์ ๋ฐฐ์น ๊ณํ์ exhaustive search + profiling ๊ธฐ๋ฐ ์ถ๋ก ์ ํผํฉํ ํ์ด๋ธ๋ฆฌ๋ ๋ฐฉ์์ผ๋ก ๊ตฌํ๋จ
๐ก M2N ํต์ ๊ตฌ์กฐ ๊ณ ๊ธ ๋ถ์
๐ฉ ๋ฌธ์ : ๊ธฐ์กด NCCL์ ํ๊ณ
| ๋ฌธ์ ์ | ์ค๋ช |
|---|---|
| ๋ถํ์ํ GPUโCPU ๋ณต์ฌ | NCCL์ proxy๋ฅผ ํตํด ํต์ ์ copy ๋ฐ์ |
| Group operation ์ ํ | 8๊ฐ ๋จ์๋ก ์ฒ๋ฆฌ๋์ด ๋ง์ receiver์ผ ๋ ์ฑ๋ฅ ์ ํ |
| Latency instability | high percentile latency (P99) ๋งค์ฐ ๋์ |
| Setup overhead | ์ผ๋ฐ ๋ชฉ์ ์งํฉ ์ฐ์ฐ์ ์ํ ๋ถํ์ํ ์ด๊ธฐํ ํฌํจ |
โ MegaScale-Infer์ ํด๊ฒฐ์ฑ : Custom M2N Library
๐ง ๋์์ธ ํน์ง
| ๊ตฌ์ฑ ์์ | ์ค๋ช |
|---|---|
| Core Sender | CPU ๊ธฐ๋ฐ RDMA write ์ฌ์ฉ + GPUDirect๋ก GPU ๋ฉ๋ชจ๋ฆฌ ์ง์ ์ก์ |
| Send-control Kernel | stream-oriented trigger ๊ธฐ๋ฐ์ผ๋ก send flag ์ฒ๋ฆฌ |
| Receiver | RDMA pre-registered buffer์์ polling์ผ๋ก ์์ |
| Copy kernel | ํ์์ ๋ฐ์ ๋ฐ์ดํฐ๋ฅผ ์ฌ์ฉ์ layout์ผ๋ก ๋ณต์ฌ |
โ๏ธ ํต์ ์๊ฐ ๋ชจ๋ธ (Eq. 6)
\[ T_c = \max\left( \frac{bahK}{tpa W_a \cdot \text{Util}()}, \frac{beh}{tpe W_e \cdot \text{Util}()} \right) \]- \(bah\): attentionโexpert๋ก ๋ณด๋ผ tensor ํฌ๊ธฐ
- \(beh\): ๋ฐ๋ ๋ฐฉํฅ ํฌ๊ธฐ
- \(W_a, W_e\): attention/expert GPU์ NIC bandwidth
- Util(): empirical bandwidth utilization ํจ์
๐ ์ฑ๋ฅ ๊ฐ์ ์์น ์์ฝ (NCCL vs MegaScale-Infer)
| ์ธก์ ํญ๋ชฉ | ๊ฐ์ ์์น |
|---|---|
| Median Latency (256KB) | 68.2% ๊ฐ์ |
| P99 Latency (256KB) | 92.9% ๊ฐ์ |
| Throughput (256KB) | 4.2ร ํฅ์ |
| M,N ํ์ฅ์ฑ (16โ32) | Tail Latency: ์ต๋ 96.9% ๊ฐ์, Throughput: 5.8ร ํฅ์ |
๐ ๊ฒฐ๋ก
- Algorithm 1์ pipeline ์กฐ๊ฑด์ ๋ง์กฑํ๋ฉด์ GPU resource๋ฅผ ๊ฐ์ฅ ํจ์จ์ ์ผ๋ก ์ฐ๋ ๋ฐฐ์น ๊ณํ์ searchํฉ๋๋ค.
- M2N ํต์ ์ต์ ํ๋ ๊ธฐ์กด ๋ผ์ด๋ธ๋ฌ๋ฆฌ์ ๋ณ๋ชฉ์ ์ ๊ฑฐํ๊ณ , ๋๊ท๋ชจ expert ๋ถ์ฐ ํ๊ฒฝ์์ ์์ ์ ์ด๊ณ ๊ณ ์ฑ๋ฅ ํต์ ์ ๊ฐ๋ฅํ๊ฒ ํฉ๋๋ค.
ํต์ฌ ์๊ณ ๋ฆฌ์ฆ์ ์์๋ก ๋ฐ๋ผ๊ฐ๊ธฐ
๋ฌผ์ด๋ณธ ๊ฒ: ๋ ผ๋ฌธ์ ํต์ฌ์๊ณ ๋ฆฌ์ฆ์ ์์๋ฅผ ๋ค์ด ์์ฃผ ์์ธํ๊ฒ ์ค๋ช ํด์ฃผ๊ณ ์ ๋ฆฌํด ์ค
๋ ผ๋ฌธ์ ํต์ฌ ์๊ณ ๋ฆฌ์ฆ์ ํฌ๊ฒ ๋ ๋ถ๋ถ์ผ๋ก ๊ตฌ์ฑ๋ฉ๋๋ค:
- Disaggregated Expert Parallelism ๊ธฐ๋ฐ์ ๋ฐฐ์น ์ต์ ํ ์๊ณ ๋ฆฌ์ฆ (Algorithm 1)
- Attention โ FFN ๊ฐ Ping-Pong Pipeline Parallelism
์ด ๋ ๊ฐ์ง๋ ์๋น ์ฑ๋ฅ์ GPU ๋ฐ ๋คํธ์ํฌ ์์ ์ ์ฝ ํ์ ์ต์ ํํ๊ธฐ ์ํ ํต์ฌ ์ค๊ณ์ ๋๋ค. ์๋์์ ์์๊ณผ ํจ๊ป ์์ ๊ธฐ๋ฐ์ผ๋ก ์ ์ฒด ์คํ ํ๋ฆ์ ์์ธํ ์ค๋ช ๋๋ฆฌ๊ฒ ์ต๋๋ค.
๐ง ํต์ฌ ์ปจ์ ์ ๋ฆฌ
| ๊ตฌ์ฑ ์์ | ๋ชฉ์ | ๊ธฐ์ ์์ฝ |
|---|---|---|
| Disaggregated Expert Parallelism | Attention๊ณผ FFN์ ๋ถ๋ฆฌํ์ฌ ๊ฐ์ ๋ ๋ฆฝ์ ์ผ๋ก ์ต์ ํ | Attention์ Data Parallelism, FFN์ Expert Parallelism ์ ์ฉ |
| Ping-Pong Pipeline Parallelism | ํต์ -๊ณ์ฐ ์ค๋ฒ๋ฉ, ์์ ์ ํด ์๊ฐ ์ ๊ฑฐ | Micro-batching๊ณผ Layer-Interleaving ํ์ฉ |
| Algorithm 1 | ์ ๊ตฌ์กฐ์์ Throughput per Dollar๊ฐ ์ต๋ํ๋๋๋ก ๋ฐฐ์น ์ ๋ต ํ์ | Tensor parallelism ํฌ๊ธฐ, micro-batch ์, attention node ์ ํ์ |
๐งช ์์ ๊ธฐ๋ฐ ์ค๋ช
โณ๏ธ ๊ฐ์ ์ค์
- ๋ชจ๋ธ: Mixtral 8ร22B (hidden size = 6144, intermediate dim = 16384, 56 layers)
- top-k = 2, expert ์ = 8
- GPU: A100 (TFLOPS = 312, Bandwidth = 2 TB/s)
- batch size = 156, micro-batch ๊ฐ์ \( m = 4 \)
- TP size: attention = 2, expert = 2
๐ช ์ ์ฒด ์๊ณ ๋ฆฌ์ฆ ์คํ ํ๋ฆ (์ ๋ฆฌ)
โ [์ ์ฒ๋ฆฌ] Attention/FFN ์ฐ์ฐ ํน์ฑ ๋ถ์
- Attention (QKV projection + Attention output):
- input: \((b_a, h)\), param: \((h, h \cdot (1 + 2/g) / tpa)\)
- FFN (top-k expert๋ก ๋ถ๊ธฐ๋ sub-batch):
- input: \((b_e, h)\), param: \((h, h' / tpe)\)
โก [๊ณ์ฐ ์๊ฐ ๋ชจ๋ธ๋ง] Pipeline ๊ท ํ ์กฐ๊ฑด ๊ณ์ฐ
Ta, Te๋ ๋ค์๊ณผ ๊ฐ์ด ๋ชจ๋ธ๋ง:
\[ T_a = k_1 b_a + k_2, \quad T_e = k_3 b_e + k_4 \]์ฌ๊ธฐ์ \( b_e = \frac{B \cdot K}{E} = \frac{156 \cdot 2}{8} = 39 \)
๊ท ํ ์กฐ๊ฑด:
\[ T_a \approx T_e \Rightarrow n_a = \frac{k_1 E}{k_3 K} \]โ ์ด๋ก๋ถํฐ attention node ์ \(n_a\) ๊ฒฐ์ (ex: 2๊ฐ ์ ๋๋ก ๊ณ์ฐ๋ ์ ์์)
โข [Pipeline ์กฐ๊ฑด ๊ณ์ฐ] micro-batch ์ \(m\) ์ ํ
ํต์ ์๊ฐ < ๊ณ์ฐ ์๊ฐ (Ta, Te) ์ด๋ผ๊ณ ๊ฐ์ ํ๋ฉด
\[ m \geq 2 \left(1 + \frac{T_c}{T_f}\right), \quad \text{where } T_f = \max(T_a, T_e) \]์: \( \frac{T_c}{T_f} = 0.3 \) ์ด๋ฉด โ \( m \ge 2(1 + 0.3) = 2.6 \) โ ์ต์ 3๊ฐ์ micro-batch ํ์
โฃ [๋ฐฐ์น ๊ณํ ํ๊ฐ] SIMULATE(plan)
- ์ต๋ batch size \(B\) ์ถ์ (latency ์ ํ ๊ณ ๋ ค, ์ 5 ๊ธฐ๋ฐ):
์: \(T_a = T_e = 2\)ms, \(T_c = 0.5\)ms, \(m = 4\), \(L = 56\)
\[ T_{total} \approx 2 + 2 + 1 + 2 \cdot (4 \cdot 56 - 1) = 447 \text{ ms} \]- ์ด ๊ฒฐ๊ณผ๋ฅผ ํตํด latency SLA ๋ง์กฑ ์ฌ๋ถ ํ๊ฐ ํ, Throughput per Dollar ๊ณ์ฐ:
โค [์ต์ข ์ ํ] ๊ฐ์ฅ ๋์ tpd๋ฅผ ๊ฐ์ง๋ ๋ฐฐ์น ์ ํ
๋ฐ๋ณต์ ์ผ๋ก \(tpa, tpe, m\) ์กฐํฉ์ ๋ฐ๊พธ๋ฉฐ ํ์ํ์ฌ ์ต๊ณ ๊ณํ ์ ํ.
๐ Ping-Pong Pipeline ์คํ ๊ตฌ์กฐ ์๊ฐํ
๊ฐ micro-batch๋ attention โ FFN โ attention โ … ์์๋ก layer๋ฅผ ํ๊ณ ํ๋ฆ.
Layer 1 Layer 2 Layer 3
Micro 1: [Attn1]โ[FFN1]โ[Attn2]โ[FFN2]โ[Attn3]โ[FFN3]
Micro 2: [Attn1]โ[FFN1]โ[Attn2]โ[FFN2]โ...
Micro 3: [Attn1]โ[FFN1]โ[Attn2]โ...
Micro 4: [Attn1]โ[FFN1]โ...
=> ๊ณ์ฐ๊ณผ ํต์ ์ด ์ค๋ฒ๋ฉ๋์ด attention/FFN ๋ชจ๋ idle ์๊ฐ์ด ๊ฑฐ์ ์์๐งพ ์์ฝ ์ ๋ฆฌํ
| ํญ๋ชฉ | ์ค๋ช |
|---|---|
| ๋ชฉ์ | FFN-GPU๊ฐ compute-intensive, Attention-GPU๊ฐ memory-intensive์ธ ์ฑ๊ฒฉ์ ๋ง๊ฒ ๋ ๋ฆฝ ์ต์ ํ |
| ์ ๋ ฅ | ๋ชจ๋ธ ํฌ๊ธฐ, MoE ๊ตฌ์กฐ, top-k, GPU ๋ฉ๋ชจ๋ฆฌ, latency SLO |
| ์๊ณ ๋ฆฌ์ฆ ์ถ๋ ฅ | ์ต์ tensor parallelism (TP), micro-batch ์, attention node ์ |
| ํต์ฌ ์์ด๋์ด | \(T_a \approx T_e\) ์กฐ๊ฑด๊ณผ ํ์ดํ๋ผ์ธ ์กฐ๊ฑด์ ๋ง์กฑํ๋ฉด์ throughput/cost ์ต๋ํ |
| ๊ฒฐ๊ณผ | ์ต๋ 1.9ร throughput ํฅ์, 1.86ร cost-per-token ๊ฐ์ |
์คํ ๊ฒฐ๊ณผ์ ๋ค๋ฅธ ๋ฐฉ๋ฒ๋ก ๊ณผ์ ์ฐจ์ด
๋ฌผ์ด๋ณธ ๊ฒ: ์ด ๋ ผ๋ฌธ์์ ์ ์ํ ๊ฒฐ๊ณผ๋ฅผ ์์ธํ๊ฒ ๋ณด๊ณ ๋ค๋ฅธ ๋ฐฉ๋ฒ๋ก ์ ๋นํ์ฌ ํน์ถ๋ ์ ๊ณผ ๋ ผ๋ฌธ์์ ์ ๊ธฐํ๋ ์ด๋ ํ ๋ฐฉ๋ฒ์ด ์ด๋ฌํ ๊ฒฐ๊ณผ๋ฅผ ๋์ถํ๊ฒ ๋์๋์ง ๋ ผ๋ฌธ์์ ์ ์ํ๋ ์ด์ ์ ๋์ ์๊ฐ์ ์๋ ค์ค
์์ฃผ ์ข์ ์ง๋ฌธ์ ๋๋ค. ์ด ๋ ผ๋ฌธ์ด ์ ์ํ๋ ๊ฒฐ๊ณผ๋ MoE ๊ธฐ๋ฐ LLM ์๋น์์ ์ฑ๋ฅ ๋๋น ๋น์ฉ ํจ์จ์ฑ(per-cost throughput)์ ๊ทน๋จ์ ์ผ๋ก ๋์๋ค๋ ์ ์์ ๋งค์ฐ ์๋ฏธ ์์ต๋๋ค. ๋ค์๊ณผ ๊ฐ์ ๋ถ์ ํ๋ ์์ผ๋ก ์ค๋ช ๋๋ฆด๊ฒ์:
โ 1. ๋ ผ๋ฌธ์์ ์ ์ํ ์ ๋์ ์ฑ๋ฅ ๊ฒฐ๊ณผ ์์ฝ
๐ง ์คํ ํ๊ฒฝ
- ๋ชจ๋ธ: Mixtral-8ร22B, DBRX, Scaled-MoE(317B)
- ํ๊ฒฝ: A100 ๊ธฐ๋ฐ homogeneous, H20 + L40S ๊ธฐ๋ฐ heterogeneous cluster
- Metric: Per-GPU throughput (token/s/GPU), Per-dollar throughput
๐ ์ฃผ์ ๋น๊ต ๊ฒฐ๊ณผ
[1] Homogeneous Deployment (A100 ๊ธฐ์ค)
| ๋ชจ๋ธ | vLLM ๋๋น | TensorRT-LLM ๋๋น | MegaScale-Infer ๊ธฐ์ฌ์ |
|---|---|---|---|
| Mixtral 8x22B | 2.56ร โ | 1.28ร โ | FFN compute ์ง์ฝํ + attention ๋ถ์ฐ |
| DBRX | 1.70ร โ | 1.30ร โ | ๋ฐฐ์น ์ต์ ํ ํตํ pipeline ๊ท ํ |
| Scaled-MoE | 7.11ร โ | 1.90ร โ | multi-node์ ์ต์ ํ๋ ํต์ ๊ตฌ์กฐ |
[2] Heterogeneous Deployment (H20: Attention / L40S: FFN)
| ๋ชจ๋ธ | MegaScale vs. vLLM(H20) | MegaScale vs. TRT-LLM(H20) |
|---|---|---|
| Mixtral 8x22B | 3.24ร per-cost โ | 1.86ร per-cost โ |
โ 2. ์ฑ๋ฅ ํฅ์์ ํต์ฌ ์์ธ: ๋ ผ๋ฌธ์ด ์ ์ํ ๊ธฐ์ฌ์ ๊ณผ ๊ทผ๊ฑฐ
| ๋ ผ๋ฌธ ์ ์ | ๊ฒฐ๊ณผ์ ๊ธฐ์ฌํ ๋ฐฉ์ | ๋ ผ๋ฌธ ๋ด ๊ทผ๊ฑฐ |
|---|---|---|
| 1. AttentionโFFN ๋ถ๋ฆฌ (Disaggregation) | FFN ์ชฝ์ ๋ฐฐ์น๋ ํ ํฐ ์ ์ฆ๊ฐ โ GPU utilization ์ฆ๊ฐ | ยง3, ยง4: โFFNs transition from memory- to compute-intensiveโ |
| 2. Ping-Pong Pipeline | FFN/Attention idle time ๊ฐ์ โ ์์ utilization ์ฆ๊ฐ | ยง4.1: โhide communication latency & balance compute timeโ |
| 3. M2N ํต์ ์ต์ ํ | Token routing ๋ณ๋ชฉ ์ ๊ฑฐ โ Tail latency ๊ฐ์ โ Throughput ์ฆ๊ฐ | ยง5, Figure 10โ11: ์ต๋ 4.2ร throughput ์ฆ๊ฐ |
| 4. Heterogeneous Deployment ์ ๋ต | ๋น์ฉ ๋๋น ์ต์ ํ๋ ํ๋์จ์ด ๋งค์นญ (L40S๋ ์ฐ์ฐ, H20์ memory) | ยง4.3, Table 3, Figure 9: “maximize cost-effective memory vs compute” |
| 5. ๋ฐฐ์น ๊ณํ ํ์ ์๊ณ ๋ฆฌ์ฆ | ๋ชจ๋ ๊ตฌ์ฑ ์กฐํฉ ์ค throughput/cost ์ต์ plan ํ์ | ยง4.2: Algorithm 1 ๊ธฐ๋ฐ ๊ณํ ์๋ฆฝ |
๐ก 3. ๋ด ์๊ฐ: ๋ค๋ฅธ ๋ฐฉ๋ฒ๋ก ๋๋น ํน์ถ๋ ์
๐ฅ ๊ธฐ์กด ์์คํ (vLLM, TensorRT-LLM)์ ํ๊ณ
| ์์คํ | ํ๊ณ |
|---|---|
| vLLM | ํตํฉํ ๊ตฌ์กฐ๋ก ์ธํด FFN ์ชฝ์ token sparsity ๋ฐ์ โ GPU ํ์ฉ๋ฅ โ |
| TensorRT-LLM | kernel-level ์ต์ ํ๋ ์ ๋์ด ์์ผ๋ FFN๊ณผ Attention์ ๋ถ๋ฆฌํ์ง ์์ |
๐งจ MegaScale-Infer์ ํน์ถ๋ ์
| ์ฐจ๋ณ์ | ์ ํน์ถ๋๊ฐ? |
|---|---|
| ๊ตฌ์กฐ์ ๋ถํด (attention vs FFN) | ๋จ์ํ kernel-level ์ต์ ํ๊ฐ ์๋ ์์คํ ๊ตฌ์กฐ ์์ฒด๋ฅผ ์ฌ์ค๊ณํจ |
| ํต์ ์ต์ ํ M2N ์ค๊ณ | ๊ธฐ์กด NCCL์ด ์ ๊ณตํ์ง ์๋ MoE token routing ํจํด์ ์ํด purpose-built ๊ตฌํ |
| ๋จ๊ฐ ์ต์ ๋ฐฐ์น ์ค๊ณ | ๋จ์ throughput์ด ์๋ throughput per dollar ์ต์ ํ๋ฅผ ๋ฌ์ฑํจ |
| ํ์ดํ๋ผ์ธ ํ์ฑํ | m micro-batch ์ค๊ณ๋ก ๊ณ์ฐ-ํต์ ์ค๋ฒ๋ฉ์ ๊ทน๋ํํด idle ์ ๊ฑฐํจ |
๐งฉ ๊ฒฐ๋ก ์ ๋ฆฌ
| ํญ๋ชฉ | ๋ด์ฉ |
|---|---|
| ๋ ผ๋ฌธ์ด ์ ์ํ ๋ฐฉ๋ฒ๋ก | Disaggregated expert parallelism + ping-pong pipeline + M2N ํต์ ์ต์ ํ |
| ์ฑ๊ณผ์ ์ ๋์ ์ฆ๊ฑฐ | ์ต๋ 7.11ร throughput โ, 3.24ร per-cost throughput โ |
| ์ฐจ๋ณํ๋ ์ค๊ณ ํฌ์ธํธ | ์๋น ๊ตฌ์กฐ ์์ฒด๋ฅผ ๋ถํดํ๊ณ , ์ ์ฒด ํ์ดํ๋ผ์ธ์ compute/memory bottleneck ๋ถ๋ฆฌ ์ฒ๋ฆฌ |
| ๋ด ์๊ฐ | “๊ตฌ์กฐ์ ์ค๊ณ ๋ณ๊ฒฝ + ํต์ ์ต์ ํ + ๋ฐฐ์น ํ์"์ด LLM ์๋น์ ์๋ก์ด ์ค์ฉ์ ๋ฐฉํฅ์ ์ ์ํจ. ๋จ์ํ CUDA ์ต์ ํ๋ณด๋ค ์ด์ฒด์ ์์คํ ์์ง๋์ด๋ง ๊ด์ ์ด ๋๋ณด์ |
์ ๋ ฅ ๋ฐ์ดํฐ์ ์ถ๋ก ๊ณผ์ , ๋ชจ๋ธ ๊ตฌ์กฐ์ ์์ ์๊ตฌ๋
๋ฌผ์ด๋ณธ ๊ฒ: ์ด ๋ ผ๋ฌธ์ ์ ๋ ฅ๋ฐ์ดํฐ์ ์ถ๋ก ๊ณผ์ ์ ๋ํด์ ์์๋ฅผ ๋ค์ด ์์ฃผ ์์ธํ๊ฒ ์ค๋ช ํด ์ฃผ๊ฒ ๋? ์ถ๊ฐ์ ์ผ๋ก ๋ชจ๋ธ์ํคํ ์ฒ์ ๊ตฌ์ฑ ๋ฐ ๋ชจ๋ธ์ ๊ตฌ์ฑํ๋ ์ฐ์ฐ๊ณผ ๋ฉ๋ชจ๋ฆฌ ์๊ตฌ๋ ์ปดํจํ ์๊ตฌ๋ ๋ํด์๋ ์ค๋ช ํด์ค
์ข์ต๋๋ค. ์ด ๋ ผ๋ฌธ์์ ์ฌ์ฉํ๋ LLM์ MoE ๊ธฐ๋ฐ์ Transformer๋ก ๊ตฌ์ฑ๋์ด ์์ผ๋ฉฐ, ๋ ผ๋ฌธ ์ ์ฒด๊ฐ ์๋น ์์คํ (ํนํ decoding phase)์ ์ต์ ํ์ ์ง์ค๋์ด ์์ต๋๋ค. ๋ฐ๋ผ์ ์๋ ๋ด์ฉ์ ์ค์ฌ์ผ๋ก ์ ๋ฆฌํ๊ฒ ์ต๋๋ค:
๐ ์ค๋ช ๊ตฌ์กฐ ์์ฝ
- ์ ๋ ฅ ๋ฐ์ดํฐ ์์
- ์ถ๋ก ๊ณผ์ (Prefill vs Decoding)
- ๋ชจ๋ธ ์ํคํ ์ฒ ๊ตฌ์ฑ
- ์ฃผ์ ์ฐ์ฐ ๋ฐ ๋ฉ๋ชจ๋ฆฌ/์ปดํจํ ์๊ตฌ๋ ๋ถ์
1. ๐ฅ ์ ๋ ฅ ๋ฐ์ดํฐ ์์
๋ ผ๋ฌธ ๊ธฐ์ค ์คํ ์ค์ ์์:
- Prompt input ๊ธธ์ด (median): 571 tokens
- Output ๊ธธ์ด (median): 159 tokens
๐ฏ ์์
Input prompt:
"Once upon a time, there was a kingdom where people communicated only using code..."
Tokenized: [1012, 4021, 1029, 4890, 8923, 3401, ...]
Total input tokens = 5712. โ๏ธ ์ถ๋ก ๊ณผ์ : Prefill vs Decoding
๐ข [Phase 1] Prefill
- ๋ชฉ์ : ์ ๋ ฅ ์ํ์ค ์ ์ฒด(571 tokens)์ attention์ ํ ๋ฒ์ ๊ณ์ฐ
- ์ฐ์ฐ ํน์ง:
- Attention: ๋ชจ๋ token ๊ฐ ๊ด๊ณ ๊ณ์ฐ โ ๋งค์ฐ compute-intensive
- FFN: ๋ชจ๋ token์ ๋์ผํ๊ฒ ์ ์ฉ (sparseํ์ง ์์)
- Key-Value (KV) cache ์์ฑ: attention ๊ฒฐ๊ณผ ์ ์ฅ
๐ต [Phase 2] Decoding
- ๋ชฉ์ : 1-step autoregressive token ์์ฑ ๋ฐ๋ณต (์: 159ํ ๋ฐ๋ณต)
- ์ฐ์ฐ ํน์ง:
- Attention: KV cache ์ฝ๊ธฐ โ memory-intensive
- FFN: top-k expert๋ง ํ์ฑํ๋จ โ sparse, compute volume โ, GPU utilization โ
- โ ์ด ๋ฌธ์ ๋ฅผ MegaScale-Infer๊ฐ ํด๊ฒฐํจ
3. ๐งฑ ๋ชจ๋ธ ์ํคํ ์ฒ ๊ตฌ์ฑ
๋ ผ๋ฌธ์์ ์ฌ์ฉํ๋ ๋ชจ๋ธ์ ์ผ๋ฐ์ ์ธ MoE ๊ธฐ๋ฐ Transformer Layer์ ๋๋ค.
| ๊ตฌ์ฑ ์์ | ์ค๋ช |
|---|---|
| Layer ์ | 48~56 (Mixtral: 56) |
| Hidden dim | 6144 (e.g., Mixtral) |
| Intermediate dim | 16384 |
| Attention | Grouped-query Attention (GQA) |
| FFN | Mixture-of-Experts (MoE), top-k = 2 |
| Expert ์ | 8, 16, 32 ๋ฑ ๊ตฌ์ฑ์ ๋ฐ๋ผ ๋ค๋ฆ |
๐ ์: Mixtral 8x22B ๊ตฌ์กฐ
- ์ด 56๊ฐ layer
- ๊ฐ layer์๋
- GQA attention
- MoE FFN (top-2 of 8 experts ์ฌ์ฉ)
4. ๐ ์ฃผ์ ์ฐ์ฐ ๋ฐ ์์ ์๊ตฌ๋
โ Attention ์ฐ์ฐ
| ์ฐ์ฐ | ์ ๋ ฅ | ํ๋ผ๋ฏธํฐ ํฌ๊ธฐ | ์ฐ์ฐ๋ |
|---|---|---|---|
| QKV Projection | \((b_a, h)\) | \((h, h \cdot (1+2/g)) / tpa\) | GEMM |
| Attention Output | \((b_a, h/tpa)\) | \((h/tpa, h)\) | GEMM |
- ๋ฉ๋ชจ๋ฆฌ ์๊ตฌ๋ (per token): KV cache (2รhiddenรsequence length)
- ํน์ง: prefill์ compute, decoding์ memory access bottleneck
โ FFN (MoE) ์ฐ์ฐ
| ์ฐ์ฐ | ์ ๋ ฅ | ํ๋ผ๋ฏธํฐ ํฌ๊ธฐ | ์ฐ์ฐ๋ |
|---|---|---|---|
| FFN Input | \((b_e, h)\) | \((h, h')/tpe\) | GEMM |
| FFN Output | \((b_e, h')/tpe\) | \((h', h)\) | GEMM |
- top-2 experts๋ง ํ์ฑํ๋จ โ b_e = B ร topk / #experts
- โ FFN์ GPU ํ์ฉ๋ฅ ์ด ๊ธ๊ฒฉํ ๊ฐ์ํ๋ ๋ฌธ์
๐พ ์์ ์๊ตฌ๋ ์์น ์์ (Mixtral 8x22B, A100 ๊ธฐ์ค)
๐ ๊ณ์ฐ๋ ์ต์ batch size (Roofline Model ๊ธฐ์ค)
A100 GPU:
- Compute: 312 TFLOPS
- Memory BW: 2 TB/s
โ \( b_{min} = \frac{F}{B} = \frac{312}{2} = 156 \)
๐ FFN ๋น token ์ (MoE sparsity ๋ฐ์)
\[ b_{e} = 156 \times \frac{2}{8} = 39 \text{ tokens per expert} \]โ ์ด๋ GEMM์ compute ํจ์จ์ฑ์ด ๋งค์ฐ ๋จ์ด์ง๋ ์์ค
โ GPU utilization: 25% ์ดํ
๐ข KV Cache ์ฉ๋ ์์
- KV per token (bfloat16): \(2 \cdot h = 2 \cdot 6144 \cdot 2 \text{ bytes} = 24 KB/token\)
- Sequence 571 + generated 159 = 730 tokens
โ KV cache per sequence: 730 ร 24 KB = 17.5 MB
โ batch size 156 ๊ธฐ์ค: ~2.7 GB (GPU memory์ ์๋นํ ๋ถ๋ด)
๐งพ ์ ๋ฆฌ
| ํญ๋ชฉ | ์ค๋ช |
|---|---|
| ์ ๋ ฅ | ํ๊ท 571 token์ prompt |
| ์ถ๋ก ๋จ๊ณ | Prefill (dense) โ Decoding (sparse, memory-bound) |
| ์ํคํ ์ฒ | 56-layer Transformer, GQA + MoE (top-2 of 8 experts) |
| ์ฐ์ฐ ๋ณ๋ชฉ | Prefill: attention compute Decoding: FFN compute (sparse), attention memory access |
| ์ฐ์ฐ๋ | FFN: ๋งค์ฐ ์์์ ธ compute utilizationโ, attention์ KV cache access๋ก bandwidth โ |
| ์์ ์ด์ | A100 ๊ธฐ์ค batch 156 ์ด์ ํ์, ๊ทธ๋ฌ๋ sparse MoE ๊ตฌ์กฐ์์๋ 39 tokens/expert ์์ค๋ง ์ฌ์ฉ๋จ |
โ ์ด ๋ชจ๋ ๊ตฌ์กฐ์์ MegaScale-Infer์ ์ต์ ํ ์์
| ๋ฌธ์ | MegaScale-Infer ํด๊ฒฐ ๋ฐฉ์ |
|---|---|
| FFN์ ์ ์ ํ ํฐ ๋ถ๋ฐฐ โ GPU ๋ญ๋น | attention module ๋ถ๋ฆฌ โ ์ฌ๋ฌ attention์ด FFN์๊ฒ token์ aggregate |
| attentionโFFN ๊ฐ ํต์ ๋ณ๋ชฉ | M2N ํต์ ๊ตฌ์กฐ๋ก latency/tail ์์ ํ |
| KV cache memory โ โ batch size ์ ์ฝ | attention node๋ฅผ ๋ฉ๋ชจ๋ฆฌ ์ค์ฌ GPU(H20)๋ก ๋ถ๋ฆฌ |
ํ์ํ์๋ฉด KV cache ์์ถ, expert load balancing ๋ฐฉ์, or ๊ฐ ๋จ๊ณ๋ณ PyTorch ์ฐ์ฐ trace ์์๋ ๋์๋๋ฆด ์ ์์ต๋๋ค. ์ด๋ค ๋ฐฉํฅ์ผ๋ก ๋ ํ์ฅํด๋ณผ๊น์?
ํ๊ณ๋ฅผ ๋์ผ๋ ค๋ ์ฐ๊ตฌ ํ๋ฆ
๋ฌผ์ด๋ณธ ๊ฒ: ์ด ๋ ผ๋ฌธ์ ํ๊ณ๋ฅผ ๊ทน๋ณตํ๊ธฐ ์ํ ๋ฐฉ๋ฒ์ผ๋ก ์ด๋ค ์ฐ๊ตฌํ๋ฆ์ด ์๋์ง ์ ๋ฆฌํด์ ์์ธํ๊ฒ ์ค๋ช ํด ์ค
์ข์ต๋๋ค. ์ด ๋ ผ๋ฌธ์ ํ๊ณ๋ฅผ ๊ทน๋ณตํ๊ฑฐ๋ ๋ณด์ํ ์ ์๋ ๋ฐฉํฅ์ ํฌ๊ฒ ๋ค์๊ณผ ๊ฐ์ด 3๊ฐ์ง ์ฐ๊ตฌ ํ๋ฆ์ผ๋ก ๋๋ ์ ์์ต๋๋ค:
๐ [A] Fine-Grained Expert Scheduling & Load Balancing
๐ ํ๊ณ:
- MegaScale-Infer๋ top-k expert์ token์ ์ผ๊ด์ ์ผ๋ก ๋ผ์ฐํ ํ๋ ๊ตฌ์กฐ์ด๋ฉฐ, load imbalance๊ฐ ๋ฐ์ํ๋ฉด ํน์ expert node๊ฐ ๋ณ๋ชฉ์ด ๋จ.
- ๋ ผ๋ฌธ์์๋ ์ด๋ฅผ ํด๊ฒฐํ๊ธฐ ์ํด on-device redundancy + greedy scheduling์ ์ฌ์ฉํ์ผ๋, traffic-aware dynamic load balancing์ ์ ํ์ ์.
โ ๋์ ์ฐ๊ตฌ ํ๋ฆ:
| ์ฐ๊ตฌ | ๊ธฐ๋ฒ | ์ค๋ช |
|---|---|---|
| Tutel (Hwang et al., MLSys 2023) | Adaptive Routing | token๋ง๋ค ์ ๋ฌธ๊ฐ ์ ํ ํ๋ฅ ์ ํ์ตํ์ฌ traffic skew ์ต์ํ |
| Brainstorm (Cui et al., OSDI 2023) | Expert-level scheduling | Expert์ popularity๋ฅผ ๊ธฐ๋ฐ์ผ๋ก GPU-Expert ๋งคํ ์ต์ ํ |
| MoE-Lightning (Cao et al., arXiv 2024) | expert preloading | token traffic ํ์คํ ๋ฆฌ ๊ธฐ๋ฐ์ผ๋ก ๋ฏธ๋ฆฌ expert๋ฅผ preload ํ์ฌ cold-start ๋ฐฉ์ง |
๐ [B] Token Routing Overhead ์ํ
๐ ํ๊ณ:
- MegaScale-Infer์ M2N ํต์ ๊ตฌ์กฐ๋ ๊ณ ์ฑ๋ฅ์ด์ง๋ง ์ฌ์ ํ ๋๊ท๋ชจ ์์คํ ์์ ๋คํธ์ํฌ ๋ณ๋ชฉ ๋ฐ์ ๊ฐ๋ฅ.
- ํนํ top-k ์ ๋ฌธ๊ฐ ์ ์ฆ๊ฐ ์ โ ํต์ ๋ ์ฆ๊ฐ โ tail latency ์ฆ๊ฐ
โ ๋์ ์ฐ๊ตฌ ํ๋ฆ:
| ์ฐ๊ตฌ | ๊ธฐ๋ฒ | ์ค๋ช |
|---|---|---|
| Switch Transformer (Fedus et al., JMLR 2022) | Top-1 routing | ๋จ์ผ expert๋ง ํ์ฑํํ์ฌ ํต์ ๋ ์์ฒด๋ฅผ ์ต์ํํจ |
| Janus (Liu et al., SIGCOMM 2023) | Unified sparse communication | MoE training/inference ๋ชจ๋๋ฅผ ์ํ ํต์ abstraction layer ์ ๊ณต |
| Pre-gated MoE (Hwang et al., ISCA 2024) | Token-to-Expert mapping ์ฌ์ ๊ฒฐ์ | inference์์ ๋ผ์ฐํ ์ฐ์ฐ ์์ฒด ์ ๊ฑฐ, ์คํ๋ผ์ธ์ผ๋ก tokenโexpert ๋งคํ์ ๋ฏธ๋ฆฌ ๊ฒฐ์ ํจ |
๐ง [C] Computation-Memory Tradeoff ๋ฐ Cache ์ต์ ํ
๐ ํ๊ณ:
- MegaScale-Infer๋ attention module์ KV cache๋ฅผ ๋ถ๋ฆฌํ๊ณ ๋ณต์ ํ์ง๋ง, ์ฌ์ ํ memory bottleneck์ด ์กด์ฌํ๋ฉฐ batch size scaling์ ์ ์ฝ
โ ๋์ ์ฐ๊ตฌ ํ๋ฆ:
| ์ฐ๊ตฌ | ๊ธฐ๋ฒ | ์ค๋ช |
|---|---|---|
| PagedAttention (vLLM, SOSP 2023) | Dynamic KV Cache Paging | GPU ๋ฉ๋ชจ๋ฆฌ์ ๋ค์ด์ค์ง ์๋ KV๋ disk๋ก swapping ํ์ฌ batch size ํ๋ ๊ฐ๋ฅ |
| Mooncake (Qin et al., arXiv 2024) | KV cache centric ์๋ฒ ์ค๊ณ | KV cache๋ฅผ ์ค์ฌ์ผ๋ก ์ค๊ณ๋ ์๋น ์์คํ : streaming/eviction ์ต์ ํ |
| Splitwise (Patel et al., NSDI 2023) | Prefill / Decoding ๋ถ๋ฆฌ ํด๋ฌ์คํฐ | MegaScale-Infer๋ ์ด ์์ด๋์ด๋ฅผ ์ฐจ์ฉํ์ง๋ง, Prefill phase๊น์ง ์์ ๋ถ์ฐํ๋ ๊ตฌ์กฐ๋ก ํ์ฅ ๊ฐ๋ฅ |
๐ ์ข ํฉ ์ ๋ฆฌ
| ์นดํ ๊ณ ๋ฆฌ | MegaScale-Infer ํ๊ณ | ๋์ ์ฐ๊ตฌ ํ๋ฆ | ํต์ฌ ๊ฐ์ ์์ด๋์ด |
|---|---|---|---|
| A. Expert Load | Expert ๊ฐ ๋ถ๊ท ํ | Tutel, Brainstorm | Expert ์ ํ ํ๋ฅ ํ์ต, load-aware mapping |
| B. Token Routing | M2N ํต์ ์ฆ๊ฐ | Janus, Pre-gated MoE | Top-1 routing, ์ฌ์ expert mapping |
| C. Memory Bottleneck | KV cache ์๋ฐ | vLLM, Mooncake | KV cache paging, streaming KV cache ์ค๊ณ |
๐ก ๋ด ์๊ฒฌ: ํฅํ ์ฐ๊ตฌ ๋ฐฉํฅ
Routing-Aware Scheduling + M2N Congestion Feedback Loop
- ์ค์๊ฐ expert ๋ก๋ + ํต์ ๋์ ๊ธฐ๋ฐ์ผ๋ก token routing์ ๋์ ์ผ๋ก ์กฐ์ ํ๋ feedback ๊ธฐ๋ฐ token scheduler ํ์
KV Cache Hierarchy + Compression
- CPU DRAM / SSD / GPU์ multi-level cache hierarchy๋ฅผ ๋ง๋ค๊ณ ,
- low-rank approximation์ด๋ lossy attention์ ํ์ฉํ KV cache compression ์ ์ฉ
Expert Graph Sharding
- Expert ๋ ธ๋๋ค ๊ฐ token ํ๋ฆ์ ๊ทธ๋ํ๋ก ๋ณด๊ณ , ์ด๋ฅผ ํต์ /๊ณ์ฐ ๋ณ๋ชฉ์ ๊ธฐ์ค์ผ๋ก ํํฐ์ ๋ํ๋ ๊ตฌ์กฐ ๋์ ๊ฐ๋ฅ
๐ ๊ฒฐ๋ก ์์ฝ
MegaScale-Infer๋ attention-FFN ๋ถ๋ฆฌ + pipeline + M2N ํต์ ์ต์ ํ๋ก MoE ์๋น์ ์๋ก์ด baseline์ ์ ์ํ์ง๋ง,
load imbalance, ํต์ ์ฆ๊ฐ, ๋ฉ๋ชจ๋ฆฌ ์๋ฐ์ด๋ผ๋ ์ธ ๊ฐ์ง ๊ตฌ์กฐ์ ํ๊ณ๋ฅผ ๊ฐ์ง๋๋ค.
์ด๋ฅผ ๊ทน๋ณตํ๋ ์ฐ๊ตฌ๋ค์ ๋ผ์ฐํ ์ต์ ํ, cache paging, ๋ชจ๋ ์ฌ์ ๊ฒฐ์ , ํต์ ์ถ์ํ ๋ฑ์ ๋ฐฉ๋ฒ์ผ๋ก ์๋น ํจ์จ์ฑ, ํ์ฅ์ฑ, ๋น์ฉ ์ ๊ฐ์ ๋์์ ์ถ๊ตฌํฉ๋๋ค.
Q : MegaScale-Infer์์ attention๊ณผ FFN์ ๋ถ๋ฆฌ(disaggregate)ํจ์ผ๋ก์จ ๊ธฐ์กด MoE ์๋น ์ํคํ ์ฒ ๋๋น ์ด๋ค ์์ ํ์ฉ์ ๋ณํ๊ฐ ๋ฐ์ํ๋ฉฐ, ์ด๊ฒ์ด FFN์ GPU utilization์ ์ด๋ป๊ฒ ๊ฐ์ ํ๋๊ฐ?
์ข์ ์ง๋ฌธ์ ๋๋ค. ์๋์ ์ด ์ง๋ฌธ์ ๋ํ ๊ตฌ์กฐ์ ์ธ ํด์ค์ ์ ๊ณตํ๊ฒ ์ต๋๋ค:
๐ฏ ํต์ฌ ์์ง ์์ฝ
attentionโFFN ๋ถ๋ฆฌ(disaggregation)๋ FFN์ token sparsity ๋ฌธ์ ๋ก ์ธํด ๋ฐ์ํ๋ GPU underutilization ๋ฌธ์ ๋ฅผ ํด๊ฒฐํ๊ณ , FFN์ ๋ค์ compute-intensiveํ ์ฐ์ฐ์ผ๋ก ์ ํ์์ผ GPU ์์์ ๋์ฑ ํจ์จ์ ์ผ๋ก ํ์ฉํ๊ฒ ํฉ๋๋ค.
๐ ๊ธฐ์กด MoE ์๋น ๊ตฌ์กฐ์ ๋ฌธ์ ์
MoE ๊ตฌ์กฐ์์ FFN์ top-k experts๋ง ํ์ฑํ๋จ
- ์: top-2 of 8 experts โ ํ expert๋น batch์ 25%๋ง ์ฒ๋ฆฌ
Attention์ full batch size ์ฒ๋ฆฌ, FFN์ ๋ถ์ฐ ์ฒ๋ฆฌ
- Attention์ GPU memory-bound (KV cache access)
- FFN์ GPU compute-bound (GEMM), ๊ทธ๋ฌ๋ ํ ํฐ์ด ์ ์ด compute underutilization ๋ฐ์
์ด๋ก ์ธํด FFN์ด ๋ ์ด์ compute-intensiveํ์ง ์๊ณ , ๋๋ถ๋ถ idle
โ MegaScale-Infer์ ํด๊ฒฐ ๋ฐฉ์: attentionโFFN ๋ถ๋ฆฌ
| ํญ๋ชฉ | ์ค๋ช |
|---|---|
| ๋ถ๋ฆฌ ์ ๋ต | attention๊ณผ FFN์ ์๋ก ๋ค๋ฅธ GPU ๋ ธ๋์ ๋ฐฐ์น |
| attention | ์ฌ๋ฌ replica๋ก ๊ตฌ์ฑ โ ๋ง์ ์์ฒญ์ ๋์์ ์ฒ๋ฆฌ ๊ฐ๋ฅ |
| FFN | ์ฌ๋ฌ attention node๋ก๋ถํฐ ํ ํฐ์ aggregateํด์ ์ฒ๋ฆฌ |
| ๊ฒฐ๊ณผ | FFN์ ์ ๋ฌ๋๋ ํ ํฐ ์๊ฐ ๋์ด๋๊ณ , FFN GPU๊ฐ ๋ค์ compute-intensiveํ ์ํ๊ฐ ๋จ |
๐ ๊ตฌ์ฒด์ ์ธ ๊ฐ์ ์์น (๋ ผ๋ฌธ ๊ธฐ์ค)
Mixtral 8x22B ๊ธฐ์ค, FFN์ theoretical utilization:
๊ธฐ์กด:
\[ \text{util} = \min\left(\frac{\text{topk}}{\text{\#experts}} \cdot \frac{B F}{\text{bandwidth}}, 1\right) = \min\left(\frac{2}{8} \cdot \frac{156 F}{B}, 1\right) = 25\% \]MegaScale-Infer:
- attention์ด N๊ฐ๋ก ๋ถ์ฐ๋์ด ๋์์ ๋ ๋ง์ ์์ฒญ ์์ฑ
- FFN์ด ์ด๋ฅผ ๋ณํฉ ์ฒ๋ฆฌํ์ฌ GPU ์ฌ์ฉ๋ฅ ์ด 1.9ร ์ด์ ํฅ์๋จ
๐ก ํต์ฌ ์ธ์ฌ์ดํธ
๊ธฐ์กด MoE๋ ํ ํฐ์ sparse routing ๊ตฌ์กฐ ๋๋ฌธ์ FFN์ scale-out ํ๋๋ผ๋ GPU ํ์ฉ๋ฅ ์ด ๋ฎ์์
โ MegaScale-Infer๋ attention์ ๋ถ๋ฆฌํ์ฌ ๊ฐ FFN expert๊ฐ multi-attention source๋ก๋ถํฐ ์ถฉ๋ถํ batch๋ฅผ ๋ฐ๊ฒ ๋ง๋ค๊ณ , ์ด๋ฅผ ํตํด FFN์ ์ฐ์ฐ๋์ด ๋ค์ GPU๋ฅผ ๊ฐ๋ ์ฑ์ฐ๋๋ก ์ค๊ณํจ
๐ ์์ฝ
- attentionโFFN ๋ถ๋ฆฌ๋ FFN์ batch size๋ฅผ ํจ๊ณผ์ ์ผ๋ก ์ฆ๊ฐ์์ผ underutilization ๋ฌธ์ ํด๊ฒฐ
- ๊ฒฐ๊ณผ์ ์ผ๋ก FFN์ด ๋ค์ GPU์์ ํจ์จ์ ์ผ๋ก ์คํ๋๋๋ก ํ๊ณ , ์๋น ์ฑ๋ฅ์ ์ ๋ฐ์ ์ผ๋ก ํฅ์์ํด
- ์ด๋ FFN์ dense model ์์ค์ GPU utilization์ผ๋ก ๋๋๋ฆฌ๋ ๊ทผ๋ณธ์ ํด๊ฒฐ์ฑ ์
Q : ping-pong pipeline parallelism์ด ํต์ ์ค๋ฒํค๋๋ฅผ ์๋ํ๊ณ idle time์ ์ค์ด๊ธฐ ์ํด ์ด๋ค ์กฐ๊ฑด์ ๋ง์กฑํด์ผ ํ๋ฉฐ, ์ด ์กฐ๊ฑด๋ค์ด ์ค์ ์์คํ ๋ฐฐ์น ๊ณํ์ ์ด๋ป๊ฒ ๋ฐ์๋๋๊ฐ?
์ข์ ์ง๋ฌธ์ ๋๋ค. ์ด ์ง๋ฌธ์ MegaScale-Infer์ ์ฑ๋ฅ ์ต์ ํ ํต์ฌ ๋ก์ง์ ๊ฟฐ๋ซ๋ ๋งค์ฐ ์ค์ํ ํฌ์ธํธ์ ๋๋ค. ์๋์์ ์กฐ๊ฑด, ์์, ์ง๊ด์ ์๋ฏธ, ๋ฐฐ์น๊ณํ ๋ฐ์ ๋ฐฉ์๊น์ง ์ฐจ๋ก๋๋ก ์ค๋ช ๋๋ฆด๊ฒ์.
โ ping-pong pipeline parallelism์ ๋ชฉ์
- attention ๋ชจ๋๊ณผ FFN ๋ชจ๋์ ๋ถ๋ฆฌํ๋ฉด ์๋ก ๋ฒ๊ฐ์ ์คํ๋๋ฏ๋ก
- ๊ฐ ๋ชจ๋์ด ์๋๋ฐฉ์ ์ฐ์ฐ ๋๋ ํต์ ์ ๊ธฐ๋ค๋ฆฌ๋ฉฐ idleํ๋ ๋ฌธ์ ๊ฐ ๋ฐ์
- ๋ฐ๋ผ์ ์ด๋ฅผ ํด๊ฒฐํ๊ธฐ ์ํด micro-batch ๋จ์๋ก ์ค๋ฒ๋ฉํ๋ pipeline ๊ตฌ์กฐ๋ฅผ ๋์
๐ ํต์ ์๋ + idle ์ ๊ฑฐ๋ฅผ ์ํ ์ธ ๊ฐ์ง ์กฐ๊ฑด
MegaScale-Infer ๋ ผ๋ฌธ์์๋ ๋ค์ 3๊ฐ์ง ์ํ์ ์กฐ๊ฑด์ ํตํด pipeline ํจ์จ์ฑ์ ์ค๋ช ํฉ๋๋ค:
[์กฐ๊ฑด 1] ๊ณ์ฐ ์๊ฐ ๊ท ํ
T_a โ T_e- \(T_a\): attention ๋ ธ๋์์ micro-batch 1๊ฐ ์ฒ๋ฆฌ ์๊ฐ
- \(T_e\): expert ๋ ธ๋์์ micro-batch 1๊ฐ ์ฒ๋ฆฌ ์๊ฐ
- ๋ชฉ์ : ์ฐ์ฐ ํธํฅ์ด ์๊ธฐ์ง ์๋๋ก ํด์ผ pipeline์ ๋ณ๋ชฉ ๋ฐ์ ์ํจ
[์กฐ๊ฑด 2] ํต์ ์๊ฐ๋ณด๋ค ์ฐ์ฐ ์๊ฐ์ด ์ถฉ๋ถํ ๊ธธ์ด์ผ ํจ
T_c < T_f, where T_f = max(T_a, T_e)- \(T_c\): micro-batch ๋น ํต์ ์๋ณต ์๊ฐ (A2E + E2A)
- ์๋ฏธ: compute ์๊ฐ์ด ํต์ ๋ณด๋ค ๊ธธ์ด์ผ ํต์ ์ ์ค๋ฒ๋ฉํ์ฌ ์๋ ๊ฐ๋ฅ
[์กฐ๊ฑด 3] ์ถฉ๋ถํ micro-batch ์
m โฅ 2 ร (1 + T_c / T_f)- \(m\): micro-batch ๊ฐ์
- ์๋ฏธ: pipeline์ด ์ถฉ๋ถํ ์ฑ์์ง๋ ค๋ฉด ์ด ์ ์ด์์ด์ด์ผ ํจ
- ์: \(T_c/T_f = 0.3\)์ด๋ฉด \(m โฅ 2.6 \Rightarrow 3๊ฐ ํ์\)
๐งฎ ์์น ์์ (A100 ๊ธฐ์ค)
๊ฐ์ :
- \(T_a = 2\)ms, \(T_e = 2.5\)ms โ \(T_f = 2.5\)ms
- \(T_c = 0.5\)ms
์ ์ฉ:
m โฅ 2 ร (1 + 0.5 / 2.5) = 2.4 โ ์ต์ m = 3๐ง ์์คํ ๋ฐฐ์น ๊ณํ ๋ฐ์ ๋ฐฉ์ (Algorithm 1)
MegaScale-Infer๋ ์ด ์กฐ๊ฑด๋ค์ ๊ณ ๋ คํ์ฌ deployment plan์ ๋ค์๊ณผ ๊ฐ์ด ์๋์ผ๋ก ๊ตฌ์ฑํฉ๋๋ค:
์กฐ๊ฑด 1์ ๋ง์กฑ์ํค๊ธฐ ์ํด:
balance(G, tpa, tpe)ํจ์์์ attention node ์ \(n_a\)๋ฅผ ์กฐ์ ํ์ฌ- \(T_a \approx T_e\) ๋ง์กฑํ๋๋ก ์ค๊ณ
โ ์์:PLAINTEXTn_a = (k1 ร E) / (k3 ร K)n_a = (k1 ร E) / (k3 ร K)
์กฐ๊ฑด 2, 3์ ๋ง์กฑ์ํค๊ธฐ ์ํด:
- ํต์ ์ฑ๋ฅ ๊ธฐ๋ฐ \(T_c\) ์ถ์ (Eq. 6)
- ์ ์กฐ๊ฑด์ผ๋ก๋ถํฐ ํ์ํ ์ต์ micro-batch ์ \(m\) ๊ณ์ฐ
SIMULATE(plan)์์ ์ค์ ์ฑ๋ฅ ๋ชจ๋ธ์ ๋๋ ค latency ํ์ธ ํ throughput/cost ์ต๋๊ฐ์ ๊ฐ๋ plan ์ ํ
๐ ping-pong ํจ๊ณผ ์คํ ๊ฒฐ๊ณผ ์์ฝ (๋ ผ๋ฌธ Figure 12)
| micro-batch ์ | Throughput (Mixtral ๊ธฐ์ค) | ๊ฐ์ ์จ |
|---|---|---|
| m = 1 | baseline | - |
| m = 2 | 1.9ร โ | pipeline ํ์ฑ ์์ |
| m = 3 | ์ถ๊ฐ 1.10ร โ | ํต์ -์ฐ์ฐ ์์ ์ค๋ฒ๋ฉ |
โ๏ธ ์์ฝ ์ ๋ฆฌ
ping-pong pipeline์ด ํต์ /๊ณ์ฐ ์ค๋ฒ๋ฉํ๋ ค๋ฉด ๋ค์ ์ธ ์กฐ๊ฑด์ด ํ์:
- (1) attention vs FFN compute ์๊ฐ ๊ท ํ
- (2) ํต์ ์๊ฐ์ด ์ฐ์ฐ๋ณด๋ค ์์ ๊ฒ
- (3) ์ถฉ๋ถํ micro-batch ์๋ก pipeline ์ฑ์ฐ๊ธฐ
MegaScale-Infer๋ ๋ฐฐ์น ๊ณํ์ ์์ ๊ธฐ๋ฐ์ผ๋ก ํ์ํ์ฌ ์ด ์กฐ๊ฑด์ ์๋ ๋ฐ์ํจ
๊ฒฐ๊ณผ์ ์ผ๋ก attention๊ณผ FFN์ ์ฐ์ฐ ๋ฐ ํต์ idle time์ด ์ฌ๋ผ์ง๊ณ , GPU utilization๊ณผ latency ๋ชจ๋ ๊ฐ์
Q : ๊ธฐ์กด NCCL ๊ธฐ๋ฐ์ All2All ํต์ ๊ตฌ์กฐ ๋์ MegaScale-Infer๊ฐ ์ ์ํ M2N ํต์ ๊ตฌ์กฐ๋ ์ด๋ค ๊ตฌ์ฒด์ ์ธ ๋ณ๋ชฉ์ ํด๊ฒฐํ๋ฉฐ, ์ด๋ฅผ ์ํด ์ค๊ณ๋ sender/receiver path๋ ์ด๋ค ๋ฐฉ์์ผ๋ก RDMA, GPUDirect, ACK ์ต์ ํ๋ฅผ ์ํํ๋๊ฐ?
์ด ์ง๋ฌธ์ MegaScale-Infer ๋ ผ๋ฌธ์ ํต์ ์ต์ ํ ์ค๊ณ์ ๋ณธ์ง์ ์ฐ๋ฅด๋ ํต์ฌ์ ๋๋ค. ์ด ํญ๋ชฉ์ ๋จ์ํ bandwidth ๊ฐ์ ์ด ์๋๋ผ, MoE ์ถ๋ก ์ ๊ตฌ์กฐ์ ํต์ ๋ณ๋ชฉ์ ํด๊ฒฐํ๊ธฐ ์ํ ์ ์ฉ ์ค๊ณ๋ผ๋ ์ ์์ ๋งค์ฐ ์ค์ํฉ๋๋ค.
๐ฅ 1. ๊ธฐ์กด All2All (NCCL ๊ธฐ๋ฐ)์ ๋ณ๋ชฉ ๋ฌธ์
โ ๋ฐฐ๊ฒฝ: MoE์์๋ token routing์ด ํ์ํจ
- ๊ฐ token์ top-k expert๋ก๋ง ๋ถ์ฐ๋จ (์: top-2 of 8)
- ๋ฐ๋ผ์ attention node โ ์ ํ๋ expert node๋ก ๋น๊ท ์ผํ, sparse ํต์ ๋ฐ์
โ ๊ธฐ์กด NCCL์ ํ๊ณ (๋ ผ๋ฌธ ยง5, Fig. 5, 10, 11)
| ๋ณ๋ชฉ | ์ค๋ช |
|---|---|
| GPU-to-CPU ๋ณต์ฌ | NCCL์ P2P ์ ์ก ์ GPU ๋ฉ๋ชจ๋ฆฌ๋ฅผ CPU proxy๋ก ๋ณต์ฌ โ latency ์ฆ๊ฐ |
| Group operation overhead | NCCL์ group call์ 8๊ฐ ๋จ์๋ก batch ์ฒ๋ฆฌ โ N์ด ํด์๋ก queueing delay ์ฌํ |
| Tail latency โ | P99 latency๊ฐ N ์ฆ๊ฐ ์ ๊ธ์ฆ โ pipeline ์ ์ฒด latency ๋์ด๋จ |
| ACK ์ฒ๋ฆฌ ์ง์ฐ | round-robin QoS๋ก ์ธํด ACK packet์ด ์ง์ฐ๋์ด sender stall ๋ฐ์ |
| GPU sync overhead | NCCL์ ๋ด๋ถ์ ์ผ๋ก GPU sync ์ฐ์ฐ์ ์๊ตฌํจ โ multi-GPU ์ํฉ์์ ๋นํจ์จ ์ ๋ฐ |
๐ 2. MegaScale-Infer์ M2N ํต์ ๊ตฌ์กฐ: ํด๊ฒฐ์ฑ
โ ๊ตฌ์กฐ์ ์ ํ: All2All โ M2N
- All2All: ๋ชจ๋ ๋ ธ๋๊ฐ ๋ชจ๋์๊ฒ ๊ท ๋ฑํ๊ฒ ์ ์ก
- M2N: Attention M๊ฐ โ Expert N๊ฐ๋ก ์ ํ์ , ๋ถ๊ท ์ผํ ์ ์ก
- โ MoE ํ ํฐ ๋ผ์ฐํ ๊ตฌ์กฐ์ ๋ ์ ํฉํ ํต์ ํจํด
๐๏ธ 3. M2N Sender/Receiver Path ๊ตฌ์ฑ
๐ค Sender Side ๊ตฌ์ฑ (๋ ผ๋ฌธ Figure 6)
| ๊ตฌ์ฑ ์์ | ์ญํ |
|---|---|
| Compute Kernel | ์ด์ GEMM์ด ๋๋ฌ๋์ง ์ฒดํฌ (stream ๋น์ฐจ๋จ) |
| Send-control Kernel | send flag ์ธํ , ๋ฐ์ดํฐ ์ ์ก ์กฐ๊ฑด ํ๋จ |
| Core Sender (CPU) | RDMA write with immediate + GPUDirect๋ก ๋ฐ์ดํฐ ์ง์ ์ ์ก |
| QPs (Queue Pairs) | ์์ ๋์ expert N๋ช ์ ๋ํด ๊ฐ๊ฐ ๊ตฌ์ฑ๋จ |
| Poll Completion Queue | ์ ์ก ์๋ฃ ์ฌ๋ถ ํ์ธ |
โจ ์ต์ ํ ๊ธฐ์
- GPUDirect RDMA: GPU memory โ NIC โ RDMA ์ง์ ์ ์ก (CPU ํต๊ณผ ์์)
- Zero Copy: host bounce buffer ์๋ต โ throughput ์ฆ๊ฐ
- ACK ์ฐ์ ์ ์ก: ACK packet์ ๋ณ๋ high-priority queue์ ํ ๋น โ ๋น ๋ฅด๊ฒ ์์ ์๋ฃ ์๋ฆผ
๐ฅ Receiver Side ๊ตฌ์ฑ (๋ ผ๋ฌธ Figure 7)
| ๊ตฌ์ฑ ์์ | ์ญํ |
|---|---|
| Recv-control Kernel | RDMA buffer์ ์ฐ์ธ ๋ฐ์ดํฐ ๋ชจ๋ํฐ๋ง |
| Core Receiver | RDMA polling ๋ฐ optional copy ์ํ |
| Copy Kernel | pre-registered buffer โ output tensor layout ๋ณต์ฌ |
| Poll CQ | ์์ ์๋ฃ ์ํ ์ถ์ |
| Auto post recv | ๋ค์ RDMA ์์ ์ฉ ๋ฒํผ ์๋ ๋ฑ๋ก (no delay) |
๐ 4. ์ฑ๋ฅ ๊ฐ์ ์์น ์์ฝ
| ์งํ | MegaScale vs NCCL |
|---|---|
| Median Latency (256KB) | 68.2% โ |
| P99 Latency (256KB) | 92.9% โ |
| Throughput (256KB) | 4.2ร โ |
| Sender/Receiver ์ ์ฆ๊ฐ ์ (M, N = 32) | Tail Latency ์ต๋ 96.9% โ, Throughput 5.8ร โ |
โจ 5. ์ถ๊ฐ ์ต์ ํ ๊ธฐ๋ฒ
๐ High-priority ACK Queue
- ์์ธ: bidirectional ํต์ ์ ACK๊ฐ ์ผ๋ฐ data packet๊ณผ ๋์ผ ์ฐ์ ์์๋ก ์ ์ก๋จ โ ์ง์ฐ ๋ฐ์
- ํด๊ฒฐ: ACK ์ ์ฉ high-priority queue ๊ตฌ์ฑ + weight tuning์ผ๋ก ์๋ต ์ง์ฐ ๋ฐฉ์ง
๐ Congestion Control Fine-Tuning
- ๋ฌธ์ : ์ผ๋ถ expert์ ํ ํฐ์ด ๋ชฐ๋ฆด ๊ฒฝ์ฐ queueing latency ํญ์ฆ
- ํด๊ฒฐ: ์ ์ก ์๋ ์ ์ด ์๊ณ ๋ฆฌ์ฆ์ ์์ ์ฒ๋ณ queue length ๊ณ ๋ ค
๐ ์์ฝ
| ํญ๋ชฉ | ์ค๋ช |
|---|---|
| ๊ธฐ์กด ๋ณ๋ชฉ | GPU-to-CPU copy, group op overhead, tail latency, sync stall |
| ํต์ฌ ๊ตฌ์กฐ ์ ํ | All2All โ M2N, attention-to-expert direct routing |
| ๊ธฐ์ ์ ํด๊ฒฐ์ฑ | RDMA + GPUDirect + ACK priority + polling CQ |
| ์ฑ๋ฅ ํฅ์ | ์ต๋ 4.2ร throughput, P99 latency 96.9% ๊ฐ์ |
๐ก ๋ด ์๊ฒฌ
MegaScale-Infer์ M2N ํต์ ๊ตฌ์กฐ๋ ๋จ์ํ โ๋น ๋ฅธ ํต์ ๋ผ์ด๋ธ๋ฌ๋ฆฌโ๊ฐ ์๋๋ผ, MoE ์ถ๋ก ์ sparse routing ๊ตฌ์กฐ์ ๊ตฌ์กฐ์ ์ผ๋ก ๋ง์ถคํ ์ค๊ณ๊ฐ ๋์๋ค๋ ์ ์์ ํ์ ์ ์ ๋๋ค. ํนํ attention node์ ์๊ฐ ๋ง์์ง์๋ก ์ด ๊ตฌ์กฐ๋ ๋์ฑ ๊ฐ๋ ฅํ ํ์ฅ์ฑ๊ณผ ์์ ์ฑ์ ๋ณด์ฌ์ค ์ ์์ต๋๋ค.
๋๊ธ