Paper

WavePP — A Pipeline-Parallel Prefill Runtime That Burns In Prefix Reuse

TL;DR

Built on top of TensorRT-LLM, WavePP overlaps request admission (reuse consensus, cache protection, capacity reservation) asynchronously with pipeline execution, and dynamically tunes the wave size to keep the pipeline full. On the same kernels and the same topology, it pushed prefill throughput for GLM 5.2 and MiniMax M2.7 up to 2.91×/2.02×, and across the 28 Kimi K3 settings it recorded the highest throughput in all 18 with concurrency 8 or higher.

Core Idea

The real bottleneck in pipeline-parallel (PP) prefill is not GPU compute but “how quickly the next request can be prepared.” In a system where each stage independently maintains and evicts its KV cache, a local cache hit does not guarantee global reusability. WavePP’s central claim boils down to one sentence: “If you move request admission off the execution path and eliminate the moments when the GPU sits idle, you can substantially raise pipeline prefill throughput with the same kernels.” Three pillars support this claim: fast reuse hint, local lease, and capacity escrow.

Background: The Problem They Solved

As agentic workloads became widespread, LLMs process increasingly long inputs such as conversations, documents, and tool outputs (source: §1). The cost of the prefill stage surged accordingly, and in disaggregated serving, architectures that dedicate prefill-only workers became common.

As models grew too large to fit on a single GPU, three main parallelization strategies emerged. Tensor parallelism (TP) splits layers across devices, pipeline parallelism (PP) divides layers into a sequence of stage groups, and context parallelism (CP) splits the input length. PP is favorable for throughput because multiple microbatches pass through the stages in sequence and execute concurrently, but it has the structural limitation that the slowest stage dominates the whole (source: §3.5).

Once prefix caching gets involved, the problem becomes more complex. Reusing the cache state of a shared prefix can cut prefill computation by several times. For example, if 90% of a 100K context is in cache, only the last 10K tokens need to be processed (source: §3.2). In PP, however, each stage keeps its own local cache for its layers. Moreover, hybrid models like Kimi K3 have both the dense KV cache of MLA attention and a sparse recurrent snapshot of KDA (linear attention), and because the recurrent state cannot be restored by rewinding $S_r$ from $S_{r'}$ ($r'>r$), the reuse boundary is tied to the intersection of the two caches (source: §3.2, Fig. 1).

Ultimately, the research gap the authors point to is this: existing PP runtimes perform request preparation (reuse-boundary consensus → cache pin → suffix capacity acquisition → metadata construction) serially on the execution path. When this preparation delays the next forward pass, the pipeline idles even though the GPU still has work available. Tree traversal requires a mutex, and if the reuse boundary diverges across stages, a second traversal (re-calibration) becomes necessary, making the cost even larger (source: §4.1).

New Approach: WavePP

WavePP is a prefill runtime that decouples this preparation work from execution and lets stages follow a single execution schedule while managing their caches independently. The core design is as follows (source: §4).

1) Two threads and two communication planes. Each rank has two threads. The WAVE thread handles request dispatch, reuse consensus, and capacity management, while the LOOP thread handles the executor and cache block materialization. Communication is likewise split into two planes: the data plane carries the wave schedule and activations, and the rounds plane carries requests and admission results (source: §4.2).

2) Fast reuse hint. To reduce traversal cost and mutex contention, it maintains a hint index that holds no blocks. For each token block $B_i$, it maintains a chained prefix hash.

$$H_{-1}=0,\qquad H_i=\operatorname{hash}(B_i,\ \mathrm{extraKeys}_i,\ \mathrm{salt},\ H_{i-1})$$

For dense-attention prefixes it reduces probing from $O(N)$ to about $O(\log N)$ using gallop (skip by doubling) + bisect (binary search), while recurrent snapshots are sparse and are found by a reverse scan at the attention boundary. This hint neither pins the cache nor reserves capacity, so there is nothing to roll back even if a pending request is cancelled (source: §4.3). In a stress experiment with 200K tokens and concurrency 64, the median probe time dropped from 64.4 ms → 24 μs, with a p99 of 2.35 ms (source: §4.3).

3) Local lease. For a consensus reuse candidate, it does not install the actual block into the tree but protects only a reference. Other requests may share that state, but it cannot be evicted until the lease is released. If the actual pin points diverge across stages, the richer side rolls back to the common point. Since this reconciliation is monotonically decreasing, it is guaranteed to converge (source: §4.4).

4) Capacity escrow. Even if a lease guarantees reuse, no one can complete if there is no room for the suffix. So each stage reserves suffix capacity immediately after the lease. This reservation does not allocate actual blocks or push out eviction victims; it holds only a reference and defers the actual eviction/offload to the executor’s materialization point (source: §4.5). The eviction victim remains in the tree so other requests can still occupy it via reuse, and at commit time the reference count detects additional owners and selects a replacement block.

5) MPU (Maximum Pipeline Utilization) scheduling. Filling greedily reduces the number of waves and creates idle stages. Filling 2M tokens with a budget $M$ gives 2 waves, but reducing to $M/2$ gives 4 (source: Fig. 6). MPU looks at the remaining work and discretely halves the budget.

$$b_0=M,\qquad b_{j+1}=\max\left(b_{\min},\ \operatorname{alignDown}_{B}\left(b_j/2\right)\right)$$

It estimates the available number of waves as $n(b)=\max\left(W/b,\ N_{>b/2}\right)$ and chooses the largest budget that can supply the execution ring $R=2P$ (source: §4.6).

How It Works: A Concrete Example

Suppose a 200K-token prompt arrives at a 4-stage (PP4) pipeline. 90% is cached, but the actual reusable point differs per stage.

  flowchart LR
    A["Request arrival (200K tokens)"] --> B["Hint lookup: find local reuse point"]
    B --> C["All-rank min: agree on common candidate"]
    C --> D["Lease: pin common prefix + last snapshot"]
    D --> E["Escrow: reserve suffix capacity (commit)"]
    E --> F["Per-stage local materialization"]
    F --> G["Rank 0 plans chunks → wave execution"]
  1. Hint lookup. Each rank probes its own hint index to find its local reuse point. Suppose stage 0 can reuse up to 176K, stage 1 up to 180K, stage 2 up to 164K, and stage 3 up to 172K. The candidate becomes the all-rank minimum, 164K. No cache is pinned yet.
  2. Lease. Each stage pins the 164K prefix and the last recurrent snapshot at that point. Other requests may occupy these blocks but cannot evict them.
  3. Escrow. The remaining suffix is $200\text{K}-164\text{K}=36\text{K}$ tokens. Each stage reserves cache blocks for 36K tokens. It does not actually allocate blocks or push out victims — this is deferred and overlapped with GPU execution.
  4. Materialization. Each stage actually installs the blocks into its local cache tree. Stage 0 can start executing as soon as it is ready, without waiting for the other stages to finish preparing.
  5. Wave execution. Rank 0 splits the 36K suffix into chunks. With $M=16{,}384$ there are 3 waves (16K+16K+4K), but if MPU reduces to $b=8{,}192$, it splits into more waves and fills the 4 stages more evenly.

The key here is the reordering of steps. Existing runtimes perform “consensus → pin → reserve → install” serially when a request is executed, making the pipeline wait. WavePP reorders this to “hint (non-blocking) → consensus → lease/escrow (reference only) → materialization while execution overlaps,” so the next request is prepared in parallel while the GPU executes the previous request (source: Fig. 3).

Performance Validation: Main Results

1) Improvements on the Same Kernels and Same Topology (GLM 5.2 / MiniMax M2.7)

Both models were fixed to NVFP4 weights, 4 B300 GPUs, PP4 topology, and TensorRT-LLM 1.3.0rc26 kernels, with only the runtime swapped (source: §5.2). The result was that WavePP was ahead in 37 of 40 settings.

Workload$c$GLM 5.2 TRT-LLMGLM 5.2 WavePPΔMiniMax TRT-LLMMiniMax WavePPΔ
Short cold3245,73047,864+4.7%49,59654,615+10.1%
Long cold3240,84443,357+6.2%36,78940,165+9.2%
Short reuse(90%)32321,750378,951+17.8%315,446348,107+10.4%
4K suffix64835,2351,160,931+39.0%832,6451,039,376+24.8%
4K suffix128416,2941,212,572+191.3%551,1431,113,196+102.0%

(source: Table 1, units input tokens/s, including cached tokens)

The most dramatic point is short suffix + high concurrency. The faster requests finish, the faster the runtime must prepare new requests, and here admission latency directly erodes throughput. At $c=128$, WavePP achieved 2.91× (GLM) and 2.02× (MiniMax) throughput, and for the 4K suffix workload the median prefill completion time fell by about 71% on GLM and about 55% on MiniMax (source: §5.2).

Notably, the reverse case also exists. For long-prefix reuse at $c=8$ and 4K suffix MiniMax at $c=16$, WavePP was actually lower. The authors explain this as “a regime where GPU compute dominates and the admission savings are relatively small” (source: §5.2).

2) Kimi K3: TP vs PP, and Cross-Library Comparison

On Kimi K3, a 2.8T-parameter hybrid KDA/MLA MoE model, using 8 GB300 GPUs (two 4-GPU nodes), they compared WavePP (TP1×PP8) against TRT-LLM/SGLang/vLLM’s TP8/EP8 and PP8 (source: §5.3). WavePP recorded the highest throughput in 21 of 28 settings, and in all 18 with concurrency 8 or higher (source: §5.4).

  • Cold long-context: Geometric-mean throughput versus TRT-LLM TP8/EP8 was +58.8% / +49.1% / +37.8% at 65.5K/131K/262K respectively. At $c=32$, 44,643 / 35,878 / 28,031 input tok/s (source: §5.5).
  • Seeded reuse(90%): 289,096 tok/s at 131K $c=32$, and 218,743 tok/s at 262K $c=32$. Consistently ahead in both throughput and p50 completion time at concurrency 8 or higher (source: §5.6).
  • Mixed (75% warm-131K / 25% cold-262K): Geometric mean +46.1~+56.8% versus TP, and +5.0% (SGLang)/+14.6% (vLLM) versus PP. 61,190 tok/s at $c=32$ (source: §5.7).

An interesting observation is that PP is more advantageous than TP for prefill. SGLang PP8 had +26.8% higher geometric-mean throughput than SGLang TP8/EP8, and WavePP went one step further (source: §5.4).

3) Cache Retention and Ablation

In the prefix retention experiment, WavePP’s good-hit rate (reusing 80% or more of the expected prefix) was 88~100%, ahead of TP8/EP8’s 73~83% at every arrival rate (source: §5.8). This is because decoupling admission from execution secures reuse before the cache is pushed out by new requests.

The ablation isolates the contribution of each mechanism (source: §5.9, Fig. 14).

RungAdded mechanism
R1Naive reimplementation (lockstep preparation)
R2+ wave admission/transport
R3+ lease admission
R4+ MPU (planner packing, adaptive sizing, deadline)
  • Cold 131K $c=16$: From R1→R2, 8,826 → 29,692 tok/s (+236.4%). Changes afterward were marginal. MPU reduced p50 from 69.69→43.03 seconds but throughput improved only +1.4%.
  • Reuse 131K $c=8$: lease/escrow (R3) gave 96,796 → 149,021 (+54.0%), and MPU (R4) gave → 211,214 (+41.7%).

In other words, admission/transport is decisive for cold, while cache coordination and scheduling are decisive for reuse. That said, MPU does not always win: at reuse 262K $c=32$, MPU lowered throughput by -9.7% (still +24.5% higher than TP8/EP8). This amounts to an admission that planner-driven packing can be worse than static packing for certain workloads.

Our Perspective: Strengths, Limitations, and Why This Work Matters

Strengths. The paper’s biggest contribution is its redefinition of the problem. Viewing prefill throughput optimization not as a “kernel compute acceleration” problem but as an “admission cadence” problem, and offering the conceptual framework hint → lease → escrow → materialization, is directly transplantable into real serving-system design. Also, by comparing with only the runtime swapped under the same kernels and same topology, the experimental design cleanly separates kernel performance from scheduling performance. Including a reproducibility check for SGLang PP8 (+5.7% versus the day-0 report) also raises confidence.

Limitations. First, MPU is a discrete budget adjustment that relies only on token counts without estimating the actual chunk execution time, so it cannot account for the fact that chunks of the same size differ in cost depending on the attention prefix length. Indeed, it caused regressions in some settings. Second, it validated only TP1×PP4·PP8, leaving the behavior under TP2×PP4, PP16, or imbalanced layer placement unknown. Third, it handled only FCFS scheduling and a single replica, leaving practical needs such as priority, deadlines, and cache-aware routing unaddressed. Finally, this is an implementation confined to TensorRT-LLM, so portability to other runtimes (vLLM/SGLang) is unproven (source: §6).

Why it matters. As agentic workloads grow longer and hybrid (attention + recurrent) models multiply, “how to consistently agree on the reuse boundary across multiple cache types” will become an increasingly important problem. By presenting both a clear principle — separating execution from preparation — and measurable gains (up to 2.91×) for this distributed admission problem, WavePP is a body of work that can serve as a reference point for both serving-system researchers and engineers.

What’s Next?: The Road Ahead

In addition to the directions the authors themselves propose (source: §6), reasonable follow-up work can be summarized as follows.

  1. Dynamic chunking based on execution-time estimation. Instead of MPU’s discrete budget, determining the size by observing per-chunk execution time and stage progress can improve pipeline balance while reducing regressions.
  2. Diverse topologies and imbalanced placement. It is necessary to verify how admission overhead and overlap change under TP2×PP4, PP16, and inter-stage layer imbalance.
  3. Scheduling that accounts for priority, deadlines, and cache state. Policies that go beyond FCFS to optimize tail latency and SLOs are a natural extension.
  4. Multi-replica + cache-aware routing. Combining cache-aware routing and SLO curves across multiple prefill workers would allow validation of tail-latency improvements in real deployments.
  5. Porting to other runtimes. Applying the hint/lease/escrow framework to other backends such as vLLM and SGLang to confirm generalizability is also a worthwhile follow-up study.

Tables from the paper

Tables converted mechanically from the arXiv e-print LaTeX source. The numbers are the paper’s own and did not pass through a model.

Table 1. PP4 comparison on matched kernels. Aggregate prompt throughput (tokens/s, including cached tokens) is the median of three boots per system and model. Each system uses one configuration throughout. Throughput change $\Delta$ is $(\mathrm{WavePP}/\text{TRT-LLM}-1)\times100\%$. Green shading marks the higher throughput. Short and long reuse use 90% cached prefixes. The 4K-suffix cells follow cold traffic and retain its cache-pool eviction history.

Family$c$GLM 5.2 TRT-LLMGLM 5.2 WavePPGLM 5.2 $\Delta$MiniMax M2.7 TRT-LLMMiniMax M2.7 WavePPMiniMax M2.7 $\Delta$
Short cold846,65847,552+1.9%50,67754,457+7.5%
1646,20748,071+4.0%50,12754,642+9.0%
3245,73047,864+4.7%49,59654,615+10.1%
Long cold841,83943,456+3.9%37,14640,388+8.7%
1641,61043,310+4.1%37,05140,363+8.9%
3240,84443,357+6.2%36,78940,165+9.2%
Short reuse8309,301309,762+0.1%283,199303,921+7.3%
16351,119391,338+11.5%344,174375,368+9.1%
32321,750378,951+17.8%315,446348,107+10.4%
Long reuse8316,516278,418-12.0%250,815206,548-17.6%
16297,822310,786+4.4%239,657241,150+0.6%
32264,520306,497+15.9%227,220236,459+4.1%
Mixed889,84492,759+3.2%91,27099,810+9.4%
1687,97493,515+6.3%91,314100,117+9.6%
3284,38792,080+9.1%89,42799,039+10.7%
4K suffix8497,230503,184+1.2%435,088441,314+1.4%
16658,662661,230+0.4%642,896606,764-5.6%
32884,810895,482+1.2%872,496886,305+1.6%
64835,2351,160,931+39.0%832,6451,039,376+24.8%
128416,2941,212,572+191.3%551,1431,113,196+102.0%

Table 2. Configurations of the Kimi K3 comparison baselines.

BaselineParallel configurationBenchmark configuration
vLLM TP8/EP8TP8/EP8 on eight GPUsv0.28.0 with the FlashInfer TensorRT-LLM MoE path
SGLang TP8/EP8TP8/EP8 on eight GPUsSame SGLang build, 16K prefill budget, and hybrid host-cache configuration as the SGLang PP8 arm
TRT-LLM TP8/EP8TP8/EP8 on eight GPUsSame Kimi K3 engine family and shared kernel changes as WavePP; valid two-node tensor-parallel placement
vLLM PP8Eight-stage pipeline parallelismv0.28.0 with a 16K packed-token budget and prefix caching enabled
SGLang PP8Eight-stage pipeline parallelism16K prefill chunks; 128 GiB/rank hybrid MLA+Mamba host cache; write-through offload, kernel I/O, and page-first loading

Table 3. Kimi K3 workloads and offered concurrency. Input lengths are abbreviated in decimal thousands; the vLLM PP8 262K exception is described in the text.

FamilyInput tokensShared prefixRequest compositionConcurrency
Cold 65.5K65,5360%cold1, 4, 8, 16, 32
Cold 131K131,0720%cold1, 4, 8, 16, 32
Cold 262K262,1440%cold1, 4, 8, 16, 32
Reuse 131K131,07290%seeded shared prefix1, 4, 8, 16, 32
Reuse 262K262,14490%seeded shared prefix1, 4, 8, 16, 32
Mixed––75/25 warm-131K/cold-262K8, 16, 32

Table 4. Cold aggregate prompt throughput (input tok/s). Positive deltas indicate higher throughput for WavePP.

Context, $c$vLLM TP8/EP8SGLang TP8/EP8TRT-LLM TP8/EP8vLLM PP8SGLang PP8WavePP$\Delta_{\mathrm{TP}}$$\Delta_{\mathrm{PP}}$
65.5K, 123,33020,37822,97620,36116,69122,089-5.3%+8.5%
65.5K, 423,92222,21523,52441,55439,92637,467+56.6%-9.8%
65.5K, 823,91722,19423,49041,98040,45843,007+79.8%+2.4%
65.5K, 1623,92522,21523,53742,24240,85244,746+87.0%+5.9%
65.5K, 3223,92522,21323,53642,45641,18044,643+86.6%+5.2%
131K, 121,43019,81021,38023,03021,01420,699-3.4%-10.1%
131K, 422,05920,67321,61933,08933,66235,765+62.1%+6.2%
131K, 822,05820,66221,55433,24233,87336,030+63.3%+6.4%
131K, 1622,05220,67021,67733,32233,98436,129+63.8%+6.3%
131K, 3222,07020,66921,69133,41034,10235,878+62.6%+5.2%
262K, 118,77217,78518,93821,47921,75421,065+11.2%-3.2%
262K, 418,98918,13819,22025,22427,04127,924+45.3%+3.3%
262K, 818,98718,13319,23025,27127,11027,988+45.5%+3.2%
262K, 1618,99018,13819,24925,29627,14427,971+45.3%+3.0%
262K, 3218,99518,11819,25425,32127,17928,031+45.6%+3.1%

Table 5. Queue-inclusive completion time for cold prefill, in seconds, with WavePP deltas against the lowest-latency TP8/EP8 and PP8 results in each cell. Negative deltas indicate lower completion time for WavePP.

Context, $c$vLLM TP8/EP8 p50vLLM TP8/EP8 p95SGLang TP8/EP8 p50SGLang TP8/EP8 p95TRT-LLM TP8/EP8 p50TRT-LLM TP8/EP8 p95vLLM PP8 p50vLLM PP8 p95SGLang PP8 p50SGLang PP8 p95WavePP p50WavePP p95$\Delta_{\mathrm{TP}}$ p50$\Delta_{\mathrm{TP}}$ p95$\Delta_{\mathrm{PP}}$ p50$\Delta_{\mathrm{PP}}$ p95
65.5K, 12.82.83.23.22.92.93.23.23.93.93.03.0+7.1%+7.1%-6.2%-6.2%
65.5K, 410.811.511.811.811.111.16.16.16.36.36.87.1-37.0%-36.0%+11.5%+16.4%
65.5K, 821.622.323.523.822.322.512.212.212.612.611.812.4-45.4%-44.4%-3.3%+1.6%
65.5K, 1643.944.047.147.344.544.524.424.425.125.122.823.2-48.1%-47.3%-6.6%-4.9%
65.5K, 3287.387.994.494.489.089.148.848.850.150.245.445.9-48.0%-47.8%-7.0%-5.9%
131K, 16.16.26.66.66.16.25.75.76.26.26.36.4+3.3%+3.2%+10.5%+12.3%
131K, 424.024.125.325.324.224.215.615.615.315.314.214.8-40.8%-38.6%-7.2%-3.3%
131K, 847.448.150.750.948.449.831.231.330.630.628.529.1-39.9%-39.5%-6.9%-4.9%
131K, 1694.895.5101.3101.696.796.862.462.561.161.156.857.6-40.1%-39.7%-7.0%-5.7%
131K, 32189.6190.3202.9202.9193.2193.5124.9124.9122.2122.2114.4116.3-39.7%-38.9%-6.4%-4.8%
262K, 114.014.014.714.713.814.012.212.212.112.112.412.8-10.1%-8.6%+2.5%+5.8%
262K, 455.255.257.857.854.654.641.341.338.538.537.137.7-32.1%-31.0%-3.6%-2.1%
262K, 8110.4110.5115.6115.8109.0109.182.682.777.077.074.274.8-31.9%-31.4%-3.6%-2.9%
262K, 16220.8220.9231.1231.4217.7218.2165.3165.3153.9153.9149.0149.7-31.6%-31.4%-3.2%-2.7%
262K, 32441.0441.7462.5462.7435.4435.9330.5330.6307.8307.9297.8298.9-31.6%-31.4%-3.2%-2.9%

Table 6. Seeded-reuse aggregate prompt throughput (input tok/s). Positive deltas indicate higher throughput for WavePP.

Context, $c$vLLM TP8/EP8SGLang TP8/EP8TRT-LLM TP8/EP8vLLM PP8SGLang PP8WavePP$\Delta_{\mathrm{TP}}$$\Delta_{\mathrm{PP}}$
131K, 1155,002127,746142,60743,19741,02256,810-63.3%+31.5%
131K, 4167,165172,924183,856130,420134,161124,682-32.2%-7.1%
131K, 8173,271173,266186,690171,714228,012255,406+36.8%+12.0%
131K, 16173,510172,581185,931212,477249,227262,740+41.3%+5.4%
131K, 32173,780172,847186,064219,649253,844289,096+55.4%+13.9%
262K, 1133,140122,545131,38249,23845,12764,674-51.4%+31.3%
262K, 4142,686142,359153,418118,899140,713176,960+15.3%+25.8%
262K, 8142,617142,854152,211139,579180,199190,842+25.4%+5.9%
262K, 16142,661141,581151,376141,134183,182220,899+45.9%+20.6%
262K, 32142,715141,497133,965141,248184,753218,743+53.3%+18.4%

Table 7. Queue-inclusive completion time for seeded reuse, in seconds, with WavePP deltas against the lowest-latency TP8/EP8 and PP8 results in each cell. Negative deltas indicate lower completion time for WavePP.

Context, $c$vLLM TP8/EP8 p50vLLM TP8/EP8 p95SGLang TP8/EP8 p50SGLang TP8/EP8 p95TRT-LLM TP8/EP8 p50TRT-LLM TP8/EP8 p95vLLM PP8 p50vLLM PP8 p95SGLang PP8 p50SGLang PP8 p95WavePP p50WavePP p95$\Delta_{\mathrm{TP}}$ p50$\Delta_{\mathrm{TP}}$ p95$\Delta_{\mathrm{PP}}$ p50$\Delta_{\mathrm{PP}}$ p95
131K, 10.80.91.01.00.90.93.03.03.23.22.32.3+187.5%+155.6%-23.3%-23.3%
131K, 43.13.73.13.82.83.03.64.63.84.74.24.5+50.0%+50.0%+16.7%-2.2%
131K, 86.26.35.67.05.26.95.77.94.16.23.75.5-28.8%-12.7%-9.8%-11.3%
131K, 1612.412.512.012.710.512.38.912.48.09.37.39.6-30.5%-22.0%-8.8%+3.2%
131K, 3223.924.823.824.722.622.817.920.315.916.613.217.0-41.6%-25.4%-17.0%+2.4%
262K, 12.02.02.12.12.02.05.35.35.85.84.14.1+105.0%+105.0%-22.6%-22.6%
262K, 46.88.06.98.17.78.58.510.27.18.45.76.2-16.2%-22.5%-19.7%-26.2%
262K, 814.714.714.815.013.115.214.516.411.212.410.612.7-19.1%-13.6%-5.4%+2.4%
262K, 1629.429.529.530.927.928.428.929.022.422.418.119.3-35.1%-32.0%-19.2%-13.8%
262K, 3258.958.959.360.661.966.057.958.044.144.935.839.7-39.2%-32.6%-18.8%-11.6%

Table 8. Mixed-workload aggregate prompt throughput (input tok/s). Positive deltas indicate higher throughput for WavePP.

$c$vLLM TP8/EP8SGLang TP8/EP8TRT-LLM TP8/EP8vLLM PP8SGLang PP8WavePP$\Delta_{\mathrm{TP}}$$\Delta_{\mathrm{PP}}$
840,79738,98641,94553,12958,07961,063+45.6%+5.1%
1640,79039,06242,06453,55158,46861,436+46.1%+5.1%
3240,84939,07441,69553,67858,34861,190+46.8%+4.9%

License

Author: Jaehun Ryu

Link: https://jaehun.me/en/posts/paper-2609-35263v1/

License: CC BY 4.0

This work is licensed under the Creative Commons Attribution 4.0 International License. You are free to use it for any purpose, including commercial use, as long as you provide proper attribution.

Comments