plot: extend batch-scaling curves with saturated-throughput ceiling
Beyond a mapping's SIP capacity (16/C) the throughput plateaus: extra users run in waves at the saturated rate, they are not un-runnable. Draw a dashed horizontal extension at the ceiling so 1-kv-per-cube (cap 2) reads as saturating, not stopping, at B=2. Caption updated. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Binary file not shown.
Binary file not shown.
|
Before Width: | Height: | Size: 69 KiB After Width: | Height: | Size: 72 KiB |
@@ -271,10 +271,13 @@ long-context batch sweep is 2H work (\S\ref{sec:future}).
|
||||
\si{\micro\second}) versus the number of concurrent decode users on one SIP,
|
||||
each user on a disjoint CUBE group. Each mapping scales almost linearly with
|
||||
$B$ up to its SIP capacity $16/C$---\textsf{1-kv-per-cube} saturates at two
|
||||
users, \textsf{8-kv-per-cube} at sixteen. The dense mapping reaches the
|
||||
highest throughput; the spread mapping is latency-optimal but
|
||||
capacity-limited. Concurrent users overlap to within
|
||||
\SI{3}{}--\SI{5}{\percent} of the single-user latency.}
|
||||
users, \textsf{8-kv-per-cube} at sixteen. Solid segments are measured
|
||||
concurrent runs; the dashed extensions are the saturated-throughput ceiling
|
||||
beyond capacity, where further users run in waves at that sustained rate
|
||||
(so the SIP is full, not idle). The dense mapping reaches the highest
|
||||
ceiling; the spread mapping is latency-optimal but capacity-limited.
|
||||
Concurrent users overlap to within \SI{3}{}--\SI{5}{\percent} of the
|
||||
single-user latency.}
|
||||
\label{fig:gqa-batch}
|
||||
\end{figure}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user