Baseten Maps LLM Inference Frontier
- •Baseten frames LLM inference as latency-throughput efficient frontier problem for production deployments
- •Batch sizing, TP, EP, ADP, and quantization manage latency, throughput, quality, and cost tradeoffs
- •Kernel optimization, speculative decoding, and P/D disaggregation push systemwide inference efficiency outward
Baseten published a September 2, 2026 engineering guide that frames LLM inference as an “efficient frontier” problem, where teams trade latency, throughput, cost, and sometimes model quality when serving models such as GLM-5.3 or Kimi K3 for agentic coding. Philip Kiely writes that inference engineers use 2 categories of techniques: configurations that move a deployment along a latency-throughput frontier, and improvements that push the whole frontier outward so gains can be spent on lower latency, higher throughput, or both.
The guide says production tuning often depends less on a novel method than on finding the right configurations for a traffic pattern. Batch sizing is the clearest tradeoff: small batches give strong per-user latency but generate fewer total tokens per GPU, raising cost per token, while larger batches worsen per-user latency and improve total throughput. The article adds that token-level continuous batching avoids delay from waiting for batches to start, but the configured batch size still shapes both user latency and system throughput.
Parallelism strategy is another frontier-moving choice because today’s LLMs can contain hundreds of billions or trillions of parameters and must be spread across multiple GPUs. Tensor Parallelism, or TP, can lower latency despite expensive all-to-all communication because high-bandwidth NVLink interconnects make those operations fast. Expert Parallelism, or EP, can support latency or throughput depending on degree: lower EP often improves latency, while wide EP, including EP across a full rack of GPUs, generally supports higher throughput. Attention Data Parallelism, or ADP, replicates attention layers for parallel computation and increases throughput at the cost of per-request speed.
Quantization changes the picture by improving both latency and throughput when a model runs at lower precision in weights, activations, or KV cache values. The article says quantization pushes out the serving frontier but creates a separate quality-versus-efficiency tradeoff, with “little-to-no reduction” in model quality possible when using microscaling floating-point formats such as MXFP4 and NVFP4.
Baseten lists kernel optimization, runtime improvements, speculative decoding, and P/D disaggregation as techniques that push the whole frontier outward. CUDA kernel optimization reduces resources needed for each generated token. Speculative decoding guesses possible tokens and validates them; newer methods such as EAGLE-3, DSpark, and DFlash still compete for resources but can skip forward passes and increase tokens per second per user, especially for code generation. P/D disaggregation separates prefill and decode onto dedicated workers so high-volume deployments can tune worker ratios to input length, output length, and cache hit rates.