Compare AIFind AIAI NewsAI How-To
About Us
PrivacyTermsFAQContactContact
AIB Inc.Company info
© 2026 AIB Inc.

Optimizing Ling-2.6-1T Inference on TPU v7x

Optimizing Ling-2.6-1T Inference on TPU v7x

LMSYS·Thursday, June 18, 2026
  • •SGLang-JAX optimization reduces Ling-2.6-1T MoE prefill latency by 53% to 2.42 ms on TPU v7x.
  • •Fused MoE V2 kernel achieves 1.29×–1.77× the output throughput of 16 H200 GPUs in decode benchmarks.
  • •Optimization targets HBM and ICI data bottlenecks by overlapping memory movement with MXU compute windows.
  • •SGLang-JAX optimization reduces Ling-2.6-1T MoE prefill latency by 53% to 2.42 ms on TPU v7x.
  • •Fused MoE V2 kernel achieves 1.29×–1.77× the output throughput of 16 H200 GPUs in decode benchmarks.
  • •Optimization targets HBM and ICI data bottlenecks by overlapping memory movement with MXU compute windows.

InclusionAI's Ling-2.6-1T model is now optimized for TPU v7x hardware using SGLang-JAX, a serving system that manages large language model deployment. The update addresses the Mixture-of-Experts (MoE) mechanism—a sparse architecture where only a subset of parameters is active per token—as the primary performance bottleneck. By replacing the existing MoE kernel with a new Pallas-fused version labeled Fused MoE V2, engineers reduced MoE prefill latency from 5.16 ms to 2.42 ms, a 53% improvement. In standard decode benchmarks, 16 TPU v7x chips achieve between 1.29× and 1.77× the output throughput of 16 H200 GPU equivalents.

The Fused MoE V2 kernel optimizes performance by scheduling data movement and computation in parallel. Previous implementations (V1) often fragmented hidden dimensions, causing frequent, small data transfers between HBM (high-bandwidth memory) and the chip’s local scratchpad, VMEM. By keeping routed token data and accumulators resident in VMEM and using double buffering for expert weight staging, V2 eliminates the need to spill partial results back to external memory. This architectural adjustment ensures that HBM-DMA (direct memory access) and ICI (inter-chip interconnect) tasks occur behind the Matrix Multiply Unit (MXU) compute window.

Beyond the MoE kernel overhaul, the Ling-2.6-1T deployment incorporates a hybrid memory pool for KV (key-value) and recurrent states, Gated Linear Attention (GLA) mechanisms, and single-controller data parallelism. The model itself is a 1T sparse MoE with 63B activated parameters per token, featuring 256 routed experts and one shared expert. The project provides a reproduction path for performance profiling using jax.profiler device traces, highlighting how manual scheduling of dependencies across TPU hardware units—MXU, VPU, VMEM, and ICI-DMA—remains essential for saturating high-bandwidth sparse model inference.

InclusionAI's Ling-2.6-1T model is now optimized for TPU v7x hardware using SGLang-JAX, a serving system that manages large language model deployment. The update addresses the Mixture-of-Experts (MoE) mechanism—a sparse architecture where only a subset of parameters is active per token—as the primary performance bottleneck. By replacing the existing MoE kernel with a new Pallas-fused version labeled Fused MoE V2, engineers reduced MoE prefill latency from 5.16 ms to 2.42 ms, a 53% improvement. In standard decode benchmarks, 16 TPU v7x chips achieve between 1.29× and 1.77× the output throughput of 16 H200 GPU equivalents.

The Fused MoE V2 kernel optimizes performance by scheduling data movement and computation in parallel. Previous implementations (V1) often fragmented hidden dimensions, causing frequent, small data transfers between HBM (high-bandwidth memory) and the chip’s local scratchpad, VMEM. By keeping routed token data and accumulators resident in VMEM and using double buffering for expert weight staging, V2 eliminates the need to spill partial results back to external memory. This architectural adjustment ensures that HBM-DMA (direct memory access) and ICI (inter-chip interconnect) tasks occur behind the Matrix Multiply Unit (MXU) compute window.

Beyond the MoE kernel overhaul, the Ling-2.6-1T deployment incorporates a hybrid memory pool for KV (key-value) and recurrent states, Gated Linear Attention (GLA) mechanisms, and single-controller data parallelism. The model itself is a 1T sparse MoE with 63B activated parameters per token, featuring 256 routed experts and one shared expert. The project provides a reproduction path for performance profiling using jax.profiler device traces, highlighting how manual scheduling of dependencies across TPU hardware units—MXU, VPU, VMEM, and ICI-DMA—remains essential for saturating high-bandwidth sparse model inference.

Read original (English)·Jun 17, 2026
Infra#tpu#moe#sglang#jax#inference#throughput#latency