Google Releases ML Drift Edge GPU Engine
- •Google released open-source ML Drift for GPU-accelerated on-device inference under the Apache 2.0 license.
- •Google reports up to 40% lower YouTube Shorts frame latency and 30% faster Adobe mobile photo features.
- •The legacy TensorFlow Lite GPU delegate will stop receiving feature updates as Google urges migration to LiteRT ML Drift.
Google AI Edge released ML Drift on October 8, 2026, as an open-source GPU compute engine for on-device AI and machine-learning inference. Licensed under Apache 2.0, it supports OpenGL ES, OpenCL, Metal and WebGPU, and is available through LiteRT as an accelerator or as a standalone library. Google says it is already running daily on millions of devices across Chrome, YouTube Shorts, Photos, Meet and AI Edge Gallery.
ML Drift addresses the variety of GPU architectures, driver versions and low-level APIs found on edge devices. Its tensor virtualization separates a tensor’s logical representation from its physical GPU allocation, allowing shared shader templates to serve multiple backends. A custom-operation framework gives developers low-level shader access, and an accompanying agentic SKILL.md guide helps coding agents author, register and verify custom shaders. The LiteRT accelerator also supports 5D tensors, enabling workloads such as 3D convolutional networks and spatiotemporal models including YOLO 11n, MobileViT v2 and Swin Transformer v2.
For large language models, ML Drift changes kernels and layouts between prefill, which processes the input context, and decode, which generates tokens one at a time. Its decoding path uses a convolution-aligned KV cache layout and in-kernel activation quantization to reduce repeated memory transfers. Google reports up to 12% lower memory overhead than other frameworks in Gemma benchmarks. In desktop previews, the same WebGPU code runs natively on Windows and Linux through Dawn, while macOS uses Metal; the stated primary focus remains mobile and edge devices.
Google cites customer results including up to 40% lower average frame latency for YouTube Shorts segmentation effects on Android and iOS, and up to a 2 sec speedup for Google Photos editing compared with the legacy GPU delegate. Chrome uses ML Drift for on-device Gemini Nano models in its Prompt, Summarize and Writer APIs. Adobe says mobile Lightroom and Photoshop features run up to 30% faster on-device, while Snap reports 30% lower model latency for face and style effects on Android. Snap also says the accelerator expands support to larger diffusion models.
Google worked with Arm on Mali and Immortalis GPU kernels, Intel on WebGPU and Xe Matrix Extensions for Core Ultra processors with Xe3 graphics, and Qualcomm on OpenCL kernels for Adreno GPUs. Google says the legacy TensorFlow Lite GPU delegate will receive no new feature updates and encourages migration to the backwards-compatible LiteRT ML Drift GPU accelerator. Standalone LiteRT packages already offer ML Drift for Android developers using unbundled runtimes; availability through LiteRT in Google Play Services is coming soon. Developers can use the LiteRT accelerator or build custom compute graphs with ML Drift’s standalone C++ APIs.