Compare AIFind AIAI NewsAI How-To
About Us
PrivacyTermsFAQContactContact
AIB Inc.Company info
© 2026 AIB Inc.

OpenLake Leads MLPerf Storage v3.0

OpenLake Leads MLPerf Storage v3.0

theopenlake.com·Sunday, September 6, 2026
  • •OpenLake reports 6.72 GiB/s writes and 11.55 GiB/s reads in MLPerf Storage v3.0
  • •Llama 3.1 8B checkpoint test used eight processes, 16 files, 10 writes, and 10 reads
  • •Submitted system used one client node, 400 Gb/s InfiniBand, and an NVMe-backed OpenLake gateway
  • •OpenLake reports 6.72 GiB/s writes and 11.55 GiB/s reads in MLPerf Storage v3.0
  • •Llama 3.1 8B checkpoint test used eight processes, 16 files, 10 writes, and 10 reads
  • •Submitted system used one client node, 400 Gb/s InfiniBand, and an NVMe-backed OpenLake gateway
  • •OpenLake reports 6.72 GiB/s writes and 11.55 GiB/s reads in MLPerf Storage v3.0
  • •Llama 3.1 8B checkpoint test used eight processes, 16 files, 10 writes, and 10 reads
  • •Submitted system used one client node, 400 Gb/s InfiniBand, and an NVMe-backed OpenLake gateway
  • •OpenLake reports 6.72 GiB/s writes and 11.55 GiB/s reads in MLPerf Storage v3.0
  • •Llama 3.1 8B checkpoint test used eight processes, 16 files, 10 writes, and 10 reads
  • •Submitted system used one client node, 400 Gb/s InfiniBand, and an NVMe-backed OpenLake gateway

OpenLake said on September 1, 2026, that its Infinity Core I/O Engine delivered 6.72 GiB/s writes and 11.55 GiB/s reads in the MLPerf Storage v3.0 Llama 3.1 8B checkpointing benchmark. MLCommons released the MLPerf Storage v3.0 results, and OpenLake said it achieved the highest read and write bandwidth among published Closed division S3 results for Llama 3.1 8B checkpointing with one client node and one data parallel instance.

The comparison covered five published results using the same Closed division checkpointing workload, S3 API interface, Llama 3.1 8B model, one client node, and one data parallel instance, although the article said the underlying storage configurations differ. OpenLake reported 6.72 GiB/s write bandwidth, 29.42 seconds write duration, 11.55 GiB/s read bandwidth, and 9.37 seconds read duration. NVIDIA AIStore, 6 node, posted 3.40 GiB/s writes, 30.83 s write duration, 11.08 GiB/s reads, and 9.47 s read duration; NVIDIA AIStore, 12 node, posted 3.20 GiB/s, 33.04 s, 8.33 GiB/s, and 12.90 s; NVIDIA AIStore, 3 node, posted 3.02 GiB/s, 34.67 s, 6.99 GiB/s, and 15.00 s; Nebius Object Storage posted 2.81 GiB/s, 37.24 s, 7.14 GiB/s, and 14.67 s.

MLPerf Storage is a reproducible, architecture-neutral benchmark for storage systems under representative AI training and inference workloads. Its checkpointing workload emulates the Llama 3 family from 8B to 1.25 trillion parameters. In the Llama 3.1 8B configuration, eight processes write and read a checkpoint through 16 files, with 10 checkpoint writes followed by 10 reads, then average performance is reported. The test estimates storage behavior at full load when training processes save model state or restore it after a failure.

OpenLake said the submitted system used one client node over a 400 Gb/s InfiniBand network connected to an NVMe-backed OpenLake gateway. The Infinity Core I/O Engine used asynchronous io_uring I/O (nonblocking Linux storage operations), pinned execution threads, fine-grained I/O coalescing, XFS and workload-specific storage tuning, and an S3 API interface for checkpoint writes and reads.

The article said checkpoint performance matters because large-scale model training can run for weeks across hundreds or thousands of GPUs, while checkpoints can reach hundreds of gigabytes or several terabytes. Synchronous checkpointing pauses training until data is written durably, so write speed affects idle GPU time and read speed affects recovery after failure. OpenLake said it plans to expand benchmark coverage across KV cache offloading, checkpointing, training, vector search, and context storage.

OpenLake said on September 1, 2026, that its Infinity Core I/O Engine delivered 6.72 GiB/s writes and 11.55 GiB/s reads in the MLPerf Storage v3.0 Llama 3.1 8B checkpointing benchmark. MLCommons released the MLPerf Storage v3.0 results, and OpenLake said it achieved the highest read and write bandwidth among published Closed division S3 results for Llama 3.1 8B checkpointing with one client node and one data parallel instance.

The comparison covered five published results using the same Closed division checkpointing workload, S3 API interface, Llama 3.1 8B model, one client node, and one data parallel instance, although the article said the underlying storage configurations differ. OpenLake reported 6.72 GiB/s write bandwidth, 29.42 seconds write duration, 11.55 GiB/s read bandwidth, and 9.37 seconds read duration. NVIDIA AIStore, 6 node, posted 3.40 GiB/s writes, 30.83 s write duration, 11.08 GiB/s reads, and 9.47 s read duration; NVIDIA AIStore, 12 node, posted 3.20 GiB/s, 33.04 s, 8.33 GiB/s, and 12.90 s; NVIDIA AIStore, 3 node, posted 3.02 GiB/s, 34.67 s, 6.99 GiB/s, and 15.00 s; Nebius Object Storage posted 2.81 GiB/s, 37.24 s, 7.14 GiB/s, and 14.67 s.

MLPerf Storage is a reproducible, architecture-neutral benchmark for storage systems under representative AI training and inference workloads. Its checkpointing workload emulates the Llama 3 family from 8B to 1.25 trillion parameters. In the Llama 3.1 8B configuration, eight processes write and read a checkpoint through 16 files, with 10 checkpoint writes followed by 10 reads, then average performance is reported. The test estimates storage behavior at full load when training processes save model state or restore it after a failure.

OpenLake said the submitted system used one client node over a 400 Gb/s InfiniBand network connected to an NVMe-backed OpenLake gateway. The Infinity Core I/O Engine used asynchronous io_uring I/O (nonblocking Linux storage operations), pinned execution threads, fine-grained I/O coalescing, XFS and workload-specific storage tuning, and an S3 API interface for checkpoint writes and reads.

The article said checkpoint performance matters because large-scale model training can run for weeks across hundreds or thousands of GPUs, while checkpoints can reach hundreds of gigabytes or several terabytes. Synchronous checkpointing pauses training until data is written durably, so write speed affects idle GPU time and read speed affects recovery after failure. OpenLake said it plans to expand benchmark coverage across KV cache offloading, checkpointing, training, vector search, and context storage.

Read original (English)·Sep 1, 2026
Infra#openlake#mlperf storage#llama 3 1 8b#checkpointing#s3 api#io uring#gpudirect storage#infiniband#nvme#kv cache