Compare AIFind AIAI NewsAI How-To
About Us
PrivacyTermsFAQContactContact
AIB Inc.Company info
© 2026 AIB Inc.

AWS Adds SageMaker HyperPod Model Caching

AWS Adds SageMaker HyperPod Model Caching

AWS ML Blog·Friday, September 11, 2026
  • •AWS launches model caching for SageMaker HyperPod inference to reduce large-model cold starts
  • •Local NVMe reads at approximately 7 GB/s can replace repeated network downloads
  • •Benchmarks show around 60 percent faster scale-out for 57–145 GB models
  • •AWS launches model caching for SageMaker HyperPod inference to reduce large-model cold starts
  • •Local NVMe reads at approximately 7 GB/s can replace repeated network downloads
  • •Benchmarks show around 60 percent faster scale-out for 57–145 GB models
  • •AWS launches model caching for SageMaker HyperPod inference to reduce large-model cold starts
  • •Local NVMe reads at approximately 7 GB/s can replace repeated network downloads
  • •Benchmarks show around 60 percent faster scale-out for 57–145 GB models
  • •AWS launches model caching for SageMaker HyperPod inference to reduce large-model cold starts
  • •Local NVMe reads at approximately 7 GB/s can replace repeated network downloads
  • •Benchmarks show around 60 percent faster scale-out for 57–145 GB models

AWS launched model caching for Amazon SageMaker Inference on HyperPod on September 10, 2026, to reduce cold starts when deploying large language models for inference. The feature pre-loads model weights and inference server container images onto cluster nodes before pods serve traffic, replacing repeated network downloads with reads from local NVMe storage at approximately 7 GB/s. AWS says pods can typically begin serving traffic in seconds rather than tens of minutes after caching is enabled.

Without caching, a HyperPod inference pod waits for two sequential downloads before handling requests. Kubelet first pulls multi-gigabyte inference server images such as vLLM or LMI from Amazon ECR, a step AWS says takes 5–7 minutes. The inference server then downloads model weights from Amazon S3, Amazon FSx for Lustre, or HuggingFace Hub; a 145 GB model on Amazon S3 can take another 20+ minutes, while a 600+ GB model such as DeepSeek-R1 can take upwards of 30 minutes.

The same delay repeats during scale-out. If a HorizontalPodAutoscaler requests five new pods after a traffic spike, all five pods independently pull images and weights before they can take requests. AWS says autoscaling policies may react in seconds, but the actual added serving capacity can arrive only after 25–30+ minutes because each new pod waits on downloads.

Model caching has two independent options: weights cache and image cache. The weights cache downloads model weights to local NVMe storage on each target node after users add modelCacheConfig with weightsCache enabled to an InferenceEndpointConfig or JumpStartModel resource. The HyperPod Inference Operator creates a ModelDataCacheConfig resource, downloads weights from Amazon S3, Amazon FSx for Lustre, HuggingFace Hub, or JumpStart, labels nodes as cache-ready, and waits for all target nodes before creating the inference deployment.

The image cache pre-pulls the inference server container image onto nodes after users enable imageCache in modelCacheConfig. The operator creates a DaemonSet (one pod per selected node) to pull the image, but unlike the weights cache it creates the inference deployment immediately. Pods with a cached image skip the Amazon ECR pull and save 5–7 minutes; pods that start before caching completes pull from ECR normally.

AWS says both caching modes use preferred scheduling rather than required scheduling. Pods prefer nodes with cached data, but they are not blocked if a scheduler places them on a node without a warm cache during rapid scale-out. In that case, the pod reads weights from the original Amazon S3 or Amazon FSx source and pulls the image from Amazon ECR, matching the normal uncached behavior.

The operator manages caching with two Custom Resource Definitions, or CRDs (Kubernetes extension objects). ModelDataCacheConfig manages model weight downloads, cache-ready labels, cache health, and cleanup when the parent resource is deleted. ModelImageCache manages image pre-pulling, image-ready labels, per-node pull status, shared image cache references across deployments, and cleanup when no deployments reference the image.

AWS benchmarked models from 57–145 GB and reported around 60 percent faster scale-out when weights caching was enabled. The image cache removed over two minutes of cold image-pull time and typically achieved up to a 97 percent reduction compared with pulling fresh from ECR on every pod start. For models over 600 GB such as DeepSeek-R1, AWS says caching removes what would otherwise be an over 30 minute download.

Model caching works with Amazon S3, Amazon FSx for Lustre, HuggingFace Hub, Amazon SageMaker JumpStart non-gated models, and Amazon SageMaker JumpStart gated models. AWS listed NVMe storage examples of 250 GB for ml.g5.xlarge, 3,800 GB for ml.g5.12xlarge, 7,600 GB for ml.g5.48xlarge, 8,000 GB for ml.p4d.24xlarge, and 30,000 GB for ml.p5.48xlarge.

AWS noted limits around local storage and cache freshness. The weights cache is per-node, so NVMe use scales with node count; initial cache population still requires one remote download; a 300 GB model cannot be cached on an instance with only 250 GB of NVMe; and source updates are not auto-detected if files change at the same Amazon S3 path without an InferenceEndpointConfig spec update. Model caching is generally available in all regions where Amazon SageMaker HyperPod is available.

AWS launched model caching for Amazon SageMaker Inference on HyperPod on September 10, 2026, to reduce cold starts when deploying large language models for inference. The feature pre-loads model weights and inference server container images onto cluster nodes before pods serve traffic, replacing repeated network downloads with reads from local NVMe storage at approximately 7 GB/s. AWS says pods can typically begin serving traffic in seconds rather than tens of minutes after caching is enabled.

Without caching, a HyperPod inference pod waits for two sequential downloads before handling requests. Kubelet first pulls multi-gigabyte inference server images such as vLLM or LMI from Amazon ECR, a step AWS says takes 5–7 minutes. The inference server then downloads model weights from Amazon S3, Amazon FSx for Lustre, or HuggingFace Hub; a 145 GB model on Amazon S3 can take another 20+ minutes, while a 600+ GB model such as DeepSeek-R1 can take upwards of 30 minutes.

The same delay repeats during scale-out. If a HorizontalPodAutoscaler requests five new pods after a traffic spike, all five pods independently pull images and weights before they can take requests. AWS says autoscaling policies may react in seconds, but the actual added serving capacity can arrive only after 25–30+ minutes because each new pod waits on downloads.

Model caching has two independent options: weights cache and image cache. The weights cache downloads model weights to local NVMe storage on each target node after users add modelCacheConfig with weightsCache enabled to an InferenceEndpointConfig or JumpStartModel resource. The HyperPod Inference Operator creates a ModelDataCacheConfig resource, downloads weights from Amazon S3, Amazon FSx for Lustre, HuggingFace Hub, or JumpStart, labels nodes as cache-ready, and waits for all target nodes before creating the inference deployment.

The image cache pre-pulls the inference server container image onto nodes after users enable imageCache in modelCacheConfig. The operator creates a DaemonSet (one pod per selected node) to pull the image, but unlike the weights cache it creates the inference deployment immediately. Pods with a cached image skip the Amazon ECR pull and save 5–7 minutes; pods that start before caching completes pull from ECR normally.

AWS says both caching modes use preferred scheduling rather than required scheduling. Pods prefer nodes with cached data, but they are not blocked if a scheduler places them on a node without a warm cache during rapid scale-out. In that case, the pod reads weights from the original Amazon S3 or Amazon FSx source and pulls the image from Amazon ECR, matching the normal uncached behavior.

The operator manages caching with two Custom Resource Definitions, or CRDs (Kubernetes extension objects). ModelDataCacheConfig manages model weight downloads, cache-ready labels, cache health, and cleanup when the parent resource is deleted. ModelImageCache manages image pre-pulling, image-ready labels, per-node pull status, shared image cache references across deployments, and cleanup when no deployments reference the image.

AWS benchmarked models from 57–145 GB and reported around 60 percent faster scale-out when weights caching was enabled. The image cache removed over two minutes of cold image-pull time and typically achieved up to a 97 percent reduction compared with pulling fresh from ECR on every pod start. For models over 600 GB such as DeepSeek-R1, AWS says caching removes what would otherwise be an over 30 minute download.

Model caching works with Amazon S3, Amazon FSx for Lustre, HuggingFace Hub, Amazon SageMaker JumpStart non-gated models, and Amazon SageMaker JumpStart gated models. AWS listed NVMe storage examples of 250 GB for ml.g5.xlarge, 3,800 GB for ml.g5.12xlarge, 7,600 GB for ml.g5.48xlarge, 8,000 GB for ml.p4d.24xlarge, and 30,000 GB for ml.p5.48xlarge.

AWS noted limits around local storage and cache freshness. The weights cache is per-node, so NVMe use scales with node count; initial cache population still requires one remote download; a 300 GB model cannot be cached on an instance with only 250 GB of NVMe; and source updates are not auto-detected if files change at the same Amazon S3 path without an InferenceEndpointConfig spec update. Model caching is generally available in all regions where Amazon SageMaker HyperPod is available.

Read original (English)·Sep 10, 2026
Infra#sagemaker hyperpod#model caching#inference#cold starts#nvme#amazon ecr#amazon s3#fsx for lustre#deepseek r1#kubernetes