Models & releases · first seen 10 Sep, updated 10 Sep
Reduce inference cold starts on Amazon SageMaker HyperPod with model caching
Amazon SageMaker HyperPod now supports model caching for inference, which pre-loads model weights and container images onto cluster nodes so pods read from local NVMe storage instead of downloading over the network. Learn how model caching…
Summary from AWS Machine Learning Blog.
Coverage 1 article · 1 outlet
-
AWS Machine Learning BlogReduce inference cold starts on Amazon SageMaker HyperPod with model caching