Tensorwire
Models & releases · first seen 10 Sep, updated 10 Sep

Reduce inference cold starts on Amazon SageMaker HyperPod with model caching

1 outlet Amazon

Amazon SageMaker HyperPod now supports model caching for inference, which pre-loads model weights and container images onto cluster nodes so pods read from local NVMe storage instead of downloading over the network. Learn how model caching…

Summary from AWS Machine Learning Blog.

Coverage 1 article · 1 outlet

  1. AWS Machine Learning Blog
    Reduce inference cold starts on Amazon SageMaker HyperPod with model caching