Products & tools · first seen 10 Sep, updated 10 Sep
Reduce LLM latency with prefix-aware routing on Amazon SageMaker Inference
Amazon SageMaker Inference now offers prefix-aware routing, a routing strategy that sends requests sharing the same prompt prefix to the same instance so the KV cache stays warm. In benchmarks on Llama 3.1 70B, it reduced P50 time-to-first-…
Summary from AWS Machine Learning Blog.
Coverage 1 article · 1 outlet
-
AWS Machine Learning BlogReduce LLM latency with prefix-aware routing on Amazon SageMaker Inference