Prefix-Aware Routing
Amazon SageMaker Inference's prefix-aware routing reduces latency in large language models by enabling cache reuse. Benchmarking with Llama 3.1 70B Instruct showed a maximum 77% reduction in P50 TTFT and a 16% increase in total throughput. This feature can help decrease inference latency and improve model throughput in Amazon SageMaker Inference.