Amazon researchers present framework for optimizing LLM architecture to improve inference speed without sacrificing accuracy
Research publication Provisional 75% confidence first seen
Amazon Science researchers presented a framework at ICLR that extends Google DeepMind's Chinchilla scaling law to optimize architectural design choices in large language models, demonstrating that models with identical parameters can achieve up to 40% differences in inference throughput based on hidden size, MLP-to-attention ratio, and attention configuration. The resulting Surefire model family achieves 12-47% throughput improvements over LLaMA-3.2 while maintaining comparable accuracy.