Promptimus: Improving already good LLM prompts with zero manual engineering
Amazon Science
Amazon researchers presented Promptimus, a method that automatically optimizes existing large language model prompts by identifying failure points and making targeted improvements without full rewrites. The system achieved best results on 16 of 20 benchmarks with an average performance score of 0.792 compared to 0.765 for the best baseline, typically reaching 90% of final performance within 300 metric evaluations. Enterprises can now improve production prompts across different LLM models without manual reengineering while preserving critical domain requirements and business logic.
Why it matters
By focusing on specific failure points and suggesting targeted solutions, a new automated prompt-engineering framework improves prompt performance without compromising existing functionality.