smevals - a small eval suite for evaluating models, prompts, and harnesses
Simon Willison's Weblog 1 month ago 47
A researcher released smevals, a framework for evaluating AI models and prompts across different configurations by running tests and grading results. The tool uses YAML-based eval suites that can be tested against multiple models like GPT and Claude, with results viewable through a web interface or static HTML reports. This enables teams to systematically benchmark model capabilities and compare performance across different versions and configurations.