Agent Seer: Synthesizing Scenarios from Specification Understanding
Apple Machine Learning Research 1 week ago 49
Agent Seer was introduced to synthesize realistic multi-turn tool-use evaluation scenarios directly from a single MCP tool specification instead of hand-built benchmarks. It was evaluated on seven MCP specifications and achieved complete tool coverage on small and medium specifications. As a result, it produces graded scenarios with synthetic outputs and improves measurements of tool-calling correctness and conversational coherence, highlighting argument value accuracy as the main failure mode and parameter schema complexity as the strongest quality correlate.