Google Research
·
3 months ago
Google released Vantage, a research experiment using generative AI to assess future-ready skills like critical thinking and collaboration through simulated conversations with AI avatars. In validation studies with 188 testers ages 18-25, the AI Evaluator's scores matched human expert raters with similar agreement levels, and in a separate study of 180 students' creative work, the system showed high correlation with OpenMic's internal experts. The approach aims to make traditionally hard-to-measure human competencies assessable at scale for integration into classroom curricula alongside subject knowledge.
Google DeepMind
·
3 months ago
Google released Gemini Robotics-ER 1.6, an upgraded AI model designed to help robots understand and reason about their physical environments for tasks like navigation and equipment monitoring. The model improves spatial reasoning capabilities including pointing accuracy, multi-view success detection, and a new ability to read analog gauges and digital instrument displays with sub-tick accuracy. Developers can now access the model via the Gemini API and Google AI Studio, with the system showing 6-10% improvement over previous versions in identifying physical safety hazards.
Import AI
·
3 months ago
Researchers demonstrated that Claude Opus 4.6 successfully reimplemented a 16,000-line bioinformatics program (gotree) by observing only its command-line interface and test cases, without access to the original source code. The task would typically require a human engineer 2–17 weeks of work, and performance improved with increased computational resources during inference. This capability suggests AI systems can already autonomously complete weeks-long coding tasks, raising questions about the pace of AI progress in software development.
Allen Institute (AI2)
·
3 months ago
AI2 released two benchmarks—ScienceWorld and DiscoveryWorld—to evaluate whether AI agents can perform scientific tasks rather than just answer questions about science. In DiscoveryWorld's 120 challenge tasks, top models complete only around 20% at higher difficulty levels while human scientists with advanced degrees solve approximately 70%. The benchmarks reveal a gap between knowing scientific concepts and applying them through experimentation, with the tools now freely available as AI agent development accelerates.
OpenAI Blog
·
3 months ago
Cloudflare has integrated OpenAI's GPT-5.4 and Codex models into its Agent Cloud platform to help enterprises build and deploy AI agents. The integration allows companies to create agents that handle real-world tasks while running on Cloudflare's infrastructure. Enterprise customers can now develop agentic workflows with access to advanced language models without building separate infrastructure.
Together AI
·
3 months ago
Google DeepMind released EinsteinArena, a platform where AI agents collaborate on open mathematical problems through shared messaging and leaderboards. On April 11, 2026, agents working together improved the kissing number lower bound in 11 dimensions from 593 to 604 spheres, with one agent proposing a construction that multiple agents then refined through 48 hours of collaborative optimization. The platform enables agents to build on each other's work rather than solving problems in isolation, resulting in 11 new state-of-the-art solutions across various mathematical problems including the Erdős minimum overlap problem and autocorrelation inequalities.