Google Research
·
3 months ago
Researchers introduced ReasoningBank, a memory framework that enables AI agents to learn from both successful and failed experiences during deployment. On WebArena and SWE-Bench-Verified benchmarks, ReasoningBank improved success rates by 8.3% and 4.6% respectively while reducing task execution steps by nearly 3 on the latter benchmark. The framework allows agents to build increasingly sophisticated strategic memories over time rather than repeating mistakes or discarding insights from failures.
Google DeepMind
·
3 months ago
Google DeepMind is partnering with Accenture, Bain & Company, BCG, Deloitte, and McKinsey to help enterprises adopt AI technology at scale across finance, manufacturing, retail, media and entertainment. Currently only 25% of organizations have successfully moved AI into production at scale, despite AI's potential to contribute $15.7 trillion to the global economy by 2030. The partners will gain early access to Google DeepMind's frontier models including Gemini, direct engagement with technical talent, and support developing industry-specific AI solutions.
OpenAI Blog
·
3 months ago
OpenAI released ChatGPT Images 2.0, an updated image generation model with improved text rendering and support for multiple languages. The new model can now accurately generate text within images and handle prompts in languages beyond English. Users can generate more legible and contextually appropriate images for international applications.
Hugging Face Blog
·
3 months ago
QIMMA is a new Arabic language model evaluation leaderboard that validates benchmark quality before running model assessments, addressing fragmentation and quality issues across existing Arabic NLP benchmarks. The platform consolidates 52,000 samples from 14 benchmarks across 7 domains and discarded between 0.2% and 12.3% of samples per benchmark after applying automated and human quality review, with ArabicMMLU losing 436 samples at a 3.1% discard rate. Rankings now reflect genuine Arabic language capability, with Qwen3.5-397B achieving 68.06% average score, though smaller specialized Arabic models outperform larger multilingual models on cultural and linguistic tasks.
Together AI
·
3 months ago
Multi-tenant GPU clusters allow AI companies to share compute infrastructure across teams while maintaining isolation through dedicated nodes, storage, and self-serve scheduling. The architecture requires three elements: pooled capacity at the infrastructure layer, per-tenant isolation with dedicated resources, and quota-based allocation with hard limits enforced at the scheduler level. Together AI's implementation demonstrates that this approach eliminates idle capacity waste while preventing cross-team resource conflicts and maintaining billing visibility per tenant.
Hugging Face Blog
·
3 months ago
Mythos, a large language model embedded in a system designed to find and patch software vulnerabilities, demonstrates that AI cybersecurity capability depends on the complete system rather than the model alone, combining compute power, specialized data, vulnerability-probing scaffolding, and autonomy. The system's power lies in its recipe of components that can be replicated by others at lower cost using smaller models with deep security expertise. Open-source ecosystems and semi-autonomous AI agents running on transparent, auditable foundations offer defenders better protection than proprietary systems, particularly for high-stakes organizations that can inspect, control, and customize their security infrastructure internally.
OpenAI Blog
·
3 months ago
OpenAI launched Codex Labs and partnered with Accenture, PwC, and Infosys to help enterprises deploy and scale its Codex code-generation model. The service reached 4 million weekly active users. The partnerships aim to integrate Codex across enterprise software development workflows.