Google DeepMind
·
3 months ago
● 2 sources
Google released Gemma 4, an open-source model family in four sizes designed for reasoning and agentic workflows. The 31B dense model ranks third on Arena AI's text leaderboard and the 26B mixture-of-experts model ranks sixth, outperforming models 20 times larger. Developers can now run frontier-class reasoning on consumer hardware, from Android phones to laptop GPUs, under a commercially permissive Apache 2.0 license.
Deep Learning Weekly
·
3 months ago
● 2 sources
This weekly newsletter covers recent developments in deep learning and AI, including Google's Gemini 3.1 Flash Live achieving 90.8% on ComplexFuncBench Audio, Cohere's open-source Transcribe model reaching 5.42% word error rate on the HuggingFace leaderboard, and research papers on efficient sparse attention and agentic image synthesis. Mistral's Voxtral supports 9 languages with 70ms latency, Meta's SAM 3.1 doubles video processing speed to 32 FPS, and Granola raised $125M at $1.5B valuation. The broader collection documents progress in LLM inference optimization, multimodal embeddings, agentic workflows, and production deployment tools.
OpenAI Blog
·
3 months ago
OpenAI has acquired TBPN, a media platform focused on AI discussions. The acquisition includes TBPN's existing audience and editorial operations dedicated to covering AI developments. OpenAI aims to use the platform to reach builders, businesses, and the broader tech community while supporting independent media coverage of AI topics.
OpenAI Blog
·
3 months ago
Codex introduced pay-as-you-go pricing for its ChatGPT Business and Enterprise tiers. The new model allows teams to pay only for usage rather than committing to fixed subscription costs upfront. This enables smaller teams or those testing the platform to adopt Codex without large initial financial commitments.
Together AI
·
3 months ago
Deepgram's speech-to-text and text-to-speech models are now available natively on Together AI's infrastructure for building real-time voice agents. The deployment includes Deepgram's Flux conversational STT model with 250ms end-of-turn detection, Nova-3 production transcription, Nova-3 Multilingual for language switching, and Aura-2 enterprise TTS, all running on Together AI's Dedicated Model Inference with 99.9% uptime SLA and HIPAA-ready support. Teams can now run their entire voice pipeline—transcription, language model reasoning, and synthesis—on a single platform, reducing latency and operational complexity for contact center, healthcare, financial services, and multilingual customer support applications.
Hugging Face Blog
·
3 months ago
● 2 sources
Google DeepMind released Gemma 4, a family of multimodal models available on Hugging Face with Apache 2 licenses that process images, text, and audio inputs. The largest dense model (31B parameters) achieved an estimated LMArena score of 1452, while a 26B mixture-of-experts variant reached 1441 using only 4B active parameters. The models are deployable across multiple frameworks and devices, from cloud infrastructure to on-device inference, supporting tasks including object detection, speech-to-text, and code completion.