Deep Learning Weekly
·
4 months ago
● 2 sources
Google launched Gemini Embedding 2, a multimodal embedding model unifying text, images, video, audio, and documents; OpenAI acquired Promptfoo for AI security and launched GPT-5.4 with 1M-token context; and Yann LeCun's AMI Labs raised $1.03 billion to develop world models. The article covers multiple AI industry developments including new models, acquisitions, funding, and research papers on agent planning and long-context video reconstruction. These launches and funding rounds represent significant advancement in multimodal AI capabilities, agent planning systems, and world model development.
One Useful Thing
·
4 months ago
AI systems have advanced from co-intelligence tools requiring human prompting to autonomous agents capable of completing hours of work in minutes, with capabilities demonstrated across benchmarks like the Google-Proof Q&A (94% accuracy) and GDPval (82% parity with top human performance). StrongDM's Software Factory represents a radical organizational experiment where AI agents write, test, and ship production code without human involvement, requiring $1,000 per day in AI token costs per engineer. As AI capabilities continue improving exponentially and companies pursue recursive self-improvement where AI systems build better AI systems, organizations face unpredictable disruption across markets, employment, and governance, creating an unstable near-term environment where current choices about AI deployment will set precedents.
Google Research
·
4 months ago
● 2 sources
Google announced Urban Flash Flood forecasts on Flood Hub using AI to predict flash flood risk in urban areas up to 24 hours in advance. The model was trained on a dataset of historical flood events extracted from news reports using Gemini and operates at 20x20 kilometer spatial resolution using global weather data from NASA, NOAA, and DeepMind. The system achieves recall and precision performance equivalent to the US National Weather Service in many flood-affected countries in the Global South, addressing a critical gap in early warning systems where less than half of developing countries have access to multi-hazard forecasting.
Google Research
·
4 months ago
● 2 sources
Google introduced Groundsource, a methodology that uses Gemini to extract flood data from global news reports and create historical disaster datasets. The initial flash flood dataset contains 2.6 million records spanning more than 150 countries from 2000 to present, with 82% of extracted events accurate enough for practical analysis. The structured data enables Google to provide urban flash flood forecasts up to 24 hours in advance through its Flood Hub service.
Together AI
·
4 months ago
Together AI launched a unified platform for building real-time voice agents with co-located speech-to-text, language model, and text-to-speech components on a single cloud infrastructure. The system achieves end-to-end latency under 500 milliseconds and now includes native integrations with Cartesia (TTS) and Deepgram (STT) models. This eliminates the need for multi-vendor setups and reduces operational complexity for teams deploying production voice systems.