TLDRocket
Sign in

Data Machina #253

Substack

Google dropped a huge AI update wave right after OpenAI's GPT-4o launch, packing in new models, tools, and safety rules. It's basically Google saying 'anything you can do' — with a 2 million token context window to prove it.

Based on reporting by Substack — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI kicked off the week with GPT-4o, a real-time multimodal model that handles audio, vision, and text together. Benchmarks came in softer than the hype suggested, and a chunk of the internet got distracted arguing about how flirty the voice mode sounded. Then Google showed up and buried the news cycle under an avalanche of releases.

The headline act is Gemini 1.5 Pro, now stretched to a 2 million token context window, with Google's own researchers claiming near-perfect retrieval accuracy — above 99 percent — out to 10 million tokens in testing. For comparison, Claude 3.0 tops out at 200,000 tokens and GPT-4 Turbo sits at 128,000. Alongside the context expansion, Google added multimodal prompting, custom function calling for real-time external interactions, configurable system instructions, and context caching to cut costs on repetitive queries. A smaller sibling, Gemini 1.5 Flash, targets high-volume, latency-sensitive tasks with a still-generous 1 million token window.

On the open-model side, Google introduced PaliGemma, a vision-language model built on the SigLIP vision encoder and Gemma language model, aimed at tasks like captioning, visual question answering, object detection, and reading text embedded in images. And Gemma 2 landed too — a 27 billion parameter open model Google says matches Mistral and Llama 3 70B performance at less than half the parameter count, which if it holds up is a genuinely useful efficiency jump for anyone running models without a data center budget.

The more speculative reveal was Project Astra, a prototype AI agent that sees, listens, and remembers context to act proactively with minimal lag. Google paired all of this with practical infrastructure: a Model Explorer tool for visualizing and debugging large ML graphs, a Responsible Generative AI toolkit with an LLM comparator, and a new Frontier Safety Framework meant to flag capabilities — like advanced cyber skills or unusual autonomy — before they become dangerous. Google also opened a $1 million developer competition built around the Gemini API, presumably to make sure none of this goes unused.

My take — AI-written commentary, not fact-checked reporting

I run an independent AI newsletter, not a PR wire, so I'll say it plainly: Google timed this dump to bury GPT-4o's mixed reviews, and it worked. But the real story isn't the spectacle — it's that a 27B open model matching 70B-class performance matters more long-term than another flirty voice demo, because efficiency, not size, is what actually gets AI out of hyperscaler data centers and into things regular people can run.

Read more about this at: Substack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.