TLDRocket
Sign in

Data Machina #245

Substack

RAG (retrieval-augmented generation) is getting a serious rethink as teams hit walls with cost, accuracy, and complexity in production. New tricks like Command-R, RAFT, and knowledge-graph RAG aim to fix what basic RAG can't.

Based on reporting by Substack — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Retrieval-augmented generation has had a rough few years of growing pains. Since Facebook AI introduced the technique back in 2021, RAG has moved from naive lookups to advanced pipelines to what people now call Modular RAG, stacking on more components, more interfaces, more moving parts. The problem is that a lot of that complexity never made it past the prototype stage. Plenty of companies poured months and budget into enterprise RAG systems only to discover they couldn't hit acceptable accuracy without costs spiraling out of control.

This week's roundup from Data Machina reads like a field guide for anyone trying to get RAG production-ready in 2024. Cohere's new Command-R model tackles the balancing act directly: long-context retrieval, tool use, and API integration, all while trying to keep latency and throughput in check at scale. Meanwhile Weaviate's Verba pushes in a different direction, offering an open-source, modular RAG framework so teams aren't locked into chaining GPT-4 calls through LangChain or Llamaindex and hoping the bill doesn't explode.

A few research threads stand out too. Researchers from Berkeley, Meta AI, and Microsoft published RAFT, a method that fuses retrieval with domain-specific fine-tuning, addressing gaps that neither approach solves alone. And a Netflix paper is pushing back on an assumption almost everyone in RAG has taken for granted: that cosine similarity is a reliable way to match queries to documents. The researchers argue it's a

My take — AI-written commentary, not fact-checked reporting

Command-R and RAFT are proof that RAG's next chapter is about ripping out lazy defaults like cosine similarity, not stacking on more abstraction layers, and I'll take that boring rigor over another flashy agent demo any day. Open tools like Verba matter more than people admit, because nobody should be locking their retrieval stack behind a single vendor's API just to ship a chatbot.

Read more about this at: Substack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.