TLDRocket
Sign in
Latest As AI reshapes the consumer journey, L’Oréal is rethinking its marketi... — Fortune Why Sam Altman takes his hardest questions to Nikesh Arora — Fortune SAP to acquire AI startup TechWolf for undisclosed sum — Sifted NetApp aims to make legacy data AI-ready without a rebuild — SiliconANGLE Mistral’s new 1T model aims to leapfrog closed and open rivals — TechCrunch The Cyber Risk Discourse is Broken — Interconnects Mistral Large 4 Trails China’s Open-Weight Leaders on Artificial Analy... — Trending Topics Pinterest’s AI now turns beauty Pins into action plans — TechCrunch

The AI intelligence platform

Every AI story that matters — and the intelligence behind it.

TLDRocket reads all relevant sources, removes duplicate coverage, and publishes a short neutral summary of every story, linking back to the original. Free, no spam, unsubscribe anytime.

Add to Slack

Every story also updates live profiles event timelines weekly rankings the AI Market Index

Sunday, 26 May 2024

Data Machina #254

Substack 2 years ago 31

Princeton Language & Intelligence released SWE-bench, a benchmark for evaluating AI coding agents on their ability to fix real GitHub repository issues, revealing that current AI agents perform poorly on the task. The leading model, Amazon Q Developer Agent, successfully solved only 13.8% of 2294 tasks, while the open-source OpenDevin achieved the highest benchmark score at 21%. These results demonstrate that despite multiple competing approaches including Devin, Devika, and GPT-Engineer, AI coding agents remain far from ready for production-scale legacy code migration and autonomous software engineering work.

Prompting Fundamentals and How to Apply them Effectively

Eugene Yan 2 years ago 30

The article explains fundamental techniques for writing effective prompts for large language models, including assigning roles, using structured input/output formats, prefilling responses, and providing multiple examples. Key techniques discussed include using XML tags for structure, providing at least 12 examples in n-shot prompting (not just 3-5), and matching example distributions to production data. The practical application of these fundamentals helps LLM users obtain more reliable and consistent outputs through better conditioning of the model's probabilistic behavior.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.