TLDRocket
Sign in
Latest Nebius looks to raise $4.5BN through bond issue — Tech.eu Also’s $3,500 e-bike is a $1 billion Trojan horse for autonomous trans... — Fortune Unitree, famous for its dancing robots, surges by 460% on its trading... — Fortune Exclusive: Replit taps OpenAI's low-cost Luna model for new 'Free Mode... — Fortune Adronite launches Codistry AI coding platform, claims half the token c... — SiliconANGLE Rundoo raises $30M to expand its AI-native operating system for small... — SiliconANGLE Temporal is in talks to raise $500M at a $12B pre-money valuation, mor... — Tech Funding News Etched raises $700M led by Jane Street, doubling to $21B and it still... — Tech Funding News

The AI intelligence platform

Every AI story that matters and the intelligence behind it.

TLDRocket reads all relevant sources, removes duplicate coverage, and publishes a short neutral summary of every story, linking back to the original. Free, no spam, unsubscribe anytime.

Add to Slack

Every story also updates live profiles event timelines weekly rankings the AI Market Index

Wednesday, 31 January 2024

Building an early warning system for LLM-aided biological threat creation

OpenAI 2 years ago 10

Researchers created an evaluation framework to assess whether large language models could help someone develop biological threats, testing GPT-4 with biology experts and students. The testing found GPT-4 provided at most a mild accuracy improvement for biological threat creation tasks, a difference too small to be statistically conclusive. The work establishes a baseline methodology for ongoing assessment of LLM biosecurity risks and calls for further investigation across different models and threat scenarios.

Introducing the Enterprise Scenarios Leaderboard: a Leaderboard for Real World Use Cases

Hugging Face 2 years ago 40

Patronus announced the Enterprise Scenarios Leaderboard, which evaluates language models on real-world business tasks rather than academic benchmarks. The leaderboard includes 6 tasks—FinanceBench, Legal Confidentiality, Creative Writing, Customer Support Dialogue, Toxicity, and Enterprise PII—with metrics ranging from accuracy to relevance and toxicity scores. To prevent test-set gaming, four of the six datasets are kept closed source, with only validation sets released to users.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.