TLDRocket
Sign in

Qwen

47 summarised stories about Qwen, each linking back to the original source. Browse all topics →

+ Follow this topic

Tuesday, 23 January 2024

Introducing Qwen

GitHub Pages 2 years ago 5 2 sources

Alibaba released the Qwen series of open-source large language models, consisting of four models ranging from 1.8B to 72B parameters trained on 2-3 trillion tokens with support for 32K token context length. The largest model, Qwen-72B, achieved competitive performance on benchmarks against Llama 2 and GPT-3.5, with the research team focusing on multilingual capabilities, alignment techniques including SFT and RLHF, and tool-use abilities through agent frameworks. The release enables researchers and developers to build specialized AI applications through open-source models with support for function calling, code interpretation, and agent-based configurations.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.