TLDRocket
Sign in

Introducing Qwen

GitHub Pages Covered by 2 sources

Alibaba released the Qwen series of open-source large language models, consisting of four models ranging from 1.8B to 72B parameters trained on 2-3 trillion tokens with support for 32K token context length. The largest model, Qwen-72B, achieved competitive performance on benchmarks against Llama 2 and GPT-3.5, with the research team focusing on multilingual capabilities, alignment techniques including SFT and RLHF, and tool-use abilities through agent frameworks. The release enables researchers and developers to build specialized AI applications through open-source models with support for function calling, code interpretation, and agent-based configurations.

Why it matters

4 months after our first release of Qwen-7B, which is the starting point of our opensource journey of large language models (LLM), we now provide an introduction to the Qwen series to give you a whole picture of our work as well as our objectives. Below are important links to our opensource projects and community. PAPER GITHUB HUGGING FACE MODELSCOPE DISCORD Additionally, we have WeChat groups for chatting and we invite you to join the groups through the provided link in our GitHub readme.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.