Introducing Qwen
Qwen ● Covered by 2 sources
Alibaba released the Qwen series of open-source large language models, consisting of four models ranging from 1.8B to 72B parameters trained on 2-3 trillion tokens with support for 32K token context length. The largest model, Qwen-72B, achieved competitive performance on benchmarks against Llama 2 and GPT-3.5, with the research team focusing on multilingual capabilities, alignment techniques including SFT and RLHF, and tool-use abilities through agent frameworks. The release enables researchers and developers to build specialized AI applications through open-source models with support for function calling, code interpretation, and agent-based configurations.
Why it matters
4 months after our first release of Qwen-7B, which is the starting point of our opensource journey of large language models (LLM), we now provide an introduction to the Qwen series to give you a whole picture of our work as well as our objectives. Below are important links to our opensource projects and community. PAPER GITHUB HUGGING FACE MODELSCOPE DISCORD Additionally, we have WeChat groups for chatting and we invite you to join the groups through the provided link in our GitHub readme.