模型介绍
中文整理Qwen-72B-Chat
🤗 Hugging Face     🤖 ModelScope      📑 Paper    |   🖥️ Demo
WeChat (微信)     Discord   |   API
介绍(Introduction)
通义千问-72B**(**Qwen-72B**)是阿里云研发的通义千问大模型系列的720亿参数规模的模型。Qwen-72B是基于Transformer的大语言模型, 在超大规模的预训练数据上进行训练得到。预训练数据类型多样,覆盖广泛,包括大量网络文本、专业书籍、代码等。同时,在Qwen-72B的基础上,我们使用对齐机制打造了基于大语言模型的AI助手Qwen-72B-Chat。本仓库为Qwen-72B-Chat的仓库。
通义千问-72B(Qwen-72B)主要有以下特点:
- **大规模高质量训练语料**:使用超过3万亿tokens的数据进行预训练,包含高质量中、英、多语言、代码、数学等数据,涵盖通用及专业领域的训练语料。通过大量对比实验对预训练语料分布进行了优化。
- **强大的性能**:Qwen-72B在多个中英文下游评测任务上(涵盖常识推理、代码、数学、翻译等),效果显著超越现有的开源模型。具体评测结果请详见下文。
- **覆盖更全面的词表**:相比目前以中英词表为主的开源模型,Qwen-72B使用了约15万大小的词表。该词表对多语言更加友好,方便用户在不扩展词表的情况下对部分语种进行能力增强和扩展。
- **更长的上下文支持**:Qwen-72B支持32k的上下文长度。
- **系统指令跟随**:Qwen-72B-Chat可以通过调整系统指令,实现**角色扮演**,**语言风格迁移**,**任务设定**,和**行为设定**等能力。
如果您想了解更多关于通义千问72B开源模型的细节,我们建议您参阅GitHub代码库。
Qwen-72B** is the 72B-parameter version of the large language model series, Qwen (abbr. Tongyi Qianwen), proposed by Alibaba Cloud. Qwen-72B is a Transformer-based large language model, which is pretrained on a large volume of data, including web texts, books, codes, etc. Additionally, based on the pretrained Qwen-72B, we release Qwen-72B-Chat, a large-model-based AI assistant, which is trained with alignment techniques. This repository is the one for Qwen-72B-Chat.
The features of Qwen-72B include:
- **Large-scale high-quality training corpora**: It is pretrained on over 3 trillion tokens, including Chinese, English, multilingual texts, code, and mathematics, covering general and professional fields. The distribution of the pre-training corpus has been optimized through a large number of ablation experiments.
- **Competitive performance**: It significantly surpasses existing open-source models on multiple Chinese and English downstream evaluation tasks (including commonsense, reasoning, code, mathematics, etc.). See below for specific evaluation results.
- **More comprehensive vocabulary coverage**: Compared with other open-source models based on Chinese and English vocabularies, Qwen-72B uses a vocabulary of over 150K tokens. This vocabulary is more friendly to multiple languages, enabling users to directly further enhance the capability for certain languages without expanding the vocabulary.
- **Longer context support**: Qwen-72B supports 32k context length.
- **System prompt**: Qwen-72B can realize roly playing, language style transfer, task setting, and behavior setting by using system prompt.
For more details about the open-source model of Qwen-72B, please refer to the GitHub code repository.
要求(Requirements)
python 3.8及以上版本
pytorch 1.12及以上版本,推荐2.0及以上版本
建议使用CUDA 11.4及以上(GPU用户、flash-attention用户等需考虑此选项)
- *运行BF16或FP16模型需要多卡至少144GB显存(例如2xA100-80G或5xV100-32G);运行Int4模型至少需要48GB显存(例如1xA100-80G或2xV100-32G)**
python 3.8 and above
pytorch 1.12 and above, 2.0 and above are recommended
CUDA 11.4 and above are recommended (this is for GPU users, flash-attention users, etc.)
- *To run Qwen-72B-Chat in bf16/fp16, at least 144GB GPU memory is required (e.g., 2xA100-80G or 5xV100-32G). To run it in int4, at least 48GB GPU memory is required (e.g., 1xA100-80G or 2xV100-32G)**
依赖项(Dependency)
使用HuggingFace进行推理
运行Qwen-72B-Chat,请确保满足上述要求,再执行以下pip命令安装依赖库
To run Qwen-72B-Chat, please make sure you meet the above requirements, and then execute the following pip commands to install the dependent libraries.
另外,推荐安装 flash-attention 库(**当前已支持flash attention 2**),以实现更高的效率和更低的显存占用。
In addition, it is recommended to install the flash-attention library (**we support flash attention 2 now.**) for higher efficiency and lower memory usage.
使用vLLM进行推理
使用vLLM进行推理可以支持更长的上下文长度并获得至少两倍的生成加速。你需要满足以下要求:
Using vLLM for inference can support longer context lengths and obtain at least twice the generation speedup. You need to meet the following requirements:
pytorch >= 2.0
cuda 11.8 or 12.1
如果你使用cuda12.1和pytorch2.1,可以直接使用以下命令安装vLLM。
If you use cuda 12.1 and pytorch 2.1, you can directly use the following command to install vLLM.
否则请参考vLLM官方的安装说明,或者我们vLLM分支仓库(支持量化模型)。
Otherwise, please refer to the official vLLM Installation Instructions, or our vLLM repo for GPTQ quantization.
快速使用(Quickstart)
使用HuggingFace Transformers进行推理(Inference with Huggingface Transformers)
下面我们展示了一个使用Qwen-72B-Chat模型,进行多轮对话交互的样例:
We show an example of multi-turn interaction with Qwen-72B-Chat in the following code:
使用vLLM和类Transformers接口进行推理(Inference with vLLM and Transformers-like APIs)
在根据上方依赖性部分的说明安装vLLM后,可以下载接口封装代码到当前文件夹,并执行以下命令进行多轮对话交互。(注意:该方法当前只支持 model.chat() 接口。)
After installing vLLM according to the dependency section above, you can download the wrapper codes and execute the following commands for multiple rounds of dialogue interaction. (Note: It currently only supports the model.chat() method.)
使用vLLM和类OpenAI接口进行推理(Inference with vLLM and OpenAI-like API)
请参考我们GitHub repo中vLLM部署和OpenAI接口使用两个部分的介绍。
Please refer to the introduction of vLLM deployment and OpenAI interface usage in our GitHub repo.
如果使用2xA100-80G进行部署,可以运行以下代码:
If deploying with 2xA100-80G, you can run the following code:
注意需要 --gpu-memory-utilization 0.98 参数避免OOM问题。
Note that the --gpu-memory-utilization 0.98 parameter is required to avoid OOM problems.
关于更多的使用说明,请参考我们的GitHub repo获取更多信息。
For more information, please refer to our GitHub repo for more information.
量化 (Quantization)
用法 (Usage)
以下我们提供示例说明如何使用Int4/Int8量化模型。在开始使用前,请先保证满足要求(如torch 2.0及以上,transformers版本为4.32.0及以上,等等),并安装所需安装包:
Here we demonstrate how to use our provided quantized models for inference. Before you start, make sure you meet the requirements of auto-gptq (e.g., torch 2.0 and above, transformers 4.32.0 and above, etc.) and install the required packages:
如安装 auto-gptq 遇到问题,我们建议您到官方repo搜索合适的预编译wheel。
If you meet problems installing auto-gptq , we advise you to check out the official repo to find a pre-build wheel.
注意:预编译的 auto-gptq 版本对 torch 版本及其CUDA版本要求严格。同时,由于
其近期更新,你可能会遇到 transformers 、 optimum 或 peft 抛出的版本错误。
我们建议使用符合以下要求的最新版本:
- torch==2.1 auto-gptq>=0.5.1 transformers>=4.35.0 optimum>=1.14.0 peft>=0.6.1
- torch>=2.0, - torch==2.1 auto-gptq>=0.5.1 transformers>=4.35.0 optimum>=1.14.0 peft>=0.6.1
- torch>=2.0,
优先采用原作者提供的中文模型卡片;仅有英文说明时由闲社AI翻译整理。重要参数、许可证和商用范围请以原文为准。
你将获得什么
完整仓库方式包含该固定版本中的权重、配置、分词器及说明文件;指定版本方式只下载所选文件或整组分片。多模态模型的投影文件、配置等依赖,请按作者说明配齐。
完整仓库可能包含多个权重格式,因此需要的磁盘空间可能大于单一格式的模型大小。下载文件不等于完成模型部署。
使用前须知
运行模型需要的设备、工具与依赖,以原作者说明为准。
是否允许商用、修改和分发,请查看模型许可证。
闲社提供资料与镜像下载方法,不代表已运行模型或完成安全审计。下载的代码应先检查,再执行。
资料或下载有问题?
登录后可提交反馈。
