尧图建网站 尧图建网站 YAOTU WEB BUILD 免费咨询
ARTICLE DETAIL

资讯详情

深耕网站建设与建站编程的一线实战洞察。

vLLM自定义对话模板

vLLM自定义对话模板 为什么 vLLM 要支持自定义对话模板把大语言模型接入线上服务时我们习惯以system、user、assistant这样的结构发送消息[{role:system,content:你是一个专业的助手。},{role:user,content:请解释什么是 KV Cache。}]但对模型而言这些结构化字段并不能直接被理解。模型最终接收到的始终是一段经过分词后的文本和 token 序列。问题在于不同模型在训练时使用的“对话文本格式”并不相同有的使用 ChatML有的使用[INST]...[/INST]有的使用 Llama 3 风格的角色标记也有模型采用自己定义的特殊 token。这正是对话模板Chat Template存在的意义。它负责将接口层传入的标准消息转换为目标模型在训练阶段最熟悉的 prompt 格式。例如同样一句用户问题可能需要被拼接为|system| 你是一个专业的助手。 |user| 请解释什么是 KV Cache。 |assistant|也可能需要采用完全不同的格式s[INST] SYS 你是一个专业的助手。 /SYS 请解释什么是 KV Cache。 [/INST]比如说如果我们使用LLaMA Factory进行微调LLaMA Factory微调使用的模板是自己写好的在下面的代码里面但是使用vLLM这些推理框架的时候他们使用的是模型配置文件里面的模板如果模板与模型训练时的格式不匹配轻则回答质量下降重则出现角色混乱、系统提示词失效、重复输出标签、多轮对话错位甚至工具调用无法正常工作。因此vLLM 提供自定义对话模板的能力并不只是为了“灵活配置提示词”。更重要的是它让服务端能够准确适配不同模型、不同微调数据格式以及工具调用、多模态等更复杂的推理场景。理解这一点是正确部署和使用 vLLM 对话服务的第一步。导出LLaMA Factory的模板在LLaMA Factory的template.py文件中有很多内部方法可以生成模板我们可以使用这些方法来生成我们的模板创建导出文件在LLaMA-Factory/src/llamafactory目录下创建一个export_template.py的文件内容如下importsysimportos# 将项目根目录添加到Python路径root_diros.path.dirname(os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))))sys.path.append(root_dir)fromllamafactory.data.templateimportTEMPLATESfromtransformersimportAutoTokenizer# 初始化分词器任意支持的分词器均可tokenizerAutoTokenizer.from_pretrained(/home/gillbert/code/hugging_face_test/modelscope_test/llm/models/Qwen--Qwen3.5-2B/snapshots/master)# 获取模板对象template_nameqwen3# 该名称是LLaMA Factory template.py中有的模型名称templateTEMPLATES[template_name]# 修复分词器的Jinja模板template.fix_jinja_template(tokenizer)# 输出模板的Jinja格式print(*60)print(tokenizer.chat_template)执行该脚本会得到如下输出(llamafactory)➜ data git:(main)✗ python export_template.py{%-setimage_countnamespace(value0)%}{%-setvideo_countnamespace(value0)%}{%- macro render_content(content, do_vision_count,is_system_contentfalse)%}{%-ifcontent is string %}{{- content}}{%-elifcontent is iterable and content is not mapping %}{%-foritemincontent %}{%-ifimageinitem orimage_urlinitem or item.typeimage%}{%-ifis_system_content %}{{- raise_exception(System message cannot contain images.)}}{%- endif %}{%-ifdo_vision_count %}{%-setimage_count.valueimage_count.value 1%}{%- endif %}{%-ifadd_vision_id %}{{-Picture ~ image_count.value ~: }}{%- endif %}{{-|vision_start||image_pad||vision_end|}}{%-elifvideoinitem or item.typevideo%}{%-ifis_system_content %}{{- raise_exception(System message cannot contain videos.)}}{%- endif %}{%-ifdo_vision_count %}{%-setvideo_count.valuevideo_count.value 1%}{%- endif %}{%-ifadd_vision_id %}{{-Video ~ video_count.value ~: }}{%- endif %}{{-|vision_start||video_pad||vision_end|}}{%-eliftextinitem %}{{- item.text}}{%-else%}{{- raise_exception(Unexpected item type in content.)}}{%- endif %}{%- endfor %}{%-elifcontent is none or content is undefined %}{{-}}{%-else%}{{- raise_exception(Unexpected content type.)}}{%- endif %}{%- endmacro %}{%-ifnot messages %}{{- raise_exception(No messages provided.)}}{%- endif %}{%-iftools and tools is iterable and tools is not mapping %}{{-|im_start|system\n}}{{-# Tools\n\nYou have access to the following functions:\n\ntools}}{%-fortoolintools %}{{-\n}}{{- tool|tojson}}{%- endfor %}{{-\n/tools}}{{-\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\ntool_call\nfunctionexample_function_name\nparameterexample_parameter_1\nvalue_1\n/parameter\nparameterexample_parameter_2\nThis is the value for the second parameter\nthat can span\nmultiple lines\n/parameter\n/function\n/tool_call\n\nIMPORTANT\nReminder:\n- Function calls MUST follow the specified format: an inner function.../function block must be nested within tool_call/tool_call XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n/IMPORTANT}}{%-ifmessages[0].rolesystem%}{%-setcontentrender_content(messages[0].content, false,true)|trim %}{%-ifcontent %}{{-\n\n content}}{%- endif %}{%- endif %}{{-|im_end|\n}}{%-else%}{%-ifmessages[0].rolesystem%}{%-setcontentrender_content(messages[0].content, false,true)|trim %}{{-|im_start|system\n content |im_end|\n}}{%- endif %}{%- endif %}{%-setnsnamespace(multi_step_tooltrue,last_query_indexmessages|length -1)%}{%-formessageinmessages[::-1]%}{%-setindex(messages|length -1)- loop.index0 %}{%-ifns.multi_step_tool and message.roleuser%}{%-setcontentrender_content(message.content,false)|trim %}{%-ifnot(content.startswith(tool_response)and content.endswith(/tool_response))%}{%-setns.multi_step_toolfalse%}{%-setns.last_query_indexindex %}{%- endif %}{%- endif %}{%- endfor %}{%-ifns.multi_step_tool %}{{- raise_exception(No user query found in messages.)}}{%- endif %}{%-formessageinmessages %}{%-setcontentrender_content(message.content,true)|trim %}{%-ifmessage.rolesystem%}{%-ifnot loop.first %}{{- raise_exception(System message must be at the beginning.)}}{%- endif %}{%-elifmessage.roleuser%}{{-|im_start| message.role \n content |im_end|\n}}{%-elifmessage.roleassistant%}{%-setreasoning_content%}{%-ifmessage.reasoning_content is string %}{%-setreasoning_contentmessage.reasoning_content %}{%-else%}{%-if/thinkincontent %}{%-setreasoning_contentcontent.split(/think)[0].rstrip(\n).split(think)[-1].lstrip(\n)%}{%-setcontentcontent.split(/think)[-1].lstrip(\n)%}{%- endif %}{%- endif %}{%-setreasoning_contentreasoning_content|trim %}{%-ifloop.index0ns.last_query_index %}{{-|im_start| message.role \nthink\n reasoning_content \n/think\n\n content}}{%-else%}{{-|im_start| message.role \n content}}{%- endif %}{%-ifmessage.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}{%-fortool_callinmessage.tool_calls %}{%-iftool_call.function is defined %}{%-settool_calltool_call.function %}{%- endif %}{%-ifloop.first %}{%-ifcontent|trim %}{{-\n\ntool_call\nfunction tool_call.name \n}}{%-else%}{{-tool_call\nfunction tool_call.name \n}}{%- endif %}{%-else%}{{-\ntool_call\nfunction tool_call.name \n}}{%- endif %}{%-iftool_call.arguments is defined %}{%-forargs_name, args_valueintool_call.arguments|items %}{{-parameter args_name \n}}{%-setargs_valueargs_value|tojson|safeifargs_value is mapping or(args_value is sequence and args_value is not string)elseargs_value|string %}{{- args_value}}{{-\n/parameter\n}}{%- endfor %}{%- endif %}{{-/function\n/tool_call}}{%- endfor %}{%- endif %}{{-|im_end|\n}}{%-elifmessage.roletool%}{%-ifloop.previtem and loop.previtem.role!tool%}{{-|im_start|user}}{%- endif %}{{-\ntool_response\n}}{{- content}}{{-\n/tool_response}}{%-ifnot loop.last and loop.nextitem.role!tool%}{{-|im_end|\n}}{%-elifloop.last %}{{-|im_end|\n}}{%- endif %}{%-else%}{{- raise_exception(Unexpected message role.)}}{%- endif %}{%- endfor %}{%-ifadd_generation_prompt %}{{-|im_start|assistant\n}}{%-ifenable_thinking is defined and enable_thinking istrue%}{{-think\n}}{%-else%}{{-think\n\n/think\n\n}}{%- endif %}{%- endif %}将刚才生成的jinja模板保存到文件qwen.jinja中我们可以在vLLM启动模型的时候使用该对话模板vllm serve /home/gillbert/code/vllm_test/llm/models/Qwen--Qwen3.5-0.8B/snapshots/master\--tensor-parallel-size1\--gpu-memory-utilization0.9\--max-model-len8192\--host0.0.0.0\--port8000\--api-key123456\--enable-auto-tool-choice\--tool-call-parser hermes\--chat-template ./qwen.jinja我们可以使用Open WebUI测试一下
返回列表