Deprecated: The each() function is deprecated. This message will be suppressed on further calls in /home/zhenxiangba/zhenxiangba.com/public_html/phproxy-improved-master/index.php on line 456
CN120218182A - Model distillation method, response information generation method and device - Google Patents
[go: Go Back, main page]

CN120218182A - Model distillation method, response information generation method and device - Google Patents

Model distillation method, response information generation method and device Download PDF

Info

Publication number
CN120218182A
CN120218182A CN202510330312.7A CN202510330312A CN120218182A CN 120218182 A CN120218182 A CN 120218182A CN 202510330312 A CN202510330312 A CN 202510330312A CN 120218182 A CN120218182 A CN 120218182A
Authority
CN
China
Prior art keywords
information
sample
language model
answer
reasoning
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
CN202510330312.7A
Other languages
Chinese (zh)
Inventor
薛宪巍
何欣燃
潘秋桐
何伯磊
陈坤斌
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Baidu Netcom Science and Technology Co Ltd
Original Assignee
Beijing Baidu Netcom Science and Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Baidu Netcom Science and Technology Co Ltd filed Critical Beijing Baidu Netcom Science and Technology Co Ltd
Priority to CN202510330312.7A priority Critical patent/CN120218182A/en
Publication of CN120218182A publication Critical patent/CN120218182A/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/096Transfer learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N5/00Computing arrangements using knowledge-based models
    • G06N5/04Inference or reasoning models
    • G06N5/041Abduction

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Software Systems (AREA)
  • Physics & Mathematics (AREA)
  • Artificial Intelligence (AREA)
  • Evolutionary Computation (AREA)
  • Computing Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Mathematical Physics (AREA)
  • Data Mining & Analysis (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Medical Informatics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Biomedical Technology (AREA)
  • Biophysics (AREA)
  • General Health & Medical Sciences (AREA)
  • Molecular Biology (AREA)
  • Vaporization, Distillation, Condensation, Sublimation, And Cold Traps (AREA)
  • Feedback Control In General (AREA)

Abstract

本申请公开了模型蒸馏方法、答复信息生成方法及装置,涉及计算机技术领域,尤其涉及自然语言处理、深度学习、大模型等人工智能领域。具体实现方案为:获取样本问题信息及样本问题信息的样本答复信息;其中,样本答复信息包括样本推理过程信息及样本答案信息;根据样本问题信息,采用初始的小语言模型,获取样本问题信息对应的预测推理过程;其中,初始的小语言模型的模型规模小于大语言模型的模型规模;根据样本问题信息,采用初始的小语言模型,获取样本问题信息的预测答案信息;根据样本推理过程信息、样本答案信息、预测推理过程信息及预测答案信息,对初始的小语言模型进行训练,以获取经训练的小语言模型。

The present application discloses a model distillation method, a reply information generation method and a device, and relates to the field of computer technology, and in particular to artificial intelligence fields such as natural language processing, deep learning, and large models. The specific implementation scheme is: obtaining sample question information and sample reply information of the sample question information; wherein the sample reply information includes sample reasoning process information and sample answer information; according to the sample question information, using an initial small language model, obtaining the prediction reasoning process corresponding to the sample question information; wherein the model scale of the initial small language model is smaller than the model scale of the large language model; according to the sample question information, using an initial small language model, obtaining the prediction answer information of the sample question information; according to the sample reasoning process information, the sample answer information, the prediction reasoning process information and the prediction answer information, the initial small language model is trained to obtain a trained small language model.

Description

Model distillation method, reply information generation method and device
Technical Field
The application relates to the technical field of computers, in particular to the artificial intelligence fields of natural language processing, deep learning, large models and the like, and particularly relates to a model distillation method, a reply information generation method and a device.
Background
In the field of artificial intelligence, knowledge of a large language model can be transferred to a small language model through model distillation, so as to improve the reasoning efficiency and performance of the model, and meanwhile, maintain the performance similar to that of the large language model.
Disclosure of Invention
The application provides a model distillation method, a reply information generation method and a device. The specific scheme is as follows:
according to an aspect of the present application, there is provided a model distillation method comprising:
Sample answer information of sample question information and sample question information is obtained, wherein the sample answer information comprises sample reasoning process information and sample answer information, and the sample answer information is output after being processed by a large language model aiming at the sample question information;
according to the sample problem information, an initial small language model is adopted to acquire predictive reasoning process information corresponding to the sample problem, wherein the model scale of the initial small language model is smaller than that of the large language model;
According to the sample question information, the initial small language model is adopted to obtain the predicted answer information of the sample question information;
and training the initial small language model according to the sample reasoning process information, the sample answer information, the prediction reasoning process information and the prediction answer information to obtain a trained small language model.
According to another aspect of the present application, there is provided a reply information generation method including:
Acquiring problem information to be processed;
and according to the to-be-processed problem information, adopting an inference algorithm of a small language model to acquire the reply information of the to-be-processed problem information, wherein the small language model is acquired by adopting the distillation method.
According to another aspect of the present application, there is provided a model distillation apparatus comprising:
The system comprises a first acquisition module, a second acquisition module and a first processing module, wherein the first acquisition module is used for acquiring sample question information and sample answer information of the sample question information, the sample answer information comprises sample reasoning process information and sample answer information, and the sample answer information is output after being processed by a large language model aiming at the sample question information;
the second acquisition module is used for acquiring prediction reasoning process information corresponding to the sample problem information by adopting an initial small language model according to the sample problem information, wherein the model scale of the initial small language model is smaller than that of the large language model;
The third acquisition module is used for acquiring the predicted answer information of the sample question information by adopting the initial small language model according to the sample question information;
And the training module is used for training the initial small language model according to the sample reasoning process information, the sample answer information, the prediction reasoning process information and the prediction answer information so as to acquire a trained small language model.
According to another aspect of the present application, there is provided a reply information generation apparatus including:
The first acquisition module is used for acquiring the problem information to be processed;
The second acquisition module is used for acquiring the reply information of the to-be-processed problem information by adopting an inference algorithm of a small language model according to the to-be-processed problem information, wherein the small language model is acquired by adopting the distillation method.
According to another aspect of the present application, there is provided an electronic apparatus including:
At least one processor, and
A memory communicatively coupled to the at least one processor, wherein,
The memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of the above embodiments.
According to another aspect of the present application, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the method according to the above-described embodiments.
According to another aspect of the application, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the steps of the method described in the above embodiments.
It should be understood that the description in this section is not intended to identify key or critical features of the embodiments of the application or to delineate the scope of the application. Other features of the present application will become apparent from the description that follows.
Drawings
The drawings are included to provide a better understanding of the present application and are not to be construed as limiting the application. Wherein:
FIG. 1 is a schematic flow chart of a model distillation method according to an embodiment of the present application;
FIG. 2 is a schematic flow chart of a model distillation method according to another embodiment of the present application;
FIG. 3 is a schematic flow chart of a model distillation method according to another embodiment of the present application;
FIG. 4 is a flowchart of a reply message generation method according to an embodiment of the present application;
FIG. 5 is a flowchart of a reply information generation method according to another embodiment of the present application;
FIG. 6 is a schematic diagram of a model distillation apparatus according to an embodiment of the present application;
Fig. 7 is a schematic diagram of a reply information generating device according to an embodiment of the present application;
fig. 8 is a block diagram of an electronic device for implementing a model distillation method of an embodiment of the present application.
Detailed Description
Exemplary embodiments of the present application will now be described with reference to the accompanying drawings, in which various details of the embodiments of the present application are included to facilitate understanding, and are to be considered merely exemplary. Accordingly, those of ordinary skill in the art will recognize that various changes and modifications of the embodiments described herein can be made without departing from the scope and spirit of the application. Also, descriptions of well-known functions and constructions are omitted in the following description for clarity and conciseness.
It should be noted that, in the technical scheme of the application, the acquisition, storage, use, processing and the like of the data all conform to the relevant regulations of national laws and regulations, and the public sequence is not violated.
The model distillation method, reply information generation method, apparatus, electronic device, and storage medium of the embodiments of the present application are described below with reference to the accompanying drawings.
In some embodiments, the model may be simplified by learning the answer portion of the large language model output. However, this distillation approach can make small language models more efficient in providing simple answers, but lacks sufficient deep reasoning capability in dealing with complex questions.
Based on the above, the embodiment of the application provides a model distillation method. Fig. 1 is a schematic flow chart of a model distillation method according to an embodiment of the application.
The model distillation method of the embodiment of the application can be executed by the model distillation device of the embodiment of the application, and the device can be configured in electronic equipment.
The electronic device may be any device with computing capability, for example, may be a personal computer, a mobile terminal, a server, etc., and the mobile terminal may be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, etc., which have various operating systems, touch screens, and/or display screens.
As shown in fig. 1, the model distillation method includes:
Step 101, sample answer information of sample question information and sample question information is obtained.
The sample reply information may include sample reasoning process information and sample answer information, among others. The sample reply information may be reply information output after processing the sample question information using a large language model, for example.
That is, the output of the large language model can be divided into two parts, an inference part and an answer part in the present application. The inference component can be intermediate steps in the inference process, inference links, and internal inference logic of the model. The inference section output contains how the model is inferred through multiple intermediate inference steps, which is typically more complex and computationally intensive. The answer portion is a compact, clear answer or conclusion of the final output. The answer part is usually compressed and refined by long-time reasoning, directly answers the user query, and the structure is simpler.
Step 102, according to the sample problem information, an initial small language model is adopted to obtain the predictive reasoning process information corresponding to the sample problem information.
Wherein the model size of the small language model is smaller than the large language model. For example, the model size may be measured by the number of parameters of the model, with the number of parameters of the small language model being smaller than the number of parameters of the large language model. Alternatively, the number of parameters of the small language model and the number of parameters of the large language model may be both greater than the preset number.
In the application, the training tasks of the small language model can comprise an reasoning task and an answer task, and the small language model can be trained in parallel. By way of example, the inference task may be referred to as a thinking task.
The inference task may refer to an inference link in which a small language model is to learn a large language model in the training process, for example, how to gradually generate intermediate inference steps from input information, so as to derive a final conclusion.
The answer task may require the small language model to learn how to generate the final compact answer from the output answers of the large language model. Answer task training focuses on learning how to quickly generate accurate conclusions without the need for intermediate steps of deep reasoning.
For the reasoning task, the application can adopt the initial small language model to acquire the solution step of the sample problem information, namely the forecast reasoning process information. The prediction reasoning process information can be the reasoning process of the sample problem information and the small language model output by the pointer.
For example, prompt information for indicating output reasoning process information can be obtained based on sample problem information, and the prompt information is processed by adopting a small language model to obtain prediction reasoning process information.
For example, the "please output the solution step or solution idea of the problem Q1" is input into the small language model, so as to obtain the prediction reasoning process corresponding to the problem Q1.
And step 103, according to the sample question information, acquiring the predicted answer information of the sample question information by adopting an initial small language model.
The predicted answer information may be an answer output by the small language model, indicating sample question information.
For the answer task, the sample question information can be directly input into the small language model, so that the small language model can be used for solving the sample question information to output the predicted answer information, or the prompt information for indicating the small language model to output the answer information can be obtained based on the sample question information, and the small language model is used for processing the prompt information to obtain the predicted answer information.
And step 104, training the initial small language model according to the sample reasoning process information, the sample answer information, the prediction reasoning process information and the prediction answer information to obtain a trained small language model.
According to the application, model loss can be determined according to the prediction reasoning process information, the prediction answer information, the sample reasoning process information and the sample answer information, parameters of an initial small language model are adjusted according to the model loss, and training is continued on the small language model with the parameters adjusted until the training ending condition is met, so that a trained small language model is obtained, and distillation of a large language model is realized. Thus, the small language model can learn the reasoning algorithm of the large language model and the ability to generate answers.
The training ending condition may be that the training number reaches a preset number, or that the model loss is smaller than a preset threshold, or other conditions, which are not limited.
The model distillation method can be widely applied to a plurality of fields, for example, a trained small language model can be used for constructing a more intelligent and accurate question-answering system, or applied to an intelligent search engine, an intelligent customer service system and the like.
For example, in a question-and-answer scenario, a question input by a user may be acquired, and the question input by the user is processed by using a trained small language model to acquire answer information output by the small language model.
For another example, questions in the form of text entered by the user may be obtained, and text describing the questions may be processed using a trained small language model to obtain answer information output by the small language model.
For another example, a question input by the user may be acquired, where the description of the question includes an image, the image may be identified using a small trained language model, and the identified content may be processed to acquire reply information output by the small language model.
In the embodiment of the application, sample answer information output by a large language model aiming at sample question information is divided into sample reasoning process information and sample answer information, initial small language models are respectively adopted for the sample question information to acquire prediction reasoning process information and prediction answer information, and the initial small language models are trained based on the prediction reasoning process information, the prediction answer information, the sample reasoning process information and the sample answer information. Therefore, the predictive reasoning process information and the predictive answer information are obtained separately through the initial small language model, and the initial small language model is trained by combining the sample reasoning process information and the sample answer information, so that the small language model can learn the reasoning process information of the sample question information and also can learn the sample answer information of the sample question information, deep reasoning and answer generation can be realized by the small language model through multi-task learning, and the accuracy of answer information output by the model is improved.
Fig. 2 is a schematic flow chart of a model distillation method according to another embodiment of the present application.
As shown in fig. 2, the model distillation method includes:
step 201, sample answer information of sample question information is obtained.
Step 202, according to sample problem information, an initial small language model is adopted to obtain prediction reasoning process information corresponding to the sample problem information.
And 203, acquiring predicted answer information of the sample question information by adopting an initial small language model according to the sample question information.
In the present application, any implementation manner of the embodiments of the present application may be adopted in step 201 to step 203, so that the description thereof is omitted here.
Step 204, determining a first loss based on a difference between the predictive reasoning process information and the sample reasoning process information.
In the application, for the reasoning task, the first loss corresponding to the reasoning task can be determined by using the loss function according to the prediction reasoning process information and the sample reasoning process information. The loss function may be, for example, a mean square error, a cross entropy loss, or the like, or may be other loss functions, and may be determined according to actual needs.
Step 205, determining a second loss according to the difference between the predicted answer information and the sample answer information.
In the application, for the answer task, the second loss corresponding to the answer task can be determined by using the loss function according to the sample reasoning process information and the prediction reasoning process information. The loss function used to calculate the second loss and the first loss may be the same or different, and is not limited thereto.
Step 206, training the initial small language model according to the first loss and the second loss to obtain a trained small language model.
As a possible implementation manner, the first loss and the second loss may be added to obtain a loss sum of the two losses, parameters of the initial small language model are adjusted according to the loss sum, and training is continued on the small language model with the parameters adjusted until the training end condition is met, so as to obtain the trained small language model.
Considering that the components occupied by the reasoning task and the answer task in the model training may be different, as another possible implementation manner, a first weight corresponding to the first loss and a second weight corresponding to the second loss may be determined, the first loss and the second loss are weighted according to the first weight and the second weight, so as to obtain a total loss, and the initial small language model is trained according to the total loss, so as to obtain a trained small language model. Therefore, the learning key points of the model can be adjusted by adjusting the weights of the first loss and the second loss, so that different training requirements can be met, and the accuracy of the model is improved.
The first weight and the second weight may be the same or different, which is not limited.
Since the complexity of the questions may be different, the reasoning process of the questions may be relatively simple, and the reasoning process of the questions may be relatively complex, it is exemplary that the first attribute information of the sample question information may be determined, the first complexity of the sample question information may be determined according to the first attribute information, and then the first weight and the second weight may be determined according to the first complexity.
For example, a mapping relationship between complexity and two losses may be pre-established, and the mapping relationship is queried according to the first complexity to determine the first weight and the second weight.
By way of example, the first attribute information may include, but is not limited to, a type, a length, a number of sub-questions, etc. of the sample question information. For example, if the type of sample problem information is a query class, the first complexity is relatively low and the first weight may be less than the second weight. As another example, if the sample problem is relatively long, the first complexity is relatively high, the first weight may be greater than the second weight, etc.
Therefore, the complexity of the sample question information is determined based on the attribute information of the sample question information, the weight of the loss of the reasoning task and the weight of the loss of the answer task are determined based on the complexity of the sample question information, the accuracy of the weight is improved, and the accuracy of the model is further improved.
Because the requirements on the reasoning capacity of the model may be different in different application scenarios, such as a mathematical calculation type scenario, the requirements on the reasoning capacity of the model are relatively high, and the requirements on the reasoning capacity of the model are not high when the class scenario is simply queried, therefore, for example, an application scenario corresponding to an initial small language model can be determined, and the first weight and the second weight are determined according to the application scenario.
For example, a mapping relationship between different application scenarios and two kinds of losses may be established in advance, and then the first weight and the second weight may be determined based on the mapping relationship.
Therefore, the two lost weights are determined according to the application scene of the small language model, so that the accuracy of the weights can be improved, and the requirement of the small language model in the application scene can be met.
In the implementation of the application, the thinking task is trained based on the predictive reasoning process information and the sample reasoning process information, and the answer task is trained based on the predictive answer information and the sample answer information, so that the small language model can improve the processing efficiency of processing different tasks on the premise of ensuring the reasoning depth by the multitask training of the thinking task and the answer task, and the dual capability of providing a detailed reasoning process and fast response is realized in the same model.
In addition, by the multitask distillation scheme of reasoning tasks and answer tasks, the adaptability of the model under multitask and multi-scene can be effectively improved, and the service access speed and efficiency of the model are greatly improved.
Fig. 3 is a schematic flow chart of a model distillation method according to another embodiment of the present application.
As shown in fig. 3, the model distillation method includes:
step 301, sample answer information of sample question information is obtained.
In the present application, step 301 may be implemented by any one of the embodiments of the present application, so that the description thereof is omitted here.
And step 302, filling the reasoning task prompt template according to the sample question information and the sample answer information to generate the reasoning task prompt information.
The reasoning task prompt template can be a prompt template for indicating the model to output reasoning process information. By way of example, the inference task tip template may include slot information for questions, answers, and the like. For example, the answer to an inference task prompt template of "question [ is ].
According to the application, the corresponding slot positions in the reasoning task prompt template can be filled according to the sample question information and the sample answer information so as to obtain the reasoning task prompt information.
The reasoning task prompt information can be used for indicating the model to output a reasoning process aiming at the sample problem information.
For example, the reasoning task prompt is "the answer to question Q1 is D1, please explain how to derive this answer".
For another example, the reasoning task prompt is "the answer to question Q1 is D1, please explain the specific steps of getting this answer.
To improve accuracy, exemplary reasoning task cues may include information such as reasoning examples, model output requirements, and the like, in addition to sample question information and sample answer information. The inference examples can include reference questions, answers to the reference questions, inference processes and the like, so that the initial small language model can output inference process information corresponding to sample question information through the reference inference examples.
And 303, processing the reasoning task prompt information by adopting an initial small language model to acquire the predicted reasoning process information.
In the application, the reasoning task prompt information can be input into the initial small language model, the initial small language model is adopted to identify the reasoning task prompt information so as to determine the processing purpose, and the sample question information and the sample answer information in the reasoning task prompt information are searched, so that the prediction reasoning process information is obtained based on the searched knowledge content and the processing purpose.
And step 304, according to the sample question information, acquiring a predicted answer of the sample question information by adopting an initial small language model.
In the present application, step 304 may be implemented by any one of the embodiments of the present application, and thus will not be described herein.
In order to improve the processing efficiency of the model, for example, the answer task prompt template may be filled according to the sample question information to obtain answer task prompt information, and the answer task prompt information is input into an initial small language model, and the answer task prompt information is processed by adopting the initial small language model to obtain predicted answer information.
Wherein the answer task prompt template may be a prompt template for instructing the model to output an answer to the question. Illustratively, the answer task prompt template may include slot information such as questions.
For example, the corresponding slots in the answer task prompt model may be filled according to the sample question information to obtain answer task prompt information.
The answer task prompt information may be used to instruct the model to output answer information of sample question information. For example, the answer task prompt information is "please give an answer to the question Q1".
Optionally, in order to improve accuracy, the answer task prompt information may include information such as answer examples, model output requirements, and the like, in addition to sample question information. The answer examples may include the reference questions and answers to the reference questions, so that the initial small language model may reference the answer examples.
Therefore, based on the answer task prompt information, the initial small language model can determine the requirement of outputting an answer, and the accuracy of an output result is improved.
Step 305, training the initial small language model according to the sample reasoning process information, the sample answer information, the predictive reasoning process information and the predictive answer information to obtain a trained small language model.
In the present application, step 305 may be implemented by any one of the embodiments of the present application, and thus will not be described herein.
Because the reasoning process is a process of deducing the answer, optionally, the parameter of the initial small language model can be adjusted according to the difference between the forecast reasoning process information and the sample reasoning process information and the difference between the forecast answer information and the sample answer information and combining the difference between the answer information deduced according to the forecast reasoning process information and the sample answer information, so that the accuracy of the model can be improved.
According to the embodiment of the application, the reasoning task prompt template is filled according to the sample question information and the sample answer information to obtain the reasoning task prompt information, and the initial small language model outputs the prediction reasoning process information of the sample question information based on the reasoning task prompt information, so that the small language model is guided to pay attention to the middle reasoning process and output the reasoning process information through the reasoning task prompt information, thereby being convenient for the small language model to better learn the reasoning process so as to improve the reasoning capability of the model, for example, the capability of gradually generating the middle reasoning step from the input information of the small language model can be improved.
Fig. 4 is a flowchart of a reply information generation method according to an embodiment of the present application.
As shown in fig. 4, the reply information generation method includes:
step 401, obtaining problem information to be processed.
According to the application, the user can input the problem in the interaction interface of the artificial intelligence service based on the small language model, so that the information of the problem to be processed can be obtained. Or the problem information to be processed may be obtained by other means, which is not limited.
Step 402, according to the to-be-processed question information, adopting an inference algorithm of a small language model to obtain the reply information of the to-be-processed question information.
The small language model can be obtained by training the distillation method described in any embodiment, so that the small language model can learn an inference algorithm of the large language model. The inference algorithm can be used to gradually generate intermediate inference steps from the input information of the small language model, so as to derive a final conclusion. That is, the small language model can learn the reasoning capabilities of the large language model.
In the application, the to-be-processed problem information can be input into the small language model, and the to-be-processed problem information is solved by adopting the reasoning algorithm of the small language model so as to obtain the reply information of the to-be-processed problem information.
The answer information may include answer information of the question information to be processed, or include reasoning process information, answer information, and the like.
For example, a small language model may be used to identify the problem information to be processed to determine the type of the problem information to be processed, and the problem information to be processed is processed according to the type of the problem information to be processed to output reply information matching the type of the problem information to be processed.
For example, if the type of the question information to be processed is a query class, the small language model may directly output an answer, and if the type of the question information to be processed is a mathematical calculation class, the small language model may output reasoning process information and answer information.
In the embodiment of the application, the small language model trained by the distillation method can carry out deep reasoning, so that the reasoning algorithm of the small language model is adopted to process the problem information to be processed, and the accuracy of the reply information can be improved.
Fig. 5 is a flowchart of a reply information generation method according to another embodiment of the present application.
As shown in fig. 5, the reply information generation method includes:
Step 501, obtaining problem information to be processed.
In the present application, step 501 may be implemented by any one of the embodiments of the present application, and thus will not be described herein.
Step 502, determining a processing mode of the problem information to be processed.
The processing modes may include a deep inference mode, a fast response mode, etc.
And the small language model performs reasoning in the deep reasoning mode, and outputs a reasoning process and an answer. The deep inference mode may be applied to a scenario requiring an interpretation or inference process, such as the resolution of complex questions or the handling of multiple rounds of conversations, etc.
The small language model in the fast response mode can directly generate and output answers to the questions. The fast response mode may be suitable for simple query or direct demand scenarios.
As a possible implementation manner, operation information for a target control in the interactive interface may be acquired, and a processing mode is determined according to the operation information.
Wherein the target control may be a control related to a processing mode.
By way of example, the target control may be a control associated with a deep inference mode, which may be determined to be a deep inference mode if a trigger operation is detected for the target control, and a fast response mode if a trigger operation is not detected for the target control.
For example, after a user inputs a problem in the interactive interface, the control "deep thinking" in the interactive interface is triggered, then the processing mode of the problem can be determined to be a deep reasoning mode, and if the user inputs the problem in the interactive interface, the problem submitting control is triggered, and the control "deep thinking" is not triggered, then the processing mode of the problem can be determined to be a quick response mode.
By way of example, the target controls may include controls related to a deep inference mode and controls related to a fast response mode, the processing mode may be determined to be the deep inference mode if a trigger operation is detected for the controls related to the deep inference mode, and the processing mode may be determined to be the fast response mode if a trigger operation is detected for the controls related to the fast response mode.
Therefore, the small language model can dynamically select the processing mode according to the operation information of the target control by switching between different processing modes according to different operation information of the target control, and different problem processing requirements can be met.
As another possible implementation manner, the second attribute information of the problem information to be processed may be determined, the second complexity of the problem information to be processed may be determined according to the second attribute information, and the processing mode of the problem information to be processed may be determined according to the second complexity.
The second attribute information may include, but is not limited to, a type, a length, a number of sub-questions, etc. of the question information to be processed. For example, if the type of the problem information to be processed is a query type, the second complexity is relatively low, and the processing mode is determined to be a fast response mode. For another example, if the problem to be processed is relatively long, the second complexity is relatively high, and the processing mode is determined to be a deep inference mode.
Therefore, the complexity of the problem information to be processed is determined based on the attribute information of the problem information to be processed, and the processing mode is determined based on the complexity, so that the processing mode can be flexibly adjusted according to the complexity of the problem, and different problem processing requirements can be met.
Step 503, obtaining reply information by utilizing the reasoning algorithm of the small language model according to the processing mode and the to-be-processed question information.
In the application, different processing modes can be provided with different prompt templates, reply generation prompt information corresponding to the processing modes can be obtained according to the to-be-processed problem information and the prompt templates corresponding to the processing modes, and the reply generation prompt information is processed by adopting an inference algorithm of a small language model so as to obtain reply information in the processing modes.
The prompting template corresponding to the processing mode can be a prompting template for indicating the model to output reply information matched with the processing mode.
For example, according to the to-be-processed problem information, the processing example corresponding to the processing mode, and the like, the corresponding slot in the prompt template corresponding to the processing mode is filled to obtain the reply generation prompt information corresponding to the processing mode.
Among other examples, processing examples may include referencing a question, referring to reply information to the question, and so forth.
The answer generation prompt information corresponding to the processing mode can be used for indicating the model to output answer information of the to-be-processed problem information in the processing mode.
In some embodiments, if the processing mode is a deep inference mode, the inference algorithm of the small language model may be adopted according to the to-be-processed problem information and the deep inference mode to obtain inference process information for the to-be-processed problem information, and obtain reply information according to the inference process information.
For example, answer information may be determined based on the inference process information, and answer information may be determined based on the inference process information and the answer information. The reply information may include, among other things, reasoning process information and reply information.
For example, if the processing mode is a deep inference mode, reply generation prompt information corresponding to the deep inference mode can be obtained according to the to-be-processed question information and the prompt template corresponding to the deep inference mode, the reply generation prompt information is processed by adopting a small language model, inference process information is obtained, and answer information is obtained according to the inference process information. The answer information may include, among other things, reasoning process information and answer information for the answer.
In some embodiments, if the processing mode of the to-be-processed question information is a fast response mode, such as query-type questions, the small language model may directly generate answer information of the to-be-processed question information without invoking an inference algorithm, and output the answer information, so that the response speed can be improved, and resources can be saved. Therefore, the small language model can not only realize the deep reasoning under the complex task, but also improve the response speed under the simple task. Therefore, prompt information can be generated based on the answers corresponding to the processing modes, and the answer information matched with the processing modes can be obtained by adopting the small language model, so that the accuracy of the answer information in different processing modes can be improved.
In the embodiment of the application, the accuracy of the reply information can be improved by determining the processing mode of the to-be-processed problem information and acquiring the reply information of the to-be-processed problem information by adopting the small language model according to the to-be-processed problem information and combining the processing mode.
Optionally, in order to improve accuracy of the reply information, if complexity of the to-be-processed question information is low, such as querying a question, an inference algorithm of a small language model may be adopted to obtain inference process information for the to-be-processed question information, determine answer information of the to-be-processed question information according to the inference process information, and output the answer information.
According to the model distillation method, through the multi-task distillation scheme of the reasoning task and the answer task, not only can the deep reasoning capacity and the quick response capacity in the model reasoning process be separated, but also the processing mode can be flexibly switched to adapt to different application scenes, so that the model can provide accurate answers when the deep reasoning is required, high-efficiency answers can be provided when the quick response is required, and the balance problem between the high-efficiency reasoning and the deep reasoning is solved.
In addition, a mode switching mechanism can be introduced in the multitasking training method based on the reasoning task and the answer task, so that the model can dynamically select a processing mode to cope with different business requirements.
If the problem requires higher accuracy and more complex reasoning, the model may enter a deep reasoning mode, the model may learn and perform multi-step reasoning, and finally output detailed answers through the deep reasoning. If the question is relatively simple or quick feedback is required, the model may enter a quick response mode, quickly generate and output an answer, and may not make complex reasoning. The mode switching mechanism is introduced, so that the processing strategy can be flexibly adjusted according to requirements in different task environments, and the application efficiency and response speed of the model can be improved.
In order to realize the embodiment, the embodiment of the application also provides a model distillation device. Fig. 6 is a schematic structural diagram of a model distillation apparatus according to an embodiment of the present application.
As shown in fig. 6, the model distillation apparatus 600 includes:
The first obtaining module 610 is configured to obtain sample question information and sample answer information of the sample question information, where the sample answer information includes sample reasoning process information and sample answer information, and the sample answer information is output after being processed by a large language model for the sample question information;
a second obtaining module 620, configured to obtain, according to the sample problem information, prediction reasoning process information corresponding to the sample problem information by using an initial small language model, where a model scale of the initial small language model is smaller than a model scale of the large language model;
A third obtaining module 630, configured to obtain, according to the sample question information, predicted answer information of the sample question information by using the initial small language model;
the training module 640 is configured to train the initial small language model according to the sample reasoning process information, the sample answer information, the predictive reasoning process information, and the predictive answer information, so as to obtain a trained small language model.
Optionally, the training module 640 is configured to:
determining a first loss based on a difference between the predictive reasoning process information and the sample reasoning process information;
Determining a second loss according to the difference between the predicted answer information and the sample answer information;
Training the initial small language model according to the first loss and the second loss to obtain the trained small language model.
Optionally, the training module 640 is configured to:
Determining a first weight corresponding to the first loss and a second weight corresponding to the second loss;
weighting the first loss and the second loss according to the first weight and the second weight to obtain total loss;
Training the initial small language model according to the total loss to obtain the trained small language model.
Optionally, the training module 640 is configured to:
Determining first attribute information of the sample problem information;
Determining a first complexity of the sample problem information according to the first attribute information;
And determining the first weight and the second weight according to the first complexity.
Optionally, the training module 640 is configured to:
Determining an application scene corresponding to the initial small language model;
and determining the first weight and the second weight according to the application scene.
Optionally, the second obtaining module 620 is configured to:
filling an reasoning task prompt template according to the sample question information and the sample answer information to generate reasoning task prompt information;
And processing the reasoning task prompt information by adopting the initial small language model to acquire the prediction reasoning process.
Optionally, a third obtaining module 630 is configured to:
filling an answer task prompt template according to the sample question information to obtain answer task prompt information;
And processing the answer task prompt information by adopting the initial small language model to acquire the predicted answer information.
It should be noted that the explanation of the embodiment of the model distillation method is also applicable to the model distillation apparatus of this embodiment, and thus will not be repeated here.
In the embodiment of the application, sample answer information output by a large language model aiming at sample question information is divided into sample reasoning process information and sample answer information, initial small language models are respectively adopted for the sample question information to acquire prediction reasoning process information and prediction answer information, and the initial small language models are trained based on the prediction reasoning process information, the prediction answer information, the sample reasoning process information and the sample answer information. Therefore, the predictive reasoning process information and the predictive answer information are obtained separately through the initial small language model, and the initial small language model is trained by combining the sample reasoning process information and the sample answer information, so that the small language model can learn the reasoning process information of the sample question information and also can learn the sample answer information of the sample question information, deep reasoning and answer generation can be realized by the small language model through multi-task learning, and the accuracy of answer information output by the model is improved.
In order to achieve the above embodiments, the embodiments of the present application further provide a reply information generating device. Fig. 7 is a schematic diagram of a reply information generating apparatus according to an embodiment of the present application.
As shown in fig. 7, the reply information generation apparatus 700 includes:
a first obtaining module 710, configured to obtain information of a problem to be processed;
The second obtaining module 720 is configured to obtain, according to the to-be-processed question information, answer information of the to-be-processed question information by using an inference algorithm of a small language model, where the small language model is obtained by using the distillation method described in any one of the embodiments.
Optionally, the second obtaining module 720 is configured to:
determining a processing mode of the problem information to be processed;
Responding to the processing mode as a deep reasoning mode, and acquiring reasoning process information aiming at the problem information to be processed by utilizing a reasoning algorithm of the small language model according to the deep reasoning mode and the problem information to be processed;
and acquiring the reply information according to the reasoning process information.
Optionally, the second obtaining module 720 is configured to:
acquiring operation information aiming at a target control in an interactive interface;
and determining the processing mode according to the operation information.
Optionally, the second obtaining module 720 is configured to:
determining second attribute information of the problem information to be processed;
determining a second complexity of the problem information to be processed according to the second attribute information;
and determining the processing mode according to the second complexity.
Optionally, the second obtaining module 720 is configured to:
generating a reply generation prompt message corresponding to the depth reasoning mode according to the to-be-processed problem information and the prompt template corresponding to the depth reasoning mode;
And processing the reply generation prompt information by adopting an inference algorithm of the small language model so as to acquire inference process information.
Note that, the explanation of the foregoing embodiment of the reply information generation method is also applicable to the reply information generation apparatus of this embodiment, and therefore will not be described in detail here.
In the embodiment of the application, the small language model trained by the distillation method can carry out deep reasoning, so that the reasoning algorithm of the small language model is adopted to process the problem information to be processed, and the accuracy of the reply information can be improved.
According to embodiments of the present application, the present application also provides an electronic device, a readable storage medium and a computer program product.
FIG. 8 illustrates a schematic block diagram of an example electronic device 800 that may be used to implement an embodiment of the application. Electronic devices are intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be exemplary only, and are not meant to limit implementations of the applications described and/or claimed herein.
As shown in fig. 8, the apparatus 800 includes a computing unit 801 that can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 802 or a computer program loaded from a storage unit 808 into a RAM (Random Access Memory ) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other by a bus 804. An I/O (Input/Output) interface 805 is also connected to bus 804.
Various components in the device 800 are connected to the I/O interface 805, including an input unit 806, such as a keyboard, a mouse, etc., an output unit 807, such as various types of displays, speakers, etc., a storage unit 808, such as a magnetic disk, optical disk, etc., and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information/data with other devices via a computer network such as the internet and/or various telecommunication networks.
The computing unit 801 may be a variety of general and/or special purpose processing components having processing and computing capabilities. Some examples of computing unit 801 include, but are not limited to, a CPU (Central Processing Unit ), a GPU (Graphic Processing Units, graphics processing unit), various specialized AI (ARTIFICIAL INTELLIGENCE ) computing chips, various computing units running machine learning model algorithms, DSPs (DIGITAL SIGNAL Processor ), and any suitable Processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as a model distillation method. For example, in some embodiments, the model distillation method may be implemented as a computer software program tangibly embodied on a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program may be loaded and/or installed onto device 800 via ROM 802 and/or communication unit 809. When a computer program is loaded into RAM 803 and executed by computing unit 801, one or more steps of the model distillation method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the model distillation method by any other suitable means (e.g., by means of firmware).
Various implementations of the systems and techniques described here above can be implemented in digital electronic circuitry, integrated Circuit System, FPGA (Field Programmable GATE ARRAY ), ASIC (Application-SPECIFIC INTEGRATED Circuit, application-specific integrated Circuit), ASSP (Application SPECIFIC STANDARD Product, application-specific standard Product), SOC (System On Chip ), CPLD (Complex Programmable Logic Device, complex programmable logic device), computer hardware, firmware, software, and/or combinations thereof. These various embodiments may include being implemented in one or more computer programs that are executable and/or interpretable on a programmable system including at least one programmable processor, which may be a special or general purpose programmable processor, operable to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
Program code for carrying out methods of the present application may be written in any combination of one or more programming languages. These program code may be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus such that the program code, when executed by the processor or controller, causes the functions/operations specified in the flowchart and/or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package, partly on the machine and partly on a remote machine or entirely on the remote machine or server.
In the context of the present application, a machine-readable medium may be a tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, RAM, ROM, EPROM (ELECTRICALLY PROGRAMMABLE READ-Only-Memory, erasable programmable read-Only Memory) or flash Memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid CRYSTAL DISPLAY) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide for interaction with a user, for example, feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user may be received in any form, including acoustic input, speech input, or tactile input.
The systems and techniques described here can be implemented in a computing system that includes a background component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front-end component (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such background, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include LAN (Local Area Network ), WAN (Wide Area Network, wide area network), the Internet, and blockchain networks.
The computer system may include a client and a server. The client and server are typically remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also called a cloud computing server or a cloud host, and is a host product in a cloud computing service system, so as to solve the defects of high management difficulty and weak service expansibility in the traditional physical hosts and Virtual service (Virtual PRIVATE SERVER, virtual special servers). The server may also be a server of a distributed system or a server that incorporates a blockchain.
It should be noted that, the electronic device for implementing the reply information generation method according to the embodiment of the present application is similar to the above-mentioned electronic device, and therefore will not be described herein.
According to an embodiment of the present application, there is also provided a computer program product which, when executed by an instruction processor in the computer program product, performs the model distillation method, or reply information generation method, set forth in the above embodiment of the present application.
It should be appreciated that various forms of the flows shown above may be used to reorder, add, or delete steps. For example, the steps described in the present application may be performed in parallel, sequentially, or in a different order, so long as the desired results of the technical solution disclosed in the present application can be achieved, and are not limited herein.
The above embodiments do not limit the scope of the present application. It will be apparent to those skilled in the art that various modifications, combinations, sub-combinations and alternatives are possible, depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the scope of the present application.

Claims (17)

1.一种模型蒸馏方法,包括:1. A model distillation method, comprising: 获取样本问题信息及所述样本问题信息的样本答复信息;其中,所述样本答复信息包括样本推理过程信息及样本答案信息,所述样本答复信息是由大语言模型针对所述样本问题信息处理后输出的;Acquire sample question information and sample answer information of the sample question information; wherein the sample answer information includes sample reasoning process information and sample answer information, and the sample answer information is output after the large language model processes the sample question information; 根据所述样本问题信息,采用初始的小语言模型,获取所述样本问题信息对应的预测推理过程信息;其中,所述初始的小语言模型的模型规模小于所述大语言模型的模型规模;According to the sample question information, an initial small language model is used to obtain prediction reasoning process information corresponding to the sample question information; wherein the model scale of the initial small language model is smaller than the model scale of the large language model; 根据所述样本问题信息,采用所述初始的小语言模型,获取所述样本问题信息的预测答案信息;According to the sample question information, using the initial small language model, obtaining predicted answer information of the sample question information; 根据所述样本推理过程信息、所述样本答案信息、所述预测推理过程信息及所述预测答案信息,对所述初始的小语言模型进行训练,以获取经训练的小语言模型。The initial small language model is trained according to the sample reasoning process information, the sample answer information, the predicted reasoning process information and the predicted answer information to obtain a trained small language model. 2.如权利要求1所述的方法,其中,所述根据所述样本推理过程信息、所述样本答案信息、所述预测推理过程信息及所述预测答案信息,对所述初始的小语言模型进行训练,以获取经训练的小语言模型,包括:2. The method according to claim 1, wherein the step of training the initial small language model according to the sample reasoning process information, the sample answer information, the predicted reasoning process information and the predicted answer information to obtain a trained small language model comprises: 根据所述预测推理过程信息与所述样本推理过程信息之间的差异,确定第一损失;Determining a first loss according to a difference between the prediction reasoning process information and the sample reasoning process information; 根据所述预测答案信息与所述样本答案信息之间的差异,确定第二损失;determining a second loss according to a difference between the predicted answer information and the sample answer information; 根据所述第一损失及所述第二损失,对所述初始的小语言模型进行训练,以获取所述经训练的小语言模型。The initial small language model is trained according to the first loss and the second loss to obtain the trained small language model. 3.如权利要求2所述的方法,其中,所述根据所述第一损失及所述第二损失,对所述初始的小语言模型进行训练,以获取所述经训练的小语言模型,包括:3. The method according to claim 2, wherein the step of training the initial small language model according to the first loss and the second loss to obtain the trained small language model comprises: 确定所述第一损失对应的第一权重及所述第二损失对应的第二权重;Determine a first weight corresponding to the first loss and a second weight corresponding to the second loss; 根据所述第一权重及所述第二权重,对所述第一损失及所述第二损失进行加权,得到总损失;According to the first weight and the second weight, weighting the first loss and the second loss to obtain a total loss; 根据所述总损失,对所述初始的小语言模型进行训练,以获取所述经训练的小语言模型。The initial small language model is trained according to the total loss to obtain the trained small language model. 4.如权利要求3所述的方法,其中,所述确定所述第一损失对应的第一权重及所述第二损失对应的第二权重,包括:4. The method of claim 3, wherein determining a first weight corresponding to the first loss and a second weight corresponding to the second loss comprises: 确定所述样本问题信息的第一属性信息;Determine first attribute information of the sample question information; 根据所述第一属性信息,确定所述样本问题信息的第一复杂度;Determining a first complexity of the sample question information according to the first attribute information; 根据所述第一复杂度,确定所述第一权重及所述第二权重。The first weight and the second weight are determined according to the first complexity. 5.如权利要求3所述的方法,其中,所述确定所述第一损失对应的第一权重及所述第二损失对应的第二权重,包括:5. The method of claim 3, wherein determining a first weight corresponding to the first loss and a second weight corresponding to the second loss comprises: 确定所述初始的小语言模型对应的应用场景;Determine an application scenario corresponding to the initial small language model; 根据所述应用场景,确定所述第一权重及所述第二权重。According to the application scenario, the first weight and the second weight are determined. 6.如权利要求1-5中任一项所述的方法,其中,所述根据所述样本问题信息,采用初始的小语言模型,获取所述样本问题信息对应的预测推理过程信息,包括:6. The method according to any one of claims 1 to 5, wherein the step of acquiring the prediction reasoning process information corresponding to the sample question information by using an initial small language model according to the sample question information comprises: 根据所述样本问题信息及所述样本答案信息,对推理任务提示模板进行填充,以生成推理任务提示信息;Filling a reasoning task prompt template according to the sample question information and the sample answer information to generate reasoning task prompt information; 采用所述初始的小语言模型,对所述推理任务提示信息进行处理,以获取所述预测推理过程信息。The initial small language model is used to process the reasoning task prompt information to obtain the predictive reasoning process information. 7.如权利要求1-5中任一项所述的方法,其中,所述根据所述样本问题信息,采用所述初始的小语言模型,获取所述样本问题信息的预测答案信息,包括:7. The method according to any one of claims 1 to 5, wherein the step of obtaining predicted answer information of the sample question information by using the initial small language model according to the sample question information comprises: 根据所述样本问题信息,对答案任务提示模板进行填充,以获取答案任务提示信息;Filling the answer task prompt template according to the sample question information to obtain answer task prompt information; 采用所述初始的小语言模型,对所述答案任务提示信息进行处理,以获取所述预测答案信息。The initial small language model is used to process the answer task prompt information to obtain the predicted answer information. 8.一种答复信息生成方法,包括:8. A method for generating a reply message, comprising: 获取待处理问题信息;Get information about pending issues; 根据所述待处理问题信息,采用小语言模型的推理算法,以获取所述待处理问题信息的答复信息;其中,所述小语言模型是采用权利要求1-7中任一项所述的方法获取的。According to the question information to be processed, an inference algorithm of a small language model is used to obtain answer information of the question information to be processed; wherein the small language model is obtained by using the method described in any one of claims 1-7. 9.如权利要求8所述的方法,其中,所述根据所述待处理问题信息,采用小语言模型,以获取所述待处理问题信息的答复信息,包括:9. The method according to claim 8, wherein the step of using a small language model to obtain answer information of the question to be processed comprises: 确定所述待处理问题信息的处理模式;Determine a processing mode for the problem information to be processed; 响应于所述处理模式为深度推理模式,根据所述深度推理模式及所述待处理问题信息,利用所述小语言模型的推理算法,获取针对所述待处理问题信息的推理过程信息;In response to the processing mode being the deep reasoning mode, according to the deep reasoning mode and the problem information to be processed, using the reasoning algorithm of the small language model, obtaining reasoning process information for the problem information to be processed; 根据所述推理过程信息,获取所述答复信息。The answer information is obtained according to the reasoning process information. 10.如权利要求9所述的方法,其中,所述确定所述待处理问题信息的处理模式,包括:10. The method according to claim 9, wherein the step of determining a processing mode of the problem information to be processed comprises: 获取针对交互界面中目标控件的操作信息;Get the operation information of the target control in the interactive interface; 根据所述操作信息,确定所述处理模式。The processing mode is determined according to the operation information. 11.权利要求9所述的方法,其中,所述确定所述待处理问题信息的处理模式,包括:11. The method of claim 9, wherein the step of determining a processing mode of the problem information to be processed comprises: 确定所述待处理问题信息的第二属性信息;Determine second attribute information of the problem information to be processed; 根据所述第二属性信息,确定所述待处理问题信息的第二复杂度;Determining a second complexity of the problem information to be processed according to the second attribute information; 根据所述第二复杂度,确定所述处理模式。The processing mode is determined according to the second complexity. 12.如权利要求9所述的方法,其中,所述根据所述深度推理模式及所述待处理问题信息,利用所述小语言模型的推理算法,获取针对所述待处理问题信息的推理过程信息,包括:12. The method of claim 9, wherein the step of obtaining the reasoning process information for the problem information to be processed by using the reasoning algorithm of the small language model according to the deep reasoning mode and the problem information to be processed comprises: 根据所述待处理问题信息及所述深度推理模式对应的提示模板,生成所述深度推理模式对应的答复生成提示信息;Generate prompt information for answer generation corresponding to the deep reasoning mode according to the problem information to be processed and the prompt template corresponding to the deep reasoning mode; 采用所述小语言模型的推理算法,对所述答复生成提示信息进行处理,以获取所述推理过程信息。The answer generation prompt information is processed using the inference algorithm of the small language model to obtain the inference process information. 13.一种模型蒸馏装置,包括:13. A model distillation apparatus comprising: 第一获取模块,用于获取样本问题信息及所述样本问题信息的样本答复信息;其中,所述样本答复信息包括样本推理过程信息及样本答案信息,所述样本答复信息是由大语言模型针对所述样本问题信息处理后输出的;A first acquisition module is used to acquire sample question information and sample answer information of the sample question information; wherein the sample answer information includes sample reasoning process information and sample answer information, and the sample answer information is output after the large language model processes the sample question information; 第二获取模块,用于根据所述样本问题信息,采用初始的小语言模型,获取所述样本问题信息对应的预测推理过程信息;其中,所述初始的小语言模型的模型规模小于所述大语言模型的模型规模;A second acquisition module is used to acquire prediction and reasoning process information corresponding to the sample question information by using an initial small language model according to the sample question information; wherein the model scale of the initial small language model is smaller than the model scale of the large language model; 第三获取模块,用于根据所述样本问题信息,采用所述初始的小语言模型,获取所述样本问题信息的预测答案信息;A third acquisition module is used to acquire predicted answer information of the sample question information by using the initial small language model according to the sample question information; 训练模块,用于根据所述样本推理过程信息、所述样本答案信息、所述预测推理过程信息及所述预测答案信息,对所述初始的小语言模型进行训练,以获取经训练的小语言模型。A training module is used to train the initial small language model according to the sample reasoning process information, the sample answer information, the predicted reasoning process information and the predicted answer information to obtain a trained small language model. 14.一种答复信息生成装置,包括:14. A reply information generating device, comprising: 第一获取模块,用于获取待处理问题信息;The first acquisition module is used to obtain information about problems to be processed; 第二获取模块,用于根据所述待处理问题信息,采用小语言模型的推理算法,以获取所述待处理问题信息的答复信息;其中,所述小语言模型是采用权利要求1-7中任一项所述的方法获取的。The second acquisition module is used to use the inference algorithm of the small language model to obtain the answer information of the question to be processed according to the question information to be processed; wherein the small language model is obtained by using the method described in any one of claims 1-7. 15.一种电子设备,包括:15. An electronic device, comprising: 至少一个处理器;以及at least one processor; and 与所述至少一个处理器通信连接的存储器;其中,a memory communicatively connected to the at least one processor; wherein, 所述存储器存储有可被所述至少一个处理器执行的指令,所述指令被所述至少一个处理器执行,以使所述至少一个处理器能够执行权利要求1-12中任一项所述的方法。The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12. 16.一种存储有计算机指令的非瞬时计算机可读存储介质,其中,所述计算机指令用于使所述计算机执行根据权利要求1-12中任一项所述的方法。16. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 12. 17.一种计算机程序产品,包括计算机程序,所述计算机程序在被处理器执行时实现权利要求1-12中任一项所述方法的步骤。17. A computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 12.
CN202510330312.7A 2025-03-19 2025-03-19 Model distillation method, response information generation method and device Pending CN120218182A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
CN202510330312.7A CN120218182A (en) 2025-03-19 2025-03-19 Model distillation method, response information generation method and device

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
CN202510330312.7A CN120218182A (en) 2025-03-19 2025-03-19 Model distillation method, response information generation method and device

Publications (1)

Publication Number Publication Date
CN120218182A true CN120218182A (en) 2025-06-27

Family

ID=96116390

Family Applications (1)

Application Number Title Priority Date Filing Date
CN202510330312.7A Pending CN120218182A (en) 2025-03-19 2025-03-19 Model distillation method, response information generation method and device

Country Status (1)

Country Link
CN (1) CN120218182A (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN120611192A (en) * 2025-08-12 2025-09-09 上海库帕思科技有限公司 Training data synthesis method, device, medium and program product based on error extrapolation and inference chain analysis

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN120611192A (en) * 2025-08-12 2025-09-09 上海库帕思科技有限公司 Training data synthesis method, device, medium and program product based on error extrapolation and inference chain analysis

Similar Documents

Publication Publication Date Title
CN113377520B (en) Resource scheduling method, device, equipment and storage medium
CN112487173B (en) Man-machine conversation method, device and storage medium
CN110795569B (en) Method, device and device for generating vector representation of knowledge graph
CN113407850B (en) Method and device for determining and acquiring virtual image and electronic equipment
CN112528995B (en) Method for training target detection model, target detection method and device
CN114792359A (en) Rendering network training and virtual object rendering method, device, equipment and medium
CN111241838B (en) Semantic relationship processing method, device and equipment for text entities
CN117574868A (en) Chart generation method, device, equipment and storage medium
EP3933719A2 (en) Method, apparatus, device, storage medium and computer program product for labeling data
CN115510203B (en) Question answer determination methods, devices, equipment, storage media and program products
CN113641829A (en) Method and device for training graph neural network and knowledge graph completion
CN115456167B (en) Lightweight model training method, image processing method, device and electronic equipment
EP4123516A1 (en) Method and apparatus for acquiring pre-trained model, electronic device and storage medium
CN120218182A (en) Model distillation method, response information generation method and device
CN114819095B (en) Method, device and electronic device for generating business data processing model
US12298956B2 (en) Method and apparatus for generating index, and electronic device
EP4167096A1 (en) Task allocation method and apparatus, electronic device, and computer readable medium
CN113742457A (en) Response processing method and device, electronic equipment and storage medium
CN119960929A (en) Large model loading method, device, electronic device and storage medium
CN119849553A (en) Method, device, equipment, medium and product for generating target sequence based on large model
CN113572679A (en) Method, device, electronic device and storage medium for generating account intimacy
CN116257611B (en) Question-answering model training method, question-answering processing device and storage medium
CN119150991A (en) Knowledge question-answering method and device based on large model and intelligent body
JP7472421B2 (en) Translation method, model training method, apparatus, electronic device and storage medium
CN117421400A (en) Dialogue interaction method, device and electronic device

Legal Events

Date Code Title Description
PB01 Publication
PB01 Publication
SE01 Entry into force of request for substantive examination
SE01 Entry into force of request for substantive examination