CN121223793B - Self-adaptive strategy optimization method and system for body robot based on interactive feedback - Google Patents
Self-adaptive strategy optimization method and system for body robot based on interactive feedbackInfo
- Publication number
- CN121223793B CN121223793B CN202511533027.1A CN202511533027A CN121223793B CN 121223793 B CN121223793 B CN 121223793B CN 202511533027 A CN202511533027 A CN 202511533027A CN 121223793 B CN121223793 B CN 121223793B
- Authority
- CN
- China
- Prior art keywords
- feedback
- strategy
- robot
- standardized
- generate
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Active
Links
Landscapes
- Feedback Control In General (AREA)
Abstract
The invention discloses a self-adaptive strategy optimization method and a self-adaptive strategy optimization system for a robot with a body based on interactive feedback, and relates to the technical field of robots. The method comprises the steps of collecting operation data of the robot with the body in a task execution process, analyzing environment information, user voice, gestures and binary evaluation signals to generate standardized feedback characterization, calculating a reward function based on the environment information and the feedback characterization and generating a standardized reward value, forming a state-feedback joint representation through high-order mapping, generating a candidate strategy set, screening update strategy parameters under the constraint of strategy consistency and execution stability, generating an action sequence for the update parameters to execute operation, and forming consistency indexes and closed-loop data streams through comparison of feedback results and the standardized feedback characterization, so that strategy self-adaptive optimization is achieved. The intelligent generation, execution consistency and task stability improvement of the physical robot strategy are realized through the high-order fusion and closed-loop self-adaptive optimization of multi-mode perception and user interaction feedback.
Description
Technical Field
The invention relates to the technical field of robots, in particular to a self-adaptive strategy optimization method and system for a self-adaptive strategy of a robot body based on interactive feedback.
Background
In practical application, the robot needs to interact with a complex dynamic environment and a user, and the accuracy and the intelligence of task execution are affected by various factors. In the prior art, an autonomous robot generally relies on a single sensory data or a predefined action strategy to accomplish the operation, which strategy generates multiple dependent static rules or low-order feedback control methods. This approach has various problems in complex environments:
on one hand, the conventional strategy cannot fully consider multi-mode feedback generated when the robot performs actions in the environment, so that response to the environment state and user intention is lagged or inaccurate in the strategy execution process. Particularly when force sense, spatial attitude and complex operation tasks are involved, it is difficult for a robot to realize high-precision motion control and continuous execution.
On the other hand, the existing method lacks a systematic closed-loop mechanism in the strategy optimization process, depends on single feedback or simple rewarding evaluation, and cannot form effective self-adaptive adjustment. The non-linear relation of the user interaction intention and the environment constraint is lack of modeling and quantification, so that the problems of strategy deviation, unstable execution, inconsistent behaviors and the like of the robot in the multi-task and multi-environment operation are easy to occur.
In addition, the conventional method has limited capabilities for fusion of multi-modal data and modeling of higher-order dependency relationships. The image, depth, force sense and pose data are often processed by low-order linearity, high-order nonlinear interaction between the environmental state information and user feedback is ignored, and the strategy optimization process lacks refinement, individuation and self-adaption capability.
Therefore, a technical scheme capable of fully utilizing multi-mode sensing data, analyzing user interaction feedback in real time and realizing strategy self-adaptive optimization through closed-loop data flow is needed, so that task execution capacity, feedback consistency and strategy intellectualization level of the robot in a dynamic complex environment are improved.
Disclosure of Invention
Based on the above-mentioned shortcomings of the prior art, the present invention aims to provide an adaptive strategy optimization method and system for an autonomous robot based on interactive feedback, so as to solve the above-mentioned technical problems.
In order to achieve the above purpose, the invention provides the following technical scheme that the self-adaptive strategy optimization method of the robot with body based on interactive feedback comprises the following steps:
Acquiring image data, depth data, force sense data and pose data of the robot body in the task execution process, and generating environment state information, wherein the environment state information comprises the current spatial pose, perception characteristics and stress state of the robot body;
Analyzing and structuring the environment state information, a voice command, a gesture command and a binary evaluation signal from a user to generate a standardized feedback representation, and expressing the joint characteristics of the environment state and the user interaction intention;
calculating a reward function according to the environmental state information and the standardized feedback characterization, generating a standardized reward value, and quantifying the consistency relation between the environmental state information and the standardized feedback characterization;
Performing high-order mapping on the environmental state information and the standardized reward value to generate a state-feedback joint representation, and reflecting the interaction relation of the state constraint and the feedback constraint in the strategy generation process;
Generating a candidate strategy set based on the state-feedback joint representation, and screening and generating updated strategy parameters from the candidate strategy set under constraint conditions, wherein the constraint conditions comprise strategy consistency constraint and execution stability constraint;
Mapping the updated strategy parameters into a robot action sequence, and performing operation and environment interaction by the robot according to the action sequence to generate a feedback execution result;
Comparing a feedback execution result generated after the robot executes the action sequence with a standardized feedback characterization to generate a consistency index, judging whether strategy parameters accord with feedback constraint according to the consistency index and a preset threshold value, generating a closed-loop data stream by the feedback execution result and the consistency index, and supporting self-adaptive optimization of strategies.
The invention is further configured that the collecting data of the robot with body in the task execution process, and generating the environmental status information includes:
In the task execution process, acquiring image data, depth data, force sense data and pose data of the robot body, preprocessing the acquired data, and generating a unified data structure;
performing multi-scale feature extraction on the image data to generate visual perception feature representation;
gradient coding and neighborhood feature mapping are carried out on the depth data, and space geometric feature representation is generated;
Performing joint stress weighting and nonlinear mapping on the force sense data to generate force sense state representation;
position coding and direction coding are carried out on pose data, and a spatial pose representation is generated;
And uniformly mapping and fusing the visual perception feature representation, the space geometric feature representation, the force sense state representation and the space posture representation to generate environment state information, wherein the environment state information comprises the current space posture, the perception feature and the stress state of the robot.
The invention is further arranged such that the generating a normalized feedback representation comprises:
Analyzing and structuring the environmental state information, the voice command, the gesture command and the binary evaluation signal of the user;
performing time sequence feature extraction and nonlinear coding on the voice command to generate voice feature representation;
mapping space and channel characteristics of the gesture instruction to generate gesture characteristic representation;
performing nonlinear embedding and weighting processing on the binary evaluation signals to generate evaluation characteristic representations;
and performing high-order nonlinear joint mapping on the voice feature representation, the gesture feature representation and the evaluation feature representation and the environment state information to generate a standardized feedback representation, and expressing joint features of the environment state and the user interaction intention.
The invention is further arranged such that the generating a normalized prize value comprises:
nonlinear mapping and high-order interaction are carried out on the environmental state information and the standardized feedback characterization, so that a state-feedback coupling characteristic is formed;
Nonlinear high-order integration is carried out on the state-feedback coupling characteristic, and a reward value is generated;
and carrying out normalization processing on the generated reward value to generate a standardized reward value.
The invention is further arranged for generating a state-feedback joint representation comprising:
Performing nonlinear transformation on the environmental state information, and strengthening state constraint characteristics;
nonlinear mapping is carried out on the standardized reward value, and feedback constraint characteristics are enhanced;
and carrying out high-order interaction on the mapped environmental state characteristics and the mapped standardized reward values to form state-feedback combined characteristics.
The invention further provides that the filtering and generating updated strategy parameters from the candidate strategy set under the constraint condition comprises the following steps:
Generating a candidate policy set based on the state-feedback joint representation;
performing policy consistency constraint evaluation on each candidate policy, and quantifying the consistency of the candidate policies in the continuous execution process;
performing stability constraint evaluation on each candidate strategy, and quantifying the robustness of the candidate strategy in a dynamic environment;
And screening strategies meeting the strategy consistency constraint and the execution stability constraint from the candidate strategy set, and generating updated strategy parameters.
The invention further provides that the generating feedback execution result comprises:
converting the updated strategy parameters into a robot action sequence, and executing operation and environment interaction by the robot according to the action sequence;
In the executing process of the robot action, the environment state changes along with the higher-order influence of the action sequence, and the environment state after the action is executed is generated;
And generating a feedback execution result according to the environment state and the action sequence after the action is executed, and quantifying the relation between the action and the environment response.
The method further provides that the determining whether the policy parameter accords with the feedback constraint according to the consistency index and the preset threshold value comprises the following steps:
Performing high-order difference calculation on a feedback execution result generated after the robot executes the action sequence and a standardized feedback characterization, and quantifying nonlinear deviation between the action sequence and the user interaction intention;
Generating a consistency index according to the high-order difference calculation result, and quantifying the consistency degree of strategy parameters and feedback constraints;
Judging whether the strategy parameters accord with feedback constraint according to the consistency index and a preset threshold value, and forming a strategy parameter closed-loop judgment result;
And combining the feedback execution result with the consistency index to generate a closed-loop data stream.
The invention is further arranged that the generating of the closed-loop data stream, the adaptive optimization of the support strategy comprises:
according to the closed-loop data stream, high-order error extraction is carried out on a feedback execution result, a consistency index and a strategy parameter judgment result after the robot executes an action sequence;
Generating nonlinear optimization increment according to the high-order error, and performing self-adaptive adjustment on strategy parameters to form an updated strategy parameter set;
And generating a next round of robot action sequence according to the updated strategy parameters, and realizing the consistency of the action sequence with the user interaction intention and the environment constraint.
The invention also provides an adaptive strategy optimization system of the self-adaptive robot based on the interactive feedback, which comprises:
The multi-mode sensing data acquisition module acquires image data, depth data, force sense data and pose data of the robot body in the task execution process, and generates environment state information, wherein the environment state information comprises the current spatial pose, sensing characteristics and stress state of the robot body;
The interactive feedback analysis module analyzes and constructs the environment state information, the voice command, the gesture command and the binary evaluation signal from the user to generate a standardized feedback representation, and expresses the joint characteristics of the environment state and the interactive intention of the user;
the instant rewarding modeling module calculates a rewarding function according to the environment state information and the standardized feedback characterization, generates a standardized rewarding value and quantifies the consistency relation between the environment state information and the standardized feedback characterization;
The state-feedback joint mapping module is used for carrying out high-order mapping on the environmental state information and the standardized rewarding value to generate a state-feedback joint representation, and reflecting the interaction relation of the state constraint and the feedback constraint in the strategy generation process;
The constraint optimization strategy generation module generates a candidate strategy set based on the state-feedback joint representation, screens and generates updated strategy parameters from the candidate strategy set under constraint conditions, wherein the constraint conditions comprise strategy consistency constraint and execution stability constraint;
the action reasoning and control mapping module maps the updated strategy parameters into a robot action sequence, and the robot performs operation and environment interaction according to the action sequence to generate a feedback execution result;
And the execution feedback consistency verification module compares an execution result after the robot executes the action sequence with the standardized feedback characterization to generate a consistency index, judges whether strategy parameters accord with feedback constraint according to the consistency index and a preset threshold value, generates a closed-loop data stream according to the execution result and the consistency index, and supports self-adaptive optimization of the strategy.
The invention provides a self-adaptive strategy optimization method and system of a self-adaptive robot based on interactive feedback, wherein the method comprises the steps of acquiring image data, depth data, force sense data and pose data of the self-adaptive robot in a task execution process to generate environment state information, wherein the environment state information comprises the current spatial pose, perception characteristics and stress state of the self-adaptive robot, analyzing and structuring the environment state information and voice instructions, gesture instructions and binary evaluation signals from a user to generate standardized feedback characterization, expressing joint characteristics of the environment state and the user interactive intention, calculating a reward function according to the environment state information and the standardized feedback characterization, generating a standardized reward value, quantifying the consistency relation between the environment state information and the standardized feedback characterization, performing high-order mapping on the environment state information and the standardized reward value, generating a state-feedback joint representation, reflecting the interactive relation between the state constraint and the feedback constraint in a strategy generation process, generating a candidate strategy set based on the state-feedback joint representation, screening and generating updated strategy parameters under constraint conditions from the candidate strategy set, wherein the constraint conditions comprise consistency constraint and execution stability constraint, generating a sequence of the updated strategy parameters and the execution stability constraint parameters, generating a self-adaptive strategy performance index according to a preset strategy, comparing the updated strategy sequence with a preset strategy execution sequence with a preset motion index, and a self-adaptive performance index of the self-adaptive strategy optimization system of the self-adaptive robot, generating a self-adaptive strategy optimization system, and a self-adaptive strategy optimization method based on the sequence of the performance index is obtained by comparing the sequence of the performance of the robot with the performance of the performance with a preset sequence, the beneficial effects include:
1. the strategy self-adaptability and execution consistency are improved, namely, the consistency quantification of environment state information and user intention is realized through joint modeling of multi-mode perception data and user interaction feedback, and the strategy parameters are subjected to high-order self-adaptive optimization by combining closed-loop data flow, so that the robot with the body can continuously execute tasks under a dynamic complex environment and maintain the high consistency of strategy execution;
2. The multi-mode information fusion and high-order interaction capability is enhanced, namely, high-order fusion and nonlinear mapping are carried out on image, depth, force sense and pose data, and simultaneously, voice, gesture and binary evaluation signals are structured to form standardized feedback characterization, so that high-order interaction of state constraint and feedback constraint in the generation process of a robot strategy is realized, and the strategy generation is more refined and intelligent;
3. and the task execution stability and the robustness are improved, namely policy consistency constraint and execution stability constraint are introduced when a candidate policy set is generated, the policy parameters are adjusted through high-order nonlinear optimization increment, and the action execution stability and the robustness to environmental disturbance of the robot under a complex environment are obviously improved by combining closed-loop consistency verification.
The foregoing description is only an overview of the present application, and is intended to be implemented in accordance with the teachings of the present application in order that the same may be more clearly understood and to make the same and other objects, features and advantages of the present application more readily apparent.
Drawings
In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly described below, and it is apparent that the drawings in the following description are only some embodiments of the present invention, and other drawings may be obtained according to these drawings without inventive effort for a person skilled in the art. In the drawings:
FIG. 1 is a flow chart of an adaptive strategy optimization method for an autonomous robot based on interactive feedback, according to an exemplary embodiment of the present invention;
fig. 2 is a schematic structural diagram of an adaptive strategy optimization system of an autonomous robot based on interactive feedback according to an exemplary embodiment of the present invention.
Detailed Description
Further advantages and effects of the present invention will become readily apparent to those skilled in the art from the disclosure herein, by referring to the accompanying drawings and the preferred embodiments. The invention may be practiced or carried out in other embodiments that depart from the specific details, and the details of the present description may be modified or varied from the spirit and scope of the present invention. It should be understood that the preferred embodiments are presented by way of illustration only and not by way of limitation.
It should be noted that the illustrations provided in the following embodiments merely illustrate the basic concept of the present invention by way of illustration, and only the components related to the present invention are shown in the drawings and are not drawn according to the number, shape and size of the components in actual implementation, and the form, number and proportion of the components in actual implementation may be arbitrarily changed, and the layout of the components may be more complicated.
In the following description, numerous details are set forth in order to provide a more thorough explanation of embodiments of the present invention, it will be apparent, however, to one skilled in the art that embodiments of the present invention may be practiced without these specific details, in other embodiments, well-known structures and devices are shown in block diagram form, rather than in detail, in order to avoid obscuring the embodiments of the present invention.
Embodiment one:
the self-adaptive strategy optimization method of the robot with body based on the interaction feedback, as shown in fig. 1, comprises the following steps:
Acquiring image data, depth data, force sense data and pose data of the robot body in the task execution process, and generating environment state information, wherein the environment state information comprises the current spatial pose, perception characteristics and stress state of the robot body;
Analyzing and structuring the environment state information, a voice command, a gesture command and a binary evaluation signal from a user to generate a standardized feedback representation, and expressing the joint characteristics of the environment state and the user interaction intention;
calculating a reward function according to the environmental state information and the standardized feedback characterization, generating a standardized reward value, and quantifying the consistency relation between the environmental state information and the standardized feedback characterization;
Performing high-order mapping on the environmental state information and the standardized reward value to generate a state-feedback joint representation, and reflecting the interaction relation of the state constraint and the feedback constraint in the strategy generation process;
Generating a candidate strategy set based on the state-feedback joint representation, and screening and generating updated strategy parameters from the candidate strategy set under constraint conditions, wherein the constraint conditions comprise strategy consistency constraint and execution stability constraint;
Mapping the updated strategy parameters into a robot action sequence, and performing operation and environment interaction by the robot according to the action sequence to generate a feedback execution result;
Comparing an execution result after the robot executes the action sequence with the standardized feedback characterization to generate a consistency index, judging whether strategy parameters accord with feedback constraint according to the consistency index and a preset threshold value, generating a closed-loop data stream by the execution result and the consistency index, and supporting self-adaptive optimization of strategies.
The invention is further configured that the collecting data of the robot with body in the task execution process, and generating the environmental status information includes:
In the task execution process, collecting image data, depth data, force sense data and pose data of the robot, preprocessing the collected data to generate a unified data structure, and specifically, setting the robot to perform time steps The data of each mode is: , wherein, Is an image matrix, is acquired by a robot camera, and has the size ofRepresenting the spatial visual information, and the information of the spatial visual information,Is a depth data array, the size isDistance information indicating each point of the robot and the environment,For force sense data sequence, the length isIndicating the stress state of each joint or end effector of the robot,Is a pose matrix with the size ofWhereinFor the number of joints, 6 represents the spatial position and orientation information of each joint, and the data are subjected to unified preprocessing to ensure that the input formats are unified and avoid the influence of the cross-modal difference on the subsequent mapping;
extracting multi-scale features of image data to generate visual perception feature representation, in particular to an image matrix And (3) multi-scale convolution feature extraction: , wherein, For the number of layers of the convolution scale,Is the firstLayer convolution operations, including convolution kernels, nonlinear activation and normalization operations,Is the firstThe layer convolves the weight matrix with the same,Representing element-wise multiplication, for fusing features between convolutions,Representing a matrix for image features, the operation extracting local edge, global texture and shape information at multiple scales to form a visual feature matrix;
Gradient encoding and neighborhood feature mapping are carried out on depth data to generate space geometrical feature representation, and the depth data is specificConversion to high dimensional spatial features: , wherein, For the depth gradient operator, depth variation information is calculated,For the depth local neighborhood weighting function, depth neighborhood information is encoded,For a nonlinear mapping function, two-dimensional depth data is embedded into a high-dimensional feature space,Capturing the geometric information of the environment space through gradient operators and neighborhood coding for representing the depth characteristics, and converting the two-dimensional depth information into high-dimensional characteristics capable of participating in state characterization;
The force sense data is subjected to joint force weighting and nonlinear mapping to generate force sense state representation, and specifically, the force sense sequence is subjected to force sense Constructing a force sense state representation: , wherein, For the number of force sensing sensors,The weight coefficient is used for representing the force sense influence of different joints,For the force sense nonlinear mapping function, the original force sense numerical value is mapped to the feature space,For force sense characteristic representation, weighting nonlinear mapping is carried out on stress information of each joint or actuator to form a unified force sense characteristic matrix which reflects the stress state of the robot;
position coding and direction coding are carried out on pose data to generate space pose representation, and in particular, pose matrix A spatial mapping is performed and the spatial mapping is performed,, wherein,For joint position coding functions, spatial pose information is captured,For the joint orientation encoding function, spatial orientation information is captured,The representation is fused element by element,The position and orientation codes are fused to obtain complete space gesture features for the gesture feature matrix, and a gesture base is provided for the environment state information;
The visual perception feature representation, the space geometrical feature representation, the force sense state representation and the space posture representation are subjected to unified mapping and high-order fusion to generate environment state information, wherein the environment state information comprises the current space posture, the perception feature and the stress state of the robot, and in particular, the multi-mode feature is obtained And performing unified mapping to generate environment state information: , wherein, The characteristic stitching operation is represented as such,For high-order nonlinear mapping operators, multi-modal features are embedded into a unified state space,The system is environment state information and comprises space gestures, perception features and stress states, vision, depth, force sense and pose information are unified into high-dimensional environment state representation through feature splicing and high-order mapping, and complete environment state information of coverage space, stress and perception is formed through multi-mode fusion of vision, depth, force sense and pose.
The invention is further arranged such that the generating a normalized feedback representation comprises:
analyzing and structuring the environment state information, the voice command, the gesture command and the binary evaluation signal of the user, particularly, in the time step The environmental status information of (a) isThe user interaction signal is: , wherein, Is a voice instruction vector sequence with the length of,The feature matrix is a gesture instruction feature matrix with the size of,In order to be of a height, the height,In the form of a width, the width,In order to provide the number of channels,For binary evaluation of signal sequences, the length is;
Extracting time sequence features and non-linear coding to phonetic instruction to generate phonetic feature representation, and specifically, phonetic vector sequenceThe speech features are generated by a non-linear embedding mapping,, wherein,Is the firstFeature transformation functions of the individual speech vectors, extracting timing semantic information,For a nonlinear attention weighting function, weights are assigned according to speech importance,Representing an element-by-element multiplication,Capturing key information related to task execution in a voice instruction through element-by-element nonlinear weighting for voice characteristic representation;
mapping the space and channel characteristics of the gesture command to generate a gesture characteristic representation, and specifically, mapping the gesture matrix Performing space one-channel convolution mapping: , wherein, Extracting motion direction and amplitude characteristics for the gesture local nonlinear mapping function,For the convolution weight matrix, the importance of different spatial locations and channels is adjusted,For gesture feature representation, capturing key information of gestures on a space and a channel through space-channel weighted mapping, and realizing representation of associated features of the gestures and environmental states;
Performing nonlinear embedding and weighting processing on the binary evaluation signals to generate evaluation characteristic representations, specifically, performing nonlinear embedding and weighting processing on binary evaluation sequences Constructing a nonlinear response map: , wherein, Is a non-linear embedded function of binary signal, willOr (b)The evaluation is mapped to a high-dimensional feature space,As the weight coefficient, the importance of each evaluation time is reflected,For the binary feedback feature, mapping the simple binary evaluation into a high-dimensional feature so as to be fused with the voice and gesture feature, and reflecting the guiding function of the evaluation on policy optimization;
Performing high-order nonlinear joint mapping on the voice characteristic representation, the gesture characteristic representation, the evaluation characteristic representation and the environment state information to generate a standardized feedback representation, expressing joint characteristics of the environment state and the user interaction intention, and particularly, the environment state information High-order nonlinear joint mapping is carried out on the voice, the gesture and the binary evaluation characteristics: , wherein, The characteristic stitching operation is represented as such,Representing element-by-element interaction mapping, preserving the coupling relationship between state and feedback,For high-order nonlinear mapping operators, multi-modal features are embedded into a unified feature space,For standardized feedback characterization, the joint characteristics of the environment state and the user interaction intention are reflected, and the environment state and the multi-mode user interaction characteristics are coupled through element-by-element interaction and high-order mapping to form the joint characterization which can directly participate in policy optimization.
The invention is further arranged such that the generating a normalized prize value comprises:
the method comprises the steps of carrying out nonlinear mapping and high-order interaction on environment state information and standardized feedback characterization to form a state-feedback coupling characteristic, specifically carrying out element-by-element interaction and high-order mapping on the environment state and the feedback characteristic to generate the state-feedback coupling characteristic: , wherein, The key space gesture, the perception characteristic and the stress state are enhanced for the nonlinear mapping function of the environment state,To feedback the key information characterizing the nonlinear mapping function, enhance the user's intent,The characteristic stitching operation is represented as such,Representing element-by-element interactions, capturing the coupling relationship between state and feedback,Output state-feedback coupling feature for high order nonlinear mapping operatorEmbedding the combined information of the environment state and the user feedback into a unified feature space through element-by-element interaction and high-order mapping, quantifying the consistency potential, realizing the deep fusion of the environment state and the user feedback through state-feedback high-order coupling mapping, and ensuring that the reward function is highly sensitive to key features;
nonlinear high-order integration is performed on the state-feedback coupling characteristic to generate a reward value, and the state-feedback coupling characteristic is specifically performed Generating a prize value by performing a nonlinear high order integration: , wherein, Is the firstDimension of environmental status, 1State at the feedback feature dimension-a feedback coupling feature,For the number of dimensions of the environmental state,For the feedback of the number of feature dimensions,For a nonlinear index, control the higher order sensitivity,The importance of a feedback combination of different states is reflected as a weight matrix,Mapping the result of the integration to a prize value as a non-linear mapping functionThe multi-dimensional information is converted into a single rewarding index through high-order integral and weighted mapping of the state-feedback coupling characteristic, consistency between the environment state and user feedback is captured, and a nonlinear coupling relation is extracted from a multi-dimensional state-feedback space through high-order consistency integral, so that the problem that complex interaction information is lost by a simple weighting method is avoided;
normalizing the generated reward value to generate a normalized reward value, specifically, generating the reward value Mapping to a standardized interval to form a rewarding index capable of directly participating in policy optimization: , wherein, In order to normalize the mapping operator,Is small constant, prevents the abnormality of zero removal,The method is used for standardized reward values for subsequent state-feedback joint mapping and strategy generation, the higher-order reward values are mapped to a unified scale, stability and comparability of the reward values in the optimization process are guaranteed, the reward values have the unified numerical scale through normalized reward construction, and the reward values can be stably applied to strategy optimization under different tasks and different feedback conditions.
The invention is further arranged for generating a state-feedback joint representation comprising:
The method comprises the steps of carrying out nonlinear transformation on environment state information, strengthening state constraint characteristics, specifically, carrying out high-order nonlinear mapping on the environment state information, and highlighting the state constraint characteristics: , wherein, For the number of dimensions of the environmental state,For the number of state sub-feature dimensions,The importance of different state dimensions and sub-features is reflected for the weight coefficient matrix,Is a nonlinear exponential matrix, enhances the high-order sensitivity,For the mapped environment state characteristic representation, constraint information of the environment state in strategy generation is strengthened through exponential weighting and high-order nonlinear mapping;
non-linear mapping is carried out on the standardized rewards value, the feedback constraint characteristic is enhanced, and specifically, the standardized rewards are obtained High order nonlinear mapping is performed to emphasize feedback constraint features: , wherein, In order to map the number of terms,To weight the coefficients, reflect the contributions of the different higher-order terms to the joint representation,Is nonlinear index, controls the sensitivity of high-order feedback,Converting the standardized rewards into high-order feedback constraint features through nonlinear mapping for mapped rewards features, and coupling the high-order feedback constraint features with state features in a joint representation;
High-order interaction is carried out on the mapped environmental state characteristics and the mapped standardized rewards value to form state-feedback combined characteristics, and the environmental state characteristics are specifically With rewards featurePerforming element-by-element high-order interaction and nonlinear mapping: , wherein, The representation is an element-by-element interaction,For the bonus feature to be a non-linear spread function,The weight tensor is combined for the state-one prize,As a non-linear exponential tensor,For the state-feedback joint representation, the state constraint and the feedback constraint are fully coupled through element-by-element interaction and high-order nonlinear mapping, so that the joint representation is realized, and high-order constraint information is provided for strategy generation.
The invention further provides that the filtering and generating updated strategy parameters from the candidate strategy set under the constraint condition comprises the following steps:
Generating a candidate strategy set based on the state-feedback joint representation, and specifically, constructing a multidimensional strategy generation operator based on the state-feedback joint representation: , wherein, The number of dimensions that are jointly represented for the state-feedback,For the number of joint feature sub-dimensions,Generating a weight matrix for the strategies, adjusting the contribution of the joint features to different candidate strategies,For a nonlinear exponential matrix, control the higher order sensitivity,Random perturbation parameters generated for the strategy, for increasing candidate strategy diversity,Is the firstCandidate strategies, definitionGenerating diversified candidate strategies for the candidate strategy set through the high-order nonlinear combined joint representation features, and ensuring that the strategy set covers a potential optimal solution space;
carrying out policy consistency constraint evaluation on each candidate policy, and quantifying the consistency of the candidate policies in the continuous execution process: , wherein, For the dimension of the policy action,In order to be a consistent weight matrix,Is a nonlinear index matrix, enhances the sensitivity of high-order differences,For the strategy consistency index, quantifying the consistency of the strategy in the continuous action sequence, and ensuring that the generated strategy cannot cause action jump or instability;
Performing execution stability constraint evaluation on each candidate strategy, and quantifying the robustness of the candidate strategy in a dynamic environment: , wherein, As a dimension of the action,As a factor of the weight of the stability,Is a nonlinear index, reflects the sensitivity of high-order disturbance,Is an environment disturbance simulation value,To perform the stability index, quantifying the robustness of the strategy in a dynamic environment;
Screening policies meeting the policy consistency constraint and the execution stability constraint from the candidate policy set, generating updated policy parameters, specifically, screening policies meeting the constraint conditions from the candidate policy set, and generating updated policy parameters: , wherein, And is also provided withTo meet the policy index set of consistency and stability thresholds,To filter the policy weights, reflect the policy contribution,Is nonlinear index, enhances the high-order strategy characteristic,And generating final strategy parameters for the updated strategy parameters by high-order weighted combination of candidate strategies, and ensuring that the strategies are optimized under the constraint of consistency and stability.
The invention further provides that the generating feedback execution result comprises:
The method comprises the steps of converting updated strategy parameters into a robot action sequence, and enabling a robot to execute operation and environment interaction according to the action sequence, specifically, constructing a high-order action inference operator based on the updated strategy parameters to generate the action sequence: , wherein, The dimensions of the policy parameters and the sub-dimensions are respectively,Generating a weight matrix for the action, controlling the contribution of the policy parameters to the action generation,Is a nonlinear exponential matrix, enhances the high-order sensitivity,Generating random disturbance parameters for the motion, for increasing the motion diversity,For the environmental adaptation correction, reflecting the environmental impact of the action during execution,Is the robot NoIndividual actions, generating a complete sequence of actionsGenerating an action sequence through high-order nonlinear combination strategy parameters, and simultaneously adding random disturbance and environmental correction to realize action diversity and environmental adaptability;
in the process of executing the robot action, the environment state changes along with the higher-order influence of the action sequence to generate the environment state after the action is executed, specifically, the robot follows the action sequence Operating the environment, and updating the environment state: , wherein, The function is updated for the state of the environment,For the higher order impact operator of the action on the environmental state, reflecting the nonlinear change of the action execution on the environmental state,For the environment state after the action is executed, applying high-order nonlinear influence on the environment through the action sequence to form environment response associated with the action sequence;
generating a feedback execution result according to the environment state after the action is executed and the action sequence, quantifying the relation between the action and the environment response, and specifically generating the feedback execution result according to the environment state after the action is executed: , wherein, For feeding back the weight matrix, the importance of the corresponding relation between the environment state and the action is reflected,Is nonlinear index, enhances the sensitivity of high-order difference,And the feedback execution result is formed by quantifying the difference between the action execution and the environment response through a high-order nonlinear function, so as to provide a data basis for closed loop self-adaption.
The method further provides that the determining whether the policy parameter accords with the feedback constraint according to the consistency index and the preset threshold value comprises the following steps:
The method comprises the steps of carrying out high-order difference calculation on a feedback execution result generated after the robot executes an action sequence and a standardized feedback representation, quantifying nonlinear deviation between the action sequence and user interaction intention, and specifically carrying out nonlinear high-order difference quantification on the feedback execution result and the standardized feedback representation: , wherein, The dimensions of the action are represented and,Representing the dimension of the state of the environment,Is a nonlinear index matrix, enhances the sensitivity of high-order differences,As a weight matrix, reflects the importance of different actions and environmental dimensions,Is the firstThe action dimension is at the firstCapturing a small deviation and a complex interaction effect through the high-order difference in the individual environment dimensions and the difference between the execution result and the user expectation by the high-order nonlinear quantitative feedback;
Generating a consistency index according to the calculation result of the high-order difference, and quantifying the consistency degree of strategy parameters and feedback constraint Performing multidimensional nonlinear synthesis to generate a consistency index: , wherein, For the number of dimensions of the motion,For the number of dimensions of the environmental state,Is a second-order nonlinear exponential matrix, enhances the sensitivity to high-order coupling,Adjusting the contribution of each dimension to the consistency index for the weight matrix,For quantifying a high-order index of strategy parameter and feedback constraint consistency, the multidimensional nonlinear combination ensures that the consistency index fully reflects a complex interaction relation between an action sequence and a user expectation;
judging whether the strategy parameters meet feedback constraint according to the consistency index and a preset threshold value to form a strategy parameter closed-loop judgment result: , wherein, In order to meet the policy parameters of the feedback constraint,In order to trigger parameters for policy adjustment or optimization,For consistency judgment threshold, strategy parameter closed-loop judgment is realized by comparing the high-order consistency index with the threshold, so that the action sequence is ensured to meet the user intention and the environmental requirement;
combining the feedback execution result with the consistency index to generate a closed-loop data stream, and specifically, combining the feedback execution result with the consistency index to form the closed-loop data stream: , wherein, Is a nonlinear index matrix, enhances the sensitivity of closed-loop data to higher-order consistency,The method is a closed-loop data stream, and is used for the next round of strategy self-adaptive optimization, and the closed-loop data stream reflects the relation between strategy execution and user expectations in a high-order mode.
The invention is further arranged that the generating of the closed-loop data stream, the adaptive optimization of the support strategy comprises:
according to the closed-loop data stream, high-order error extraction is carried out on a feedback execution result, a consistency index and a strategy parameter judgment result after the robot executes an action sequence; specifically, policy bias information is extracted from a closed-loop data stream to form an error matrix: , wherein, As a dimension of the action,As a dimension of the environment,Is the first in the closed loop data streamThe action dimension is at the firstThe integrated feedback information in the individual environmental dimensions,As a function of the current policy parameters,Is a nonlinear index matrix, enhances sensitivity to higher order deviations,For the weight matrix, the contributions of different actions and environmental dimensions in the optimization are adjusted,The optimization sensitivity and precision are ensured by amplifying the small deviation in a high-order nonlinear way for the quantization strategy parameter deviating from the high-order error of the closed-loop data;
Generating nonlinear optimization increment according to the high-order error, performing self-adaptive adjustment on strategy parameters to form an updated strategy parameter set, and specifically, calculating the nonlinear optimization increment according to the high-order error: , wherein, Taking into account the higher-order dependency between motion and environmental dimensions as a nonlinear coupling function,Controlling the adjustment amplitude, preventing overshoot or oscillation,For the update amplitude of the action parameters in the round of optimization, nonlinear coupling ensures that adjustment not only depends on the dimension error, but also can respond to feedback of other dimensions, and global self-adaptive optimization is realized;
Generating a next round of robot action sequence according to the updated strategy parameters, realizing the consistency of the action sequence with the user interaction intention and the environment constraint, and specifically generating the next round of strategy parameters through closed-loop data flow and optimized increment: When the policy parameters are determined to be valid At the time of incrementThe adjustment amplitude is smaller, the stability is kept, and when the strategy parameters are judged to be invalidAt the time of incrementThe adjustment amplitude is increased, the parameters deviating from the feedback constraint are corrected rapidly, and the strategy parameters are optimized in a continuous self-adaptive mode through closed-loop error quantization and nonlinear increment adjustment, so that the action sequence is ensured to approach the user interaction intention gradually.
Embodiment two:
Referring to fig. 2, the exemplary adaptive strategy optimization system for an autonomous robot based on interactive feedback includes:
The multi-mode sensing data acquisition module acquires image data, depth data, force sense data and pose data of the robot body in the task execution process, and generates environment state information, wherein the environment state information comprises the current spatial pose, sensing characteristics and stress state of the robot body;
The interactive feedback analysis module analyzes and constructs the environment state information, the voice command, the gesture command and the binary evaluation signal from the user to generate a standardized feedback representation, and expresses the joint characteristics of the environment state and the interactive intention of the user;
the instant rewarding modeling module calculates a rewarding function according to the environment state information and the standardized feedback characterization, generates a standardized rewarding value and quantifies the consistency relation between the environment state information and the standardized feedback characterization;
The state-feedback joint mapping module is used for carrying out high-order mapping on the environmental state information and the standardized rewarding value to generate a state-feedback joint representation, and reflecting the interaction relation of the state constraint and the feedback constraint in the strategy generation process;
The constraint optimization strategy generation module generates a candidate strategy set based on the state-feedback joint representation, screens and generates updated strategy parameters from the candidate strategy set under constraint conditions, wherein the constraint conditions comprise strategy consistency constraint and execution stability constraint;
the action reasoning and control mapping module maps the updated strategy parameters into a robot action sequence, and the robot performs operation and environment interaction according to the action sequence to generate a feedback execution result;
And the execution feedback consistency verification module compares an execution result after the robot executes the action sequence with the standardized feedback characterization to generate a consistency index, judges whether strategy parameters accord with feedback constraint according to the consistency index and a preset threshold value, generates a closed-loop data stream according to the execution result and the consistency index, and supports self-adaptive optimization of the strategy.
It should be noted that, the self-adaptive strategy optimization system of the self-adaptive strategy based on the interactive feedback provided by the above embodiment and the self-adaptive strategy optimization method of the self-adaptive strategy based on the interactive feedback provided by the above embodiment belong to the same conception, the specific manner in which the respective modules and units perform operations has been described in detail in the method embodiments, and will not be described here again. In practical application, the self-adaptive strategy optimization system of the self-adaptive robot based on the interactive feedback provided by the embodiment can distribute the functions by different functional modules according to the needs, namely, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above, and the self-adaptive strategy optimization system is not limited in this place.
The foregoing is merely illustrative of the present application, and the present application is not limited thereto, and any person skilled in the art will readily recognize that variations or substitutions are within the scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims (6)
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202511533027.1A CN121223793B (en) | 2025-10-24 | 2025-10-24 | Self-adaptive strategy optimization method and system for body robot based on interactive feedback |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202511533027.1A CN121223793B (en) | 2025-10-24 | 2025-10-24 | Self-adaptive strategy optimization method and system for body robot based on interactive feedback |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| CN121223793A CN121223793A (en) | 2025-12-30 |
| CN121223793B true CN121223793B (en) | 2026-04-10 |
Family
ID=98159981
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| CN202511533027.1A Active CN121223793B (en) | 2025-10-24 | 2025-10-24 | Self-adaptive strategy optimization method and system for body robot based on interactive feedback |
Country Status (1)
| Country | Link |
|---|---|
| CN (1) | CN121223793B (en) |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114995657A (en) * | 2022-07-18 | 2022-09-02 | 湖南大学 | A multi-modal fusion natural interaction method, system and medium for intelligent robot |
| CN119260754A (en) * | 2024-09-18 | 2025-01-07 | 台州爱鑫智能科技有限公司 | An information interaction method for embodied intelligent robots |
Family Cites Families (11)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US9274607B2 (en) * | 2013-03-15 | 2016-03-01 | Bruno Delean | Authenticating a user using hand gesture |
| EP4017689B1 (en) * | 2019-09-30 | 2024-10-30 | Siemens Aktiengesellschaft | Robotics control system and method for training said robotics control system |
| CN117549310A (en) * | 2023-12-28 | 2024-02-13 | 亿嘉和科技股份有限公司 | General system of intelligent robot with body, construction method and use method |
| US20250292097A1 (en) * | 2024-03-14 | 2025-09-18 | International Business Machines Corporation | Optimizing grayscale release strategies based on multiple objectives and constraints |
| CN118567477A (en) * | 2024-05-31 | 2024-08-30 | 广东铂锶特科技有限公司 | Adaptive augmented reality diversified interaction method and system for complex scenes |
| CN120095830B (en) * | 2025-04-22 | 2025-12-23 | 节卡未来科技(上海)有限公司 | Robot control method, computer device, and computer-readable storage medium |
| CN120630670A (en) * | 2025-05-07 | 2025-09-12 | 南京康龙威科技实业有限公司 | A robot gait training method and system based on reinforcement learning |
| CN120542396A (en) * | 2025-05-20 | 2025-08-26 | 平安科技(深圳)有限公司 | Natural language driven table processing method, device, equipment and medium |
| CN120560339A (en) * | 2025-05-27 | 2025-08-29 | 上海大学 | A UAV swarm collaborative coverage method with dynamic obstacle avoidance |
| CN120807320A (en) * | 2025-05-30 | 2025-10-17 | 郑州大学第五附属医院 | Low-count PET image quality enhancement method and system based on fusion multi-input cycle consistency generation countermeasure network |
| CN120773064B (en) * | 2025-09-02 | 2026-04-10 | 深圳乾海格致科技有限公司 | Self-adaptive behavior mode adjusting system of intelligent robot with body |
-
2025
- 2025-10-24 CN CN202511533027.1A patent/CN121223793B/en active Active
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114995657A (en) * | 2022-07-18 | 2022-09-02 | 湖南大学 | A multi-modal fusion natural interaction method, system and medium for intelligent robot |
| CN119260754A (en) * | 2024-09-18 | 2025-01-07 | 台州爱鑫智能科技有限公司 | An information interaction method for embodied intelligent robots |
Also Published As
| Publication number | Publication date |
|---|---|
| CN121223793A (en) | 2025-12-30 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Hewing et al. | Learning-based model predictive control: Toward safe learning in control | |
| CN111433689B (en) | Generation of control systems for target systems | |
| KR102239186B1 (en) | System and method for automatic control of robot manipulator based on artificial intelligence | |
| CN117709185A (en) | Digital twin architecture design method for process industry | |
| EP3614314B1 (en) | Method and apparatus for generating chemical structure using neural network | |
| CN120688015A (en) | AI interactive data processing system based on multimodal perception and dynamic decision-making | |
| CN120552071A (en) | Real-time collaborative decision-making method for humanoid robots based on multimodal perception fusion | |
| CN116502069B (en) | Haptic time sequence signal identification method based on deep learning | |
| JP2020155010A (en) | Neural network model contraction device | |
| CN119974028B (en) | Self-adaptive gravity center adjusting method and system for transfer robot | |
| CN114091554A (en) | Training set processing method and device | |
| CN119442144A (en) | A reinforcement learning modeling method and device integrating multi-source heterogeneous data | |
| CN117875407A (en) | A multimodal continuous learning method, device, equipment and storage medium | |
| CN119681909A (en) | Robot control method based on multi-modal large model | |
| CN118940220A (en) | A multimodal industrial data fusion method and system for discrete manufacturing | |
| Wu et al. | A framework of improving human demonstration efficiency for goal-directed robot skill learning | |
| CN121223793B (en) | Self-adaptive strategy optimization method and system for body robot based on interactive feedback | |
| CN120388565B (en) | Voice interaction method and system based on 3D (three-dimensional) virtual | |
| CN119407779A (en) | Flexible robot control method and system based on multimodal data and heuristic graph search | |
| JP7623619B2 (en) | Learning device, estimation device, learning method, estimation method, and program | |
| Grimble et al. | Non-linear predictive control for manufacturing and robotic applications | |
| Tang | Research on Computer-Aided Architectural Design Optimization Based on Building Information Modeling (BIM) Technology | |
| CN121267944B (en) | Robot motion control strategy network training method and device based on state representation learning | |
| CN121515213A (en) | An Adaptive Control Method and System for Industrial Robots Based on Multimodal Sensor Fusion | |
| Iannotti | Development and Assessment of Structural Theories Using Machine Learning Techniques |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PB01 | Publication | ||
| PB01 | Publication | ||
| SE01 | Entry into force of request for substantive examination | ||
| SE01 | Entry into force of request for substantive examination | ||
| GR01 | Patent grant | ||
| GR01 | Patent grant |